We’re still carrying next-generation intelligence inside a last-generation interaction paradigm
09/09/2026 by Mary Li

Forty-six teams. Three-minute demos. One Agentic AI Hackathon.

When the judges sat back on the first judging round of the Alibaba Cloud x Atlas Agentic AI Hackathon Singapore 2026 to get to the shortlist of 10 submissions, the first thing they agreed on wasn’t “how strong the tech was this year.”

It was this: Agent capability has already outrun our imagination for what to do with it.

Search, ticketing, rebooking, coordinating a multi-city group trip – tasks that used to demand deep domain expertise – were pulled off by almost every team within days. That shows an entire cohort of ordinary developers clearing the same bar at once, with AI doing the heavy lifting.

And yet, across all 46 projects, almost none broke out of the old product mold. The capability underneath was brand new but the interface on top still looked a decade old.

That’s the key point: the question isn’t whether AI is smart enough anymore. It’s whether we’ve actually figured out what to do with that intelligence and how to work with it.

The hackathon drew 461 applications, shortlisted 150 teams, and ended with 46 completed submissions and demos. The first round of judges from Alibaba Cloud Qoder, Alibaba Cloud Solution Architecture, AchieveGo.ai, and Atlas spent days watching demos, reading code, and checking evidence before landing on a Top 10.

Another panel of judges, which will include travel tech leaders with domain expertise, will bring the list of 10 to three finalists, who will pitch at the WiT Bootcamp (September 30), held in conjunction with WiT Singapore.

What was interesting was how the three judges, coming at this from three different angles, all ran into the same wall.

 

Honesty was rarer than perfection

Jennifer, co-founder of AchieveGo, led the AI judging and calibration for this hackathon. The first thing she brought up wasn’t technical. It was honesty.

Some teams said this, plainly, in their demos: “This step is simulated.”

A three-minute demo shapes a judge’s entire first impression. The easy move is to make everything look like it already works. But Jennifer found that the teams willing to clearly mark the edges of what they’d actually built were, more often than not, the strongest ones.

Because the first thing a real Agent needs to prove isn’t that it can do everything – it’s that it knows exactly what it has done, and what it hasn’t.

That honesty tested the judges too. Machine judges and human judges disagreed. Different models disagreed with each other. So “who’s right” stopped being the useful question. The useful question became: why the disagreement? Was the evidence thin, or was the rubric itself unclear?

Jennifer put it this way: “Every disagreement between the machine and human judges taught us something – often, what it taught us was about our own standards.”

Her conclusion: “There’s still a long way between an impressive three-minute demo and an Agent you’d actually trust to spend real money on your behalf. Nobody has crossed that gap yet.”

 

 

When agents start taking on real responsibility

Hongwei, Atlas’s AI Application Engineer sits closest to the participants on Atlas’s side — he configured Qoder for every official team and handled the technical onboarding.

What surprised him most wasn’t the tech. It was how much the participants’ understanding of “what an Agent should do” had quietly leveled up.

In the earlier era of AI travel tools, the ask was simple: find me a flight, compare hotels, build me an itinerary. AI gave advice; humans acted on it.

This time, a number of teams weren’t satisfied with just giving advice. They made the Agent actually do the thing: completing bookings and ticketing, handling the specific needs of medical travel, coordinating a group departing from different cities to converge at one point, even auto-rebuilding an entire itinerary after a flight got cancelled.

Hongwei’s read: “What people are thinking about now is how a real product should actually work.”

That’s the shift – once AI moves from “giving advice” to “taking action”, it’s no longer just a technical problem. Behind every decision could be a real ticket, real money, a meeting that can’t be missed.

So, the next question he asked: How do you give an Agent the ability to act, while also teaching it when to stop?

When should it say, “something’s changed here”? When should it explain “here’s what happens if I proceed”? And when does the decision have to go back to a human – “this one’s yours to make”?

The stronger the capability, the more the boundaries matter. There’s a long way to go –but some teams are already walking that road seriously.

 

The most uncomfortable finding: We built something we don’t yet know how to use

Lex, CTO Advisor at Atlas, who was involved on Atlas’s side from designing the competition rules to verifying evidence in judging, kept coming back to a more fundamental question:

The capability is already running. Do we actually know what to do with it?

Two things surprised him.

First, how fast Agent engineering has spread. This year’s challenge plugged into a full-chain airline fulfillment API – not a low bar technically – yet almost every team got search, ticketing, and rebooking working end-to-end in a very short window. “Anyone can build with Agents now” isn’t an aspiration anymore in this competition. It’s a fact.

Second, how well AI grasped a specialized domain. Airline ticketing is dense with hidden knowledge and jargon, but almost no team showed meaningful gaps in the underlying concepts. Domain fluency that used to take months to build was compressed into a matter of hours.

By that logic, flattening the barrier this far should have triggered an explosion of new product forms.

Instead, Lex found the opposite:

Almost no team broke free of the traditional software interaction experience.

Most projects still shipped as conventional desktop or mobile interfaces – some even piled on more forms and dense information panels. None of them showed a truly dynamic interaction that shifted in real time with user intent and context. The Agent’s capability got neatly locked inside a traditional UI, instead of becoming the interaction itself.

This isn’t a shortfall on any team’s part – if anything, it’s precisely because everyone got so good, so fast, at making the AI do the right thing, that this gap became so visible: we’re already holding a next-generation engine but still steering it with last-generation controls.

As Lex puts it: “We’re still carrying next-generation intelligence inside a last-generation interaction paradigm.”

That’s not a knock on the participants – it’s a description of where the entire industry is standing right now. Nobody has crossed this line yet, which means whoever first figures out what “Agent-native interaction” actually looks like gets the next real ticket in.

 

 

“Everyone’s imagination has been given a pair of wings”

Robin Liu, a Cloud Solution Architect at Alibaba Cloud Singapore, watched all 46 demos start to finish.

His read came from a different angle. AI has spent the last few years proving, again and again, that it “can do a lot of things.” But one question has never been fully answered: what should we actually use that capability for?

It used to be that turning a travel product idea into reality meant first finding an engineer, building a system, standing up infrastructure — before you ever got to test whether the idea worked. That barrier is disappearing fast.

Robin summed up the shift in one line: “People from any background can now build, and everyone’s imagination has been given a pair of wings.”

Across the 46 projects, he saw more than polished product design and rehearsed demos – he saw a lot of people bringing AI into something they genuinely loved: travel. Endlessly re-searching an itinerary, freezing up when something goes wrong mid-trip, refreshing flight prices at 2 am – the burdens that used to fall entirely on the traveler might, one day, no longer have to.

What excited him even more: some teams were already thinking further ahead – asking what the entire travel experience becomes if AI turns into an actual actor within the journey, not just an advisor on the side.

That exceeded his expectations going in.

Robin had also given the participants one warning before the hackathon started: don’t let AI vibe coding spiral out of control. A lower barrier to entry doesn’t excuse dropping engineering discipline – security, compliance, and code quality don’t disappear in the Agent era; if anything, they matter more. He was glad to see many teams write security and compliance requirements directly into their own rules and hold the AI to them.

“I am glad, and proud, to have seen both.”

What he saw wasn’t just the speed AI brings – it was developers learning how to actually steer that speed.

 

AI for consistency. Humans for judgment.

There was one more interesting layer to this competition: the judging process itself tried to take AI seriously, too.

The first round brought in an AI Judge – machines reading submissions against a shared rubric, checking code and evidence, and scoring – with humans cross-checking afterward. It turned out to be far messier than expected. AI got things wrong. Humans got things wrong. AI models disagreed with each other. Humans and AI disagreed even more.

So “who’s right” wasn’t the valuable question here either. What mattered was: why the gap? Was the evidence insufficient, was the rubric unclear, or were the human judges carrying their own biases and experience into the room?

We kept recalibrating the algorithm and kept recalibrating ourselves. What we landed on was a principle: AI for consistency. Humans for judgment.

Let AI help us read, compare, and gather evidence as consistently as possible. But leave the final call to people.

In a way, this is the same problem Lex identified from another angle – AI is now capable enough that we have to seriously rethink where humans actually belong inside this new system.

 

461 → 150 → 46 → Top 10

From idea to execution, this hackathon took about a month. 461 applications, 150 teams admitted, 46 finished projects, and now a freshly minted Top 10.

Looking back at these submissions, one thing feels certain: the future of travel Agents isn’t just about making AI smarter. It’s about someone actually figuring out – now that the capability is here – what the interaction should look like, where the boundaries should sit, and when a decision needs to go back to a human.

That’s not something any single team failed to do. It’s a door the entire industry hasn’t walked through yet.

Robin left us with a line we keep coming back to: “The pioneers of this industry once walked alone, carrying a lamp through the dark. In the age of AI, many more people have joined that road.”

Pioneers used to carry that lamp alone through the dark. Maybe the biggest shift of the AI era is that far more people now have the ability to join that walk and these 46 teams just took their first steps onto it.

This week, our wider judging panel moves into the next stage – the “real human” review: watching the demos again, looking deeper into the code, and asking whether what’s on screen is not only technically interesting, but genuinely useful, executable, and meaningful for the future of travel.

That panel: Ross Veitch, Timothy O’Neil-Dunne, David Liu, Rajnish Kumar, Anthony Chen, Idan Zalzberg, Larry Fang, Jennifer Zhang, 陈天予 — and me.

Their job now: narrow 46 down to the Top 3.

 


This column is written by Mary Li (CEO and Founder of Atlas). Atlas is a Singapore-based global travel technology company focusing on intelligent LCC retailing and infrastructure, connecting 140+ low-cost carriers with Travel Sellers through a single API.


 

BACK