Google just let AI book your hotel. Researchers just showed the AI can be tricked.
The same week the shelf shrank to three options, researchers showed the agent can be fooled.
The first is a product launch. The second is a research paper. Neither one, on its own, is the whole story.
The launch
Google confirmed to Skift that it has begun a limited test in the United States in which its own AI assistant chooses and books a hotel on the user’s behalf, inside Search’s AI Mode. No comparing ten tabs. No scrolling a results page. You describe what you want, the agent decides.
The company also introduced something called the Universal Commerce Protocol, infrastructure that could let a booking be completed inside the chat itself, without redirecting the user anywhere else. In the same week, Google Maps rolled out similar agentic features, including the ability to book a hotel or order food directly from the app. The example Google itself used to demonstrate the new function is worth sitting with: find a well-reviewed, reasonably priced hotel near a gym and restaurants for a conference next weekend. That is not a hypothetical leisure query. It is, almost verbatim, what a business traveler asks every week.
This is the moment the shelf actually shrank, not the moment someone predicted it would. A recent study from Skift found that only 6% of hotels currently appear in AI-powered search, meaning 94% simply do not surface when a traveler asks an assistant for a recommendation. That statistic explains why Google is moving from suggesting options to making the choice directly: if the discovery layer already compresses thirty options into three, the company that owns the layer has an obvious incentive to own the decision too.
The paper
Now the second thing, and it complicates the first in a way that deserves more attention than it has gotten.
Researchers published a study demonstrating that vision-language agents, the same kind of system now being tested to book your hotel, can be reliably distracted by adversarial pop-ups: fake elements deliberately designed to mislead the machine rather than the person. The paper explicitly names travel booking as one of the task categories most exposed to this kind of attack, and it measured high success rates for the deception across the testing environments used. The work was presented last year at the Association for Computational Linguistics, one of the field’s most rigorous venues, which means this is peer-reviewed research, not speculation from a blog.
The mechanism is almost unsettling in how simple it is. A human ignores a fake pop-up instinctively, a skill built from years of exposure to bad ads and phishing attempts. An AI agent parsing a screenshot or an accessibility tree has no equivalent instinct yet. It processes what is in front of it, and a sufficiently well-designed fake element can sit right in the path of the action the agent is about to take.
Why both facts belong in the same sentence
I do not think the right reaction to the research is alarm, and I do not think the right reaction to Google’s test is uncritical enthusiasm. The two findings are not in tension. They are describing the same transition from two different angles.
The launch tells you that agentic booking has moved from roadmap to shipped product, tested by a company whose search engine is most people’s starting point for travel discovery. The paper tells you that the exact category of task Google just automated is one where independent researchers have already demonstrated a real, reproducible vulnerability. Put together, they describe a technology that is simultaneously more capable and less battle-tested than the pace of its rollout suggests.
That combination is not new in technology generally. It is relatively new in a category where the failure mode is not a wrong answer in a chat window, but a real charge on a real card, for a hotel a real person is going to sleep in.
What this changes for anyone managing travel
If I ran a corporate travel program, this week would change three things about how I think about agentic booking, not in the direction of avoiding it, but in the direction of asking better questions about it.
First, I would ask any vendor pitching agentic booking exactly one question: what happens between the agent’s decision and the money moving. Not whether the agent is smart. Whether there is a verification layer between intent and transaction that a manipulated screen cannot bypass. The pop-up research is specific about where the vulnerability lives, at the exact moment an agent parses what is on the screen and decides where to click. That is precisely the moment a corporate program needs a checkpoint, not an assumption.
Second, I would treat visibility and trust as two separate problems, because the industry is currently conflating them. Being one of the 6% of hotels an AI assistant actually surfaces solves a discovery problem. It says nothing about whether the transaction that follows is secure. A supplier optimizing purely for AI visibility without also hardening the booking flow against exactly this kind of manipulation is solving half the problem and calling it done.
Third, I would stop treating “the AI decided” as a satisfying answer to “why this hotel.” The whole appeal of agentic booking is removing friction, and removing friction is real value. But friction and scrutiny are not the same thing, and a program that automates away scrutiny along with friction is trading a slow process for a fast one with a new, largely unaudited failure mode.
The pattern underneath
I keep returning to a version of this argument, because I think it is the actual shape of what is happening across this industry right now, not just in hotel booking.
Capability is advancing on one curve. Trust infrastructure is advancing on a slower one. Google shipping agentic hotel booking is capability. The pop-up study is a reminder that trust infrastructure has not caught up, and that the gap between the two curves is exactly where the next round of fraud, error, and bad press in this industry is going to come from.
None of this means slow down. It means build the checkpoint before you need it, not after the first story about an agent that booked the wrong hotel, at the wrong price, because it clicked the wrong thing.
The shelf really did shrink to three options this week. The research just reminded everyone that someone can still rearrange which three you see.
About me
I am an entrepreneur with over 20 years of experience at the intersection of tourism and technology. I am co-founder and Chief Business Officer of VOLL, the largest mobile-first corporate travel and expense management platform in Latin America, and a recognized reference in the development of the corporate travel industry.
A Marketing specialist from Fundação Dom Cabral, I serve on the Tourism Council of FecomércioSP and on the Executive Council of the Latin American Association of Corporate Events and Travel Management (Alagev). A frequent traveler and close observer of human behavior in motion, I write and speak about innovation, digital transformation, entrepreneurial leadership, and the future of corporate travel.



