· 6 minute read

On August 18, 2026, a team from the University of Pennsylvania and Northeastern submitted a preprint on what happens when searchers are given no choice about Google's AI Mode. Searches routed into AI Mode produced fewer clicks to websites, lower ratings of usefulness and satisfaction, lower trust in information found on Google, fewer search sessions, and a significant rise in the share of users searching on a competitor engine. Most coverage stopped at the traffic. The last finding matters more for anyone modeling a customer. When the answer was made unavoidable, a larger share of searchers did not stay where the answer was.

What did the AI Mode field experiment test?

The AI Mode field experiment tested what happens to real search behavior when Google's generative AI features are removed, left alone, or made unavoidable. The paper, by Stephanie T. Wang, Jeffrey Gleason, Yakov Bart, Christo Wilson and Danaé Metaxa, was submitted to arXiv on August 18, 2026 as a preregistered field experiment on Google Search with 1,100 participants.

According to Search Engine Journal's account of the study, 1,444 people enrolled and the searches of 1,100 were tracked over seven days through a browser extension. Participants were randomly assigned to one of three conditions: regular Google Search with AI Overviews hidden, Google behaving normally, or searches redirected into AI Mode. The redirect worked for 94.7% of searches in the AI Mode group, so the forced condition was close to what it claimed to be.

The design matters for anyone drawing conclusions about customers. The outcomes were declared before the data came in, and the behavior was recorded on participants' own machines during their own searches, not reported in a survey about what they imagine they would do. The limits are as plain: one week, and a preprint not yet through peer review.

What happened when searches were forced into AI Mode?

Forcing searches into AI Mode reduced clicks to publishers and lowered trust. The preprint states it directly: an AI Mode only experience "reduces click-through rates and erodes user experience and trust in information found on Google." Removing AI Overviews and AI Mode did the opposite and raised click-through to publishers.

Search Engine Journal reports three further results. The AI Mode group conducted fewer search sessions. Its members rated search as less useful and less satisfying. And assignment to AI Mode significantly increased the fraction of users searching on a competitor engine: Bing, DuckDuckGo or Yahoo.

Every one of these is a condition-level result. It describes how the AI Mode group differed from the other two groups, not what any single participant did after a given answer. The accounts cited here report the direction and significance of each shift, not its size, and we claim nothing beyond that.

Why is "AI search kills organic traffic" the wrong lesson from this study?

The traffic lesson treats the click as the unit of value and tries to recover it by optimizing to be quoted inside the AI summary. That advice assumes the summary is where the decision now gets made. In the AI Mode experiment, the group that received the summary as the default route reported the lowest trust in what it found on Google, which is weak ground for moving a brand's whole investment into that summary.

The study names where more of that group searched: Bing, DuckDuckGo or Yahoo. It does not report what they were looking for there, and we will not guess. What it does show is enough for a brand. Being named in the summary did not end the search for a measurable share of the people who received it.

A brand quoted in an answer has won a mention. The study gives no reason to treat a mention as the decision.

Does an AI answer end the customer journey?

An AI answer ends the customer journey only in the diagram of the product that produced it. The Marketing Helix holds that customers remain in motion and that marketing changes state around them. The AI Mode experiment gives that position a measurable case: the forced condition delivered an answer to every query, and the group receiving it searched on other engines at a significantly higher rate.

Of the three forces the Helix names, trust, relevance and timing, the study measured only trust. Relevance and timing were not measured, and nothing here claims they moved. Trust is enough for the point. The Helix holds that alignment is temporary, and the forced condition shows trust falling while the answer itself was delivered on schedule. Delivery was not alignment.

A model that ends the journey at "answered" is describing the interface, not the customer.

What is the strongest case against reading this study as evidence about customers?

The strongest case against this reading is novelty. People were pushed into an unfamiliar interface they did not choose, for seven days. Searching on Bing or DuckDuckGo could be a reaction to losing a familiar tool, the way people react to any forced redesign, and that reaction could fade within weeks. The accounts cited here do not report whether competitor-engine searching held steady across the week or decayed toward the other groups by day seven. Nothing reported shows any participant checking an AI answer against a second source.

That objection is correct about what the study cannot show. It cannot separate irritation from preference, it cannot show verification, and it cannot predict behavior after the novelty wears off. A day by day breakdown of competitor-engine share is the test that would settle it, and until that is published, the verification reading stays unproven.

The novelty explanation does not rescue the answer-as-endpoint model, because the model fails under either motive. Whether a searcher left out of doubt or out of irritation, the searcher was still moving after the interface had delivered its answer, and a model that ends at "answered" has no place to record that movement. The trust finding does not settle it either way. In the forced condition the information found on Google was the AI answer itself, so the study cannot separate distrust of the content from dislike of the interface.

A brand quoted in an answer has won a mention. The study gives no reason to treat a mention as the decision.

The evidence is one week and one preprint, and it should be held at that weight. Within that weight it is clear. When Google made its answer unavoidable, people trusted what they found there less, rated search as less satisfying, and a significantly larger fraction of them searched on another engine. A model of customer behavior that ends at the answer will be accurate about the software and wrong about the person, and the person is the one who buys.

Further reading

  1. Wang, Gleason, Bart, Wilson and Metaxa, the preregistered field experiment (N=1,100) this essay rests on arxiv.org/abs/2608.18352
  2. Roger Montti, Search Engine Journal, reporting the study design, the three conditions, fewer search sessions and the move to competitor engines searchenginejournal.com/research-shows-google-ai-mode-sends-less-clicks
  3. Google, Search updates from I/O 2026, the source for AI Mode passing one billion monthly users blog.google/products-and-platforms/products/search/search-io-2026

Every source above was fetched and a verbatim phrase confirmed on the page before this essay published. Nothing here is paraphrased from memory.