The AI Verdict Study
What four AI engines told the same buyer about 18 brands, asked 576 times.
I wanted to know what happens after the first answer. A buyer asks an AI engine for options, gets a list, pushes back a few times, and then asks the question that matters, which is "so should I choose this one?" This study measures how that last answer comes out, for a challenger brand and for the category leader it competes with, across four engines, with the buyer's words held identical on every run.
If you want the story of how this came together (including the first version I nearly published, which was wrong), that's in the blog post. This page is the reference that includes the findings, the charts, the method, the limits, and the data.
Findings
Category leaders were rejected about twice as often as challengers. Engines told the buyer not to choose the category leader in 35% of conversations (75 of 215) and the challenger in 16% (35 of 214). Challengers got a clean yes 57% of the time and leaders 26%. In every pair the buyer was written as the challenger's ideal customer, so this is a finding about fit. The engines read the buyer's constraints and used them.
The engine mattered more than anything else measured. ChatGPT never rejected a challenger in 54 conversations. Claude rejected the category leader 74% of the time. Gemini rejected almost nobody (12% and 11%). Klaviyo got six yeses out of six from ChatGPT and six nos out of six from Claude, from the same buyer using the same words.
29% of brand and engine pairings were unstable. In 21 of 72 pairings, the same buyer got a recommendation on some runs and a rejection on others. In 11 of those, the answer went from a clean yes to a no. ChatGPT and Gemini each split on 2 brands of 18, Perplexity on 6, and Claude on 11.
Being named in the first answer went with a better ending. When the engine named the brand in turn one, before the buyer mentioned it, the conversation ended in a rejection 18% of the time. When it didn't, 44%.
Turning search off barely changed ChatGPT and Gemini, and changed Claude a lot. From memory alone, ChatGPT and Gemini matched their usual search-on verdict in 29 and 30 of 36 conversations. Claude matched in 17, said yes once, and refused to give any verdict 11 times. With search on, Claude never refused in 108 conversations.
Two category leaders never got a clean yes. Squarespace (0 of 24) was usually recast as half the answer, with a specialist gallery tool recommended for the rest. Tempur-Pedic (0 of 24) was named in the first answer once and rejected outright by ChatGPT and Claude every time.
Challenger versus leader, by engine
The gap held on each of the three days. Challengers were rejected 18%, 14%, and 17% of the time, and leaders 35%, 36%, and 34%. When an engine rejected a leader, it named the paired challenger in that same answer 55 times out of 75, though the buyer had never mentioned it.
Counts are out of 24 conversations per brand (23 where a Gemini run didn't finish). Six challengers beat their leader. Brooks and Breville held. Basecamp and Asana both lost the agency buyer.
Every verdict, every run
Each dot is one full conversation. Switch to the memory-only view to see what changed with web search off. Tap any cell to read how each run ended, in the engine's own words.
Stability
The buyer's messages are hashed, and the hash matched on every run, so any difference between two runs of the same cell comes from the engine. Two runs on the same day landed on the same side (favorable or rejected) in 183 of 213 same-day pairs. More of the movement showed up between days. As an out-of-sample check, a fourth day of Perplexity runs matched the most common verdict from the first three days 21 times out of 36.
First mention and final verdict
Turn one never names a brand. This is a correlation, and part of it is a fit signal, since an engine that thinks a brand suits the buyer will tend to both name it early and recommend it late. ChatGPT's figure for leaders is pulled down by the shoe and mattress buyers, where it often answered turn one with clarifying questions and no brands.
Memory versus search
This is the smallest sample in the study (two runs per cell, one day), so treat it as directional. Three cells show what search was doing for Claude. Huntress went from yes six times out of six with search to no both times from memory, where Claude's stated reason was "Huntress requires you to have security expertise to act on what it finds." With search, the same model described a managed service with a 24/7 human team. Saucony was told it "runs narrow" from memory and had its wide sizing cited with search. Pipedrive went the other way, a yes from memory and a no five times out of six with search, after Claude found cleaning-industry software (Jobber, QuoteIQ) and decided the buyer needed a field-service tool. I haven't tried to referee which descriptions are correct.
All 18 brands
Method
Pairs. Nine categories, each with a challenger and the brand most buyers would call the leader: Basecamp and Asana, Huntress and CrowdStrike, Omnisend and Klaviyo, Pipedrive and HubSpot, OnPay and Gusto, Pixieset and Squarespace, Saucony and Brooks, Helix and Tempur-Pedic, Gaggia and Breville.
Buyer. One buyer per pair, written as the challenger's ideal customer, with a role, a job to get done, and constraints stated as facts about the buyer (size, staffing, budget) and never as product features. Example: "founder and owner of a 12-person commercial cleaning company who handles sales personally," who needs to "move my 12-person commercial cleaning company off a lead-tracking spreadsheet and into a CRM," given "a budget under $50 per user per month and nobody on staff to administer software."
Script. Seven frozen turns: a broad question with no brand named, the brand introduced ("I've heard of..."), a skeptical question about tradeoffs given the constraints, a question about outgrowing it, a question about what people like me compare it with, a final pushback, and "should I choose it?" The leader's script is identical to the challenger's except for the brand name. Each brand's script is hashed with SHA-256, and the hash matched on every run, every day, and both search settings. The engine receives only the buyer's messages and its own earlier answers.
Engines. gpt-5 (returned as gpt-5-2025-08-07), claude-haiku-4-5 (20251001), gemini-2.5-flash, and Perplexity sonar-pro. All were called through their APIs by way of Replit AI Integrations (Perplexity through OpenRouter), with no system prompt, no date injected, and full history on every turn. Web search was each vendor's own tool. Claude was capped at five searches per turn because its API exposes a cap and the others don't. Temperature was left at provider defaults.
Runs. Search on: two runs per brand per engine per day on September 18, 19, and 20, 2026, for 432 conversations. Three Gemini runs on the third day didn't reach the final turn, so verdicts are out of 429. Search off: two runs per brand on ChatGPT, Claude, and Gemini on September 20, for 108. Thirty-six Perplexity runs launched in that batch are search-on by nature and were used only as the day-four check.
Labels. An automated classifier (gpt-5-mini) labeled each final answer yes, conditional, refused, or no, seeing only that answer and the brand name. A second reader screened all 573 final answers for labels that contradicted the answer's own wording and corrected 33. Two human raters then labeled 141 blind, with engine names hidden: all 33 corrected items plus a random 20% of the rest. On rejected (no or refused) versus not rejected, the raters agreed with each other 90% of the time (Cohen's kappa 0.78) and with the final labels 93% and 96% of the time. On the random sample, which stands in for the answers nobody hand-checked, agreement was 94% and 96%. On all four labels the raters agreed 66% of the time (kappa 0.53), with almost all disagreement falling between yes and conditional. So rejection rate is the primary measure and the clean-yes figures are secondary. Disagreements were settled by majority among the two raters and the second reader, and I broke one three-way tie. An answer of "use this brand for part of the job and something else for the rest" was labeled conditional.
Intervals. 95% Wilson intervals on the headline figures: challengers rejected 12 to 22%, leaders 29 to 41%; challengers clean yes 50 to 63%, leaders 21 to 32%.
Limits
One buyer per category, written to favor the challenger. I'm not claiming challengers beat leaders in general. A buyer written as the leader's ideal customer is a different study.
These are API results with no system prompt. The consumer apps add their own system prompts, memory, and search stack.
One model per vendor, and two of them (Haiku, Flash) are the small model in their family. I wouldn't assume the larger models behave the same way.
Three consecutive days. It says nothing about drift over weeks or months.
The buyer is a script. A real person would wander off it.
The memory-only round is two runs per cell on a single day.
The study measures what the engines responded to. It didn't test whether changing anything changes a verdict.
I built the tool that ran it (Citingly), so please do check my work.
The pilot
A first pass on September 16 and 17 ran seven of these challengers once per engine in each of two rounds, with an AI buyer that improvised its wording and no category leaders. It suggested two things that didn't survive the rerun which is that every dismissal went to the category leader, and that Claude refuses to pick. With leaders in the test, leaders were dismissed more often than challengers and Claude only refused when it couldn't search. The pilot's results are in the last chart as a separate tab and aren't combined with anything above.
Data
Every conversation, every turn, every label (the classifier's, the corrected one, and both raters'), and all 37,440 source URLs the engines cited are in one download.
Download the dataset (zip, 6 MB)
If you rerun any of this and get something different, I'd love to hear about it, especially if it disagrees with me.
How to cite
Smith, Jarred. "The AI Verdict Study" jarredsmith.com, September 2026. https://jarredsmith.com/research/ai-verdict-study
Changelog
Version 1, September 2026. Nine pairs, four engines, 432 search-on and 108 search-off conversations, audited labels with two-rater reliability check.
Planned. The same pairs with a buyer written as the leader's ideal customer. The larger model from each vendor. A before-and-after test of a specific content change on an unstable cell.
Jarred Smith is the author of Explainable: Why AI Recommends Some Brands & Ignores Others, an Amazon bestseller on AEO, GEO, and SEO. He's a marketing leader with nearly 20 years of experience across healthcare, public media, retail, and environmental services. Find him at jarredsmith.com.