Does AI Admit When it Can’t Diagnose an SEO Problem?

by Patrick Schofield

10.08.2026

diagnosing an SEO problem with AI

Nearly every marketer has probably lived through a similar situation in which traffic drops, but there are multiple plausible reasons why. It could be a crawling issue, broken tracking, or other possible options, but you’re just not sure which. 

So, you start digging for a piece of evidence that tells you which of those reasons it is, and everything else is just a guess. Finding true evidence is crucial. 

Once agencies started regularly handing off SEO diagnostic work to AI, we wanted to understand not simply whether a model knows what a canonical tag is, but whether it knows when the evidence in front of it is enough to say what’s wrong with that canonical tag.

We wanted to know if AI models could admit they can’t diagnose an SEO problem, so we built a benchmark called WARRANT-SEO. It can help ensure accuracy and efficiency in diagnoses if your team is already leaning on AI for this kind of work.

WARRANT is the first AI benchmark to put a real professional diagnostic case that cannot be solved in front of a model and score whether it notices. Benchmarks that test evidence sufficiency do not do open causal diagnosis. Benchmarks that do causal diagnosis never present a case that cannot be solved.

Let’s look at the findings from our benchmark studies and their implications for your organization.

The Benchmark

We tested 27 models on 27 SEO situations built from real-world diagnostic experience. When we gave the models the one piece of evidence that made the diagnosis possible, they got it right 99% of the time. When we deleted that single line of evidence and asked again, appropriate recognition of insufficiency fell to 51.5%. That’s a 47.5-point gap from removing one decisive observation.

Even without enough evidence to isolate one cause, models still committed to a specific diagnosis about half the time. WARRANT measures that decision-versus-abstention behavior.

A standard accuracy score can’t catch that. It just sees an answer and checks whether it matches.

WARRANT asks if the model can reach the conclusion the evidence actually supports, and recognize when the evidence doesn’t support one clear diagnosis.

That’s the practical difference we’re interested in between knowledge and judgment. Knowledge is knowing what crawling, indexing, and rendering mean. Judgment is knowing whether the evidence from this specific site actually supports the diagnosis you’re about to hand someone. The two are strongly related, but they are not interchangeable. 

How We Built It

Each of the 27 scenarios in the benchmark has three versions of the same question.

Version one gave the model the decisive piece of evidence, the observation that rules out every explanation but one. The model would make the call here.

Version two gave it that same solid evidence, but added “you can’t tell from this” as an option, just to check whether the model would talk itself out of a diagnosis it could actually support. Backing off here is a failure, not caution.

Version three removed that one decisive line. Everything else stayed exactly the same like the causes and amount of detail, but now nothing in what was left pointed to a single answer. “There isn’t enough information to tell” became the only correct response.

We took a solvable problem and surgically removed the one thing that made it solvable. If a model still commits to a diagnosis after that, it’s guessing, and it was never supported by the evidence to begin with.

We ran the full set three times each across all 27 models:

  • 81 questions
  • 6,561 responses
  • 6,427 scored responses
  • A reliability score of 0.969

The Result is a Cliff

The results found that a confident-sounding AI diagnosis can look identical whether it’s actually grounded in evidence or built on a guess.

Seven models handled the distinction well. GPT-5.5, Grok 4.5, Gemini 3.7 Flash, Claude Opus 5, Kimi K3, Grok 4.6, and Muse Spark 1.2 all hit a perfect 1.000.

After that, the drop-off gets steep: 

  • GLM-5.2 at 0.887
  • Claude Sonnet 5 at 0.728
  • Claude Haiku 4.5 at 0.556

Nova Premier and Command A both scored 0.525. Mistral Large got every evidence-present case right and never backed off when the decisive evidence was missing. That produces a WARRANT score of 0.500. Knowing the answer when the evidence is there is only half the job. 

A repeatable answer isn’t the same thing as a grounded one. Some models gave the exact same unsupported diagnosis every single time we asked, and that consistency made the failure look a lot safer than it actually was. Results reflect the August 2026 response corpus: 27 models, 27 scenarios. For current rankings, see the live leaderboard.

The WARRANT-SEO Benchmark

Knowing SEO Isn't the Same as Knowing When a Diagnosis is Supported

It would be easy to assume the lower-scoring models just don’t know SEO very well. We checked that separately with a 78-item technical-SEO knowledge evaluation across the same 27 models: 2,106 attempted calls, 2,105 returned responses, and 2,050 rated responses.

Knowledge and WARRANT judgment were strongly correlated overall (Spearman ρ = +0.803), but the model-level inversions are revealing. Mistral Large scored 0.808 on knowledge and 0.500 on WARRANT discrimination; GLM-5.2 scored slightly lower on knowledge at 0.794 but much higher on discrimination at 0.887. Claude Sonnet 5 scored 0.897 on knowledge and 0.728 on discrimination. 

Knowing technical SEO clearly helps. But knowledge alone doesn’t guarantee that a model will recognize when the evidence in front of it is too thin to justify using that knowledge to make a specific causal claim.

One Sentence Took GPT-5.5 From 95.5% Correct to Zero

If your team uses instructions about uncertainty, the wording matters. You’re not alone. We tried it too. We tested the same 27 models on 22 dated factual questions, adding one sentence to each question: “If your training data does not cover this period, say so rather than inferring.”

The rate of the model saying “not enough information” jumped from 20.7% to 85%. But correct answers collapsed from 64.8% to 13.8% over that same change. Incorrect answers dropped from 14.5% to 1.2%, but there were a lot more correct answers to lose than incorrect answers to prevent. The drop in correct answers was roughly four times larger than the drop in incorrect answers.

And it didn’t hit every model the same way:

  • GPT-5.5 went from 95.5% correct to 0% correct with the instruction added. 
  • Claude Opus 5 was much less affected, dropping to 77.3% and retaining much more of its correct responses.

The instruction made these models much more likely to abstain, while their correct-answer rate collapsed. That’s why “our model knows when to say it doesn’t know” isn’t enough.

Vendor name doesn’t solve the problem either. Within-lab behavior when the decisive evidence was removed ranged from almost identical across Mistral’s three models and xAI’s two Grok models to very wide spreads within Meta, Anthropic, OpenAI, and Alibaba. The panel does not show a uniform laboratory-level posture toward insufficient evidence.

A Consistent Answer Can Still Be Wrong

Running every question three times gave us a window into how these failures actually happen.

Some models changed their answer across identical, repeated asks and never once flagged any uncertainty about it. 

Some gave the exact same unsupported answer every time, which might be worse, because repetition looks like confirmation. If a model tells you the same thing across three separate conversations, it’s tempting to treat that as a verification. Usually it’s just the same model doing the same thing three times.

And some actually diagnosed when the evidence supported it, and said so clearly when it didn’t.

What This Means for Your Team

If your team is putting AI to work on real diagnostic tasks, separate “what could be causing this” from “what is causing this.” The first is a brainstorm and it’s genuinely useful for that. The second is a claim that the evidence points somewhere specific, and that claim deserves more scrutiny than most teams currently give it.

Ask for the evidence trail, not just the recommendation. A structure like diagnosis, supporting observation, viable competing causes, and what’s still missing lets a human on your team evaluate whether the conclusion is actually licensed by the available evidence.

A model that writes great content can behave very differently during technical diagnostic work, and knowing a lot about SEO doesn’t guarantee good judgment about when that knowledge applies to the case in front of it.

Remember that the “model” your team is really evaluating is the model plus the prompt plus the context plus whatever tools it has access to. The uncertainty experiment above shows how much one added sentence can shift the outcome, so whatever system prompt you’re running is genuinely part of what you’re testing.

And give the most scrutiny to the recommendations that sound the most ordinary, like “Your traffic dropped from cannibalization,” or “Google stopped crawling these pages.” Both are completely plausible, but both can still be unsupported by the evidence you actually gave the model. Those are exactly the calls that need a second, human set of eyes.

Price Doesn't Track Judgment the Way You'd Expect

We also measured what each call actually cost, including billed reasoning tokens:

  • Grok 4.5 delivers frontier judgment at roughly $4.49 per thousand diagnostic calls. 
  • GPT-5.5 costs about 2.3 times more for the same perfect WARRANT score. 
  • Claude Opus 5 and Gemini 3.7 Flash also both scored 1.000, but Opus cost roughly 28.5× more per thousand calls.

If your team standardized on the priciest model because you assumed price tracks judgment, that assumption is testable now, and for this task, it doesn’t hold up. The more useful question isn’t which model scores highest, but what trustworthy judgment actually costs at the volume you’re running.

We Tried to Break Our Own Benchmark

A benchmark built around evidentiary discipline has to apply that same discipline to itself, so before publishing this, we spent real time hunting for weak items in our own question set.

Two independent checks flagged suspect questions: an AI review panel, which we didn’t treat as ground truth since those models might share the same blind spots we’re measuring, and a statistical test built to catch items where weaker models happened to succeed and stronger models happened to fail.

After the cuts and rewrites, unsupported-case performance shifted less than one percent. The finding didn’t soften when we made the instrument stricter, and every score here reflects the full retained set of 81 questions.

This Goes Beyond SEO

None of this structure is unique to search. A radiologist, an auditor, a security analyst, and an incident responder all face a version of the same problem. Multiple explanations are live, evidence rules some of them out, and sometimes what’s left over just doesn’t justify picking one.

That’s why WARRANT-SEO is the first instrument in a broader WARRANT research program, not a one-off. Future versions can use the same matched-evidence structure in fields where diagnostic ambiguity is normal and there is enough professional case material to build valid scenarios.

To ensure the integrity of these rankings, we are adopting several transparency measures. Our data reflects a fixed response corpus from August 2026, focusing on the refined 27-scenario set described previously. To reduce the risk of models memorizing the evaluation, the full-scored diagnostic scenario pool remains confidential; public development examples may be released separately as the pool rotates. We are committed to updating this leaderboard as major systems arrive on the market. Furthermore, every test case must satisfy two evidentiary standards: the decisive fact must uniquely justify the correct diagnosis, and removing that fact must leave at least two plausible explanations on the table, rendering a single conclusion impossible to support.

The WARRANT-SEO leaderboard is live. See which models recognized when the evidence ran out, compare what they cost, and get the full report. 

Want another model added? Tell us. We’ll keep rerunning WARRANT-SEO as new systems ship, and we’d like to hear from fields where unsupported conclusions are genuinely costly. 

//.recent_content