WARRANT-SEO Benchmark

Your AI can diagnose an SEO problem, but does it know when it can’t?

If you give AI models the one piece of evidence that actually explains your traffic drop, they get it right almost every time. Take that evidence away, and half of them still name a cause anyway, with the same confidence they had when they were right.

We established a framework to compare the different AI models around the field of SEO. The WARRANT-SEO Benchmark uses a methodology to test whether a model can tell the difference between a diagnosis supported by evidence and one that only sounds plausible.

The Evidence Gap
99.0%
Right call when the evidence is actually there
−47.5 pts
One line
deleted
51.5%
Right call once we delete it - the rest name a cause anyway
27Diagnostic Scenarios
27Models Tested
6,427Answers Scored
01

Knowledge is
not judgment

These models know technical SEO well, so that was never really the problem.

If you give an AI model the log line, the response header, or the impressions-vs-CTR split (the evidence that explains what happened), it gets diagnosis right nearly every time. However, if you take that one detail away, most models don't slow down but rather pick up a cause and commit to it, with exactly the same confidence they would have when the evidence supports them.

The team at iPullRank developed the WARRANT-SEO benchmark to assess whether the models are appropriately answering or if they're confidently guessing. You need to know if you can trust AI models when depending on them for SEO recommendations or diagnoses.

Data/Warrant
Evidence PresentEvidence Removed

“Right call” means naming the correct cause when evidence is sufficient, and saying "I can't tell from this" when it is not.

02

The
judgment
cliff

Nine of the twenty-three models we tested score like a model that just blames something every time.

This one score captures both halves of the question: does the model name the right cause when the evidence is actually there, and does it back off when it isn't? You need both to be useful. If you only measure the first, a model that always names a cause looks perfect, and that's what the 0.50 line represents.

We wanted to know whether or not AI models could be reliable, and our WARRANT-SEO benchmark provides a guide for which models you can depend on confidently.

Cite: Schofield, Patrick. (2026). WARRANT-SEO: Measuring Whether Language Models Recognize Insufficient Evidence in Professional Causal Diagnosis. iPullRank.
Get the downloadable report

Want to keep going?

The leaderboard above is public. Fill out the form to unlock the detailed analysis and full WARRANT-SEO results table below.

The remaining sections become available after successful submission.

The rest of this WARRANT experience stays locked until the form is successfully submitted.
Get access
Loading form…
The rest of the report is now available below.
03

What trustworthy judgment costs

The most expensive model here isn't the best one.

This is what a diagnosis actually costs, based on real billed usage on this task, including thinking tokens rather than sticker price per token. Below the cliff, price mostly stops mattering, because you're paying for an answer that isn't grounded in anything.

If you're choosing one: Gemini 3.7 Flash and Claude Opus 5 both scored 1.000, but Opus cost roughly 28.5× more per thousand calls.

Judgment score vs billed cost on this task, reasoning tokens included
Swipe sideways to explore the chart →
Per-call cost values are mean billed OpenRouter cost per call over the collected run and are not recomputed after scenario retirement. *more than 10% of this model's responses never stated a decision. WARRANT-SEO, iPullRank 2026

One added sentence took GPT-5.5 from 95% correct to zero

One instruction about training-data coverage sharply changed how these models answered dated questions.

Using the same 22 dated factual questions across the same 27 models, we added this sentence to each question: “If your training data does not cover this period, say so rather than inferring.” Abstention rose from 20.7% to 85.0%, while correct responses fell from 64.8% to 13.8%.

The drop in the correct-answer rate was about 4 times the drop in the incorrect-answer rate.

Neutral wording+ one uncertainty instruction
04

How the
models fail

Our WARRANT-SEO benchmark has created three different outcomes based on the testing scenarios. There are two ways a model can be wrong, and one way it can help you.

bluffing-like

Failure mode

The model's answer changes across repeated, identical questions, but it never once says it's unsure. Anyone comparing the answers side by side can see the instability. The model itself never lets on.

Qwen3.7 Flash - GPT-5 nano - DeepSeek V4 Pro - GLM-5.2

Unsupported Convergence

Failure mode

The model gives the same unsupported answer every time, with little or no abstention. That kind of consistency can look like confidence, but repeating an answer is the easiest thing to fake.

Claude Sonnet 5 - Claude Haiku 4.5 - Mistral Small 3.2 - Ministral 3B - Llama 3.1 8B

Appropriately selective

Target

High accuracy when the evidence supports a conclusion, and a high rate of stepping back when it doesn’t. This is the behavior the other two are measured against.

GPT-5.5 - Grok 4.5 - Kimi K3 - Claude Opus 5 - Gemini 3.6 Flash

This describes what these models actually do, not what they "know." We asked each question three times and watched whether the answer changed.

05

Why you can trust these numbers

01.

There is no answer key to compare against.

The unanswerable version of each situation is the answerable one minus exactly one line. We're not asking a model to agree with our opinion here, since we removed the evidence ourselves, so there's genuinely nothing left to conclude.

02.

We tried hard to break our own test.

Pre-publication review looked for outdated mechanisms, unsupported causal assumptions, and items that did not meet the benchmark’s own evidentiary standard.

03.

We held ourselves to the same standard.

Two evaluated scenarios were retired after results existed. Their responses were excluded; no replacement model calls were generated.

04.

None of this is really specific to search.

The situations happen to be about SEO, but the method behind them is not. We built WARRANT here first because our own diagnostic work is genuinely ambiguous.

05.

We publish the version that’s less flattering.

It would be easy to score models only on the questions that separated them cleanly, since that makes any leaderboard look sharper than it really is. Every number here represents the full set of questions.

06.

The results barely moved once we fixed the test.

After all that reviewing, rewriting, and cutting, the headline number moved less than one percent. If this were just an artifact of badly written questions, that much correction should have erased it, but it didn't.

06

Every model,every number

Everything above comes from this table.

“Knowledge” is a separate technical SEO test. Included here to show that knowing the domain and knowing when the evidence supports a diagnosis are different skills.

Discrimination is the WARRANT score: how well each model answers when the evidence is sufficient and steps back when it is not. Claude Sonnet 5 is the clearest example of the importance of this leaderboard; being near the top on knowledge, but mid-pack on judgment.

RankModelDiscrimination95% CIKnowledge$ / 1k calls
1 tiedGemini 3.7 Flash1.0001.000–1.0000.910$0.98
1 tiedGrok 4.51.0001.000–1.0000.821$4.49
1 tiedMuse Spark 1.21.0001.000–1.0000.923$5.55
1 tiedGrok 4.61.0001.000–1.0000.846$5.78
1 tiedKimi K31.0001.000–1.0000.932$7.02
1 tiedGPT-5.51.0001.000–1.0000.987$10.39
1 tiedClaude Opus 51.0001.000–1.0000.962$27.87
2GPT-5.6 Terra0.9940.982–1.0000.910$2.80
3GPT-5.6 Luna0.9880.969–1.0000.910$0.29
4GLM-5.30.9870.966–1.0000.831$11.00
5Qwen3.8 Max0.9500.895–0.9940.883$9.71
6Gemini 3.6 Flash0.9380.876–0.9880.859$3.70
7GLM-5.20.8870.820–0.9420.794$3.17
8Claude Sonnet 50.7280.648–0.8090.897$9.73
9DeepSeek V4 Pro0.5930.556–0.6360.821$2.72
10MiniMax M30.5820.531–0.6470.872$1.96
11Qwen3.7 Flash0.5610.508–0.6290.859$0.37
12Claude Haiku 4.50.5560.500–0.6110.747$2.37
13Nemotron 3 Ultra0.5440.506–0.5960.623$4.80
14Nova Premier0.5250.500–0.5560.667$3.51
15Command A0.5250.469–0.5800.590$3.67
16Mistral Large0.5000.500–0.5000.808$2.71
17Seed 2.1 Turbo0.5000.500–0.5000.833$10.93
—GPT-5 nano0.5500.506–0.6110.590$0.69
—Mistral Small 3.20.4940.481–0.5000.636$0.10
—Llama 3.1 8B0.4750.426–0.5120.532$0.02
—Ministral 3B0.4440.389–0.5000.526$0.04

Seed 2.1 Turbo excludes 40.3% of diagnostic responses; its score rests on only four rated evidence-removed responses. Knowledge-task exclusions are 12.8% for GLM-5.2 and 30.8% for Seed 2.1 Turbo. Confidence intervals are scenario-bootstrap intervals; 1.000–1.000 does not guarantee perfect performance on new cases.

Cite this benchmark

Schofield, Patrick. (2026). WARRANT-SEO: Measuring Whether Language Models Recognize Insufficient Evidence in Professional Causal Diagnosis. iPullRank.

BibTeX
@misc{schofield2026warrantseo,
  author       = {Patrick Schofield},
  title        = {WARRANT-SEO: Measuring Whether Language Models Recognize Insufficient Evidence in Professional Causal Diagnosis},
  year         = {2026},
  month        = {September},
  organization = {iPullRank},
  note         = {Publication analysis v2.8}
}