bluffing-like
Failure modeThe model's answer changes across repeated, identical questions, but it never once says it's unsure. Anyone comparing the answers side by side can see the instability. The model itself never lets on.
If you give AI models the one piece of evidence that actually explains your traffic drop, they get it right almost every time. Take that evidence away, and half of them still name a cause anyway, with the same confidence they had when they were right.
We established a framework to compare the different AI models around the field of SEO. The WARRANT-SEO Benchmark uses a methodology to test whether a model can tell the difference between a diagnosis supported by evidence and one that only sounds plausible.
If you give an AI model the log line, the response header, or the impressions-vs-CTR split (the evidence that explains what happened), it gets diagnosis right nearly every time. However, if you take that one detail away, most models don't slow down but rather pick up a cause and commit to it, with exactly the same confidence they would have when the evidence supports them.
The team at iPullRank developed the WARRANT-SEO benchmark to assess whether the models are appropriately answering or if they're confidently guessing. You need to know if you can trust AI models when depending on them for SEO recommendations or diagnoses.
“Right call” means naming the correct cause when evidence is sufficient, and saying "I can't tell from this" when it is not.
This one score captures both halves of the question: does the model name the right cause when the evidence is actually there, and does it back off when it isn't? You need both to be useful. If you only measure the first, a model that always names a cause looks perfect, and that's what the 0.50 line represents.
We wanted to know whether or not AI models could be reliable, and our WARRANT-SEO benchmark provides a guide for which models you can depend on confidently.
The leaderboard above is public. Fill out the form to unlock the detailed analysis and full WARRANT-SEO results table below.
The remaining sections become available after successful submission.
This is what a diagnosis actually costs, based on real billed usage on this task, including thinking tokens rather than sticker price per token. Below the cliff, price mostly stops mattering, because you're paying for an answer that isn't grounded in anything.
If you're choosing one: Gemini 3.7 Flash and Claude Opus 5 both scored 1.000, but Opus cost roughly 28.5× more per thousand calls.
Using the same 22 dated factual questions across the same 27 models, we added this sentence to each question: “If your training data does not cover this period, say so rather than inferring.” Abstention rose from 20.7% to 85.0%, while correct responses fell from 64.8% to 13.8%.
The drop in the correct-answer rate was about 4 times the drop in the incorrect-answer rate.
Our WARRANT-SEO benchmark has created three different outcomes based on the testing scenarios. There are two ways a model can be wrong, and one way it can help you.
The model's answer changes across repeated, identical questions, but it never once says it's unsure. Anyone comparing the answers side by side can see the instability. The model itself never lets on.
The model gives the same unsupported answer every time, with little or no abstention. That kind of consistency can look like confidence, but repeating an answer is the easiest thing to fake.
High accuracy when the evidence supports a conclusion, and a high rate of stepping back when it doesn’t. This is the behavior the other two are measured against.
This describes what these models actually do, not what they "know." We asked each question three times and watched whether the answer changed.
The unanswerable version of each situation is the answerable one minus exactly one line. We're not asking a model to agree with our opinion here, since we removed the evidence ourselves, so there's genuinely nothing left to conclude.
Pre-publication review looked for outdated mechanisms, unsupported causal assumptions, and items that did not meet the benchmark’s own evidentiary standard.
Two evaluated scenarios were retired after results existed. Their responses were excluded; no replacement model calls were generated.
The situations happen to be about SEO, but the method behind them is not. We built WARRANT here first because our own diagnostic work is genuinely ambiguous.
It would be easy to score models only on the questions that separated them cleanly, since that makes any leaderboard look sharper than it really is. Every number here represents the full set of questions.
After all that reviewing, rewriting, and cutting, the headline number moved less than one percent. If this were just an artifact of badly written questions, that much correction should have erased it, but it didn't.
“Knowledge” is a separate technical SEO test. Included here to show that knowing the domain and knowing when the evidence supports a diagnosis are different skills.
Discrimination is the WARRANT score: how well each model answers when the evidence is sufficient and steps back when it is not. Claude Sonnet 5 is the clearest example of the importance of this leaderboard; being near the top on knowledge, but mid-pack on judgment.
| Rank | Model | Discrimination | 95% CI | Knowledge | $ / 1k calls |
|---|---|---|---|---|---|
| 1 tied | Gemini 3.7 Flash | 1.000 | 1.000–1.000 | 0.910 | $0.98 |
| 1 tied | Grok 4.5 | 1.000 | 1.000–1.000 | 0.821 | $4.49 |
| 1 tied | Muse Spark 1.2 | 1.000 | 1.000–1.000 | 0.923 | $5.55 |
| 1 tied | Grok 4.6 | 1.000 | 1.000–1.000 | 0.846 | $5.78 |
| 1 tied | Kimi K3 | 1.000 | 1.000–1.000 | 0.932 | $7.02 |
| 1 tied | GPT-5.5 | 1.000 | 1.000–1.000 | 0.987 | $10.39 |
| 1 tied | Claude Opus 5 | 1.000 | 1.000–1.000 | 0.962 | $27.87 |
| 2 | GPT-5.6 Terra | 0.994 | 0.982–1.000 | 0.910 | $2.80 |
| 3 | GPT-5.6 Luna | 0.988 | 0.969–1.000 | 0.910 | $0.29 |
| 4 | GLM-5.3 | 0.987 | 0.966–1.000 | 0.831 | $11.00 |
| 5 | Qwen3.8 Max | 0.950 | 0.895–0.994 | 0.883 | $9.71 |
| 6 | Gemini 3.6 Flash | 0.938 | 0.876–0.988 | 0.859 | $3.70 |
| 7 | GLM-5.2 | 0.887 | 0.820–0.942 | 0.794 | $3.17 |
| 8 | Claude Sonnet 5 | 0.728 | 0.648–0.809 | 0.897 | $9.73 |
| 9 | DeepSeek V4 Pro | 0.593 | 0.556–0.636 | 0.821 | $2.72 |
| 10 | MiniMax M3 | 0.582 | 0.531–0.647 | 0.872 | $1.96 |
| 11 | Qwen3.7 Flash | 0.561 | 0.508–0.629 | 0.859 | $0.37 |
| 12 | Claude Haiku 4.5 | 0.556 | 0.500–0.611 | 0.747 | $2.37 |
| 13 | Nemotron 3 Ultra | 0.544 | 0.506–0.596 | 0.623 | $4.80 |
| 14 | Nova Premier | 0.525 | 0.500–0.556 | 0.667 | $3.51 |
| 15 | Command A | 0.525 | 0.469–0.580 | 0.590 | $3.67 |
| 16 | Mistral Large | 0.500 | 0.500–0.500 | 0.808 | $2.71 |
| 17 | Seed 2.1 Turbo | 0.500 | 0.500–0.500 | 0.833 | $10.93 |
| — | GPT-5 nano | 0.550 | 0.506–0.611 | 0.590 | $0.69 |
| — | Mistral Small 3.2 | 0.494 | 0.481–0.500 | 0.636 | $0.10 |
| — | Llama 3.1 8B | 0.475 | 0.426–0.512 | 0.532 | $0.02 |
| — | Ministral 3B | 0.444 | 0.389–0.500 | 0.526 | $0.04 |
Seed 2.1 Turbo excludes 40.3% of diagnostic responses; its score rests on only four rated evidence-removed responses. Knowledge-task exclusions are 12.8% for GLM-5.2 and 30.8% for Seed 2.1 Turbo. Confidence intervals are scenario-bootstrap intervals; 1.000–1.000 does not guarantee perfect performance on new cases.
Schofield, Patrick. (2026). WARRANT-SEO: Measuring Whether Language Models Recognize Insufficient Evidence in Professional Causal Diagnosis. iPullRank.
@misc{schofield2026warrantseo,
author = {Patrick Schofield},
title = {WARRANT-SEO: Measuring Whether Language Models Recognize Insufficient Evidence in Professional Causal Diagnosis},
year = {2026},
month = {September},
organization = {iPullRank},
note = {Publication analysis v2.8}
}