Technical benchmarks rarely help you understand the current landscape. Sometimes they make it even murkier. We decided to figure it out ourselves, drawing exclusively on high-quality independent research.

We evaluated models across three parameters. Long context means how accurately a model handles large documents: reports, contracts, academic texts. Speed covers how quickly it responds and completes a task. Reasoning addresses how deeply and correctly it thinks through complex analytical questions. For each parameter, we used data from independent sources only: Artificial Analysis, LMArena, LiveBench, and OpenRouter. (We've listed all sources in full at the end of this article.) We took no figures from the providers themselves, since there's always far more marketing than truth in those numbers, and thus the data in our table sometimes contradicts an LLM's established reputation.

We deliberately chose not to evaluate the models' ability to generate literary prose. On one hand, we believe all of them still do it very poorly (see: "Hemingway's Writing Style, as Faked by Five AI Models"). On the other hand, we have absolutely no confidence in benchmarks in this area.

Model comparison table

Below each model's name, we've listed its price range ($ – $$$$) and three scores based on independent data: EfC (effective context), Sp (speed), and Th (thought, or reasoning ability). Each score reflects how well the model performs on that parameter, with 0 meaning poor and 3 meaning excellent.

On the EfC scale, a score of 2 is a strong result; 3 goes only to models that stand out by a clear margin. The absence of negatives doesn't mean a model is perfect, just that independent testing hasn't been extensive enough to reveal flaws yet. "Generates images*" indicates that a separate model from the same provider provides the images in question.

Flagships

Claude Fable 5
$$$$
EfC 2 · Sp 0 · Th 3
First place in LMArena's blind comparison goes to Claude's top-of-the-line model, whose answers people choose more often than those of any other. Fable thinks deeply, handles facts carefully, and holds its own on demanding tasks, though in highly specialized domains it can be just as confidently wrong as the rest of the pack, and it performs at the same level as others on long documents. Use it for serious analysis and hard decisions.

Downsides: Fable is very expensive and very slow. Safety filters sometimes trigger false positives and silently swap the model out for a weaker one (Opus). Fable goes off on tangents, rarely gets answers right on the first try, and tends to play it safe.
Claude Opus 5.0
$$$
EfC 2 · Sp 2 · Th 3
The next step down for Claude formally ranks higher than Fable on the independent intelligence index at half the price. It handles long documents with confidence, follows instructions precisely, and offers an accelerated mode, making it a sensible choice for serious work when budget matters.

Downsides: despite all of that, Opus loses to Fable 5 in blind user tests. People simply prefer the older, pricier model, even if they don't know it.
GPT-5.6 Sol
$$$
EfC 2 · Sp 2 · Th 3
generates images*
OpenAI's top performer on independent reasoning benchmarks, Sol earned first place on the GPQA Diamond science test. Use it for analysis and reports. In creative work it's drier than Claude and more tightly structured, with less tonal range.

Downsides: it consumes more tokens than the listed rate suggests, so the final bill can come as a surprise. Like all OpenAI models, it's also verbose.
Gemini 3.1 Pro
$$
EfC 2 · Sp 2 · Th 2
generates images*
Google's flagship offers the best price-to-performance ratio among its kind. The model handles long context very well and understands science and images. Use it for writing and working with large documents, just don't expect deep reasoning.

Downsides: hallucinations and sycophancy are documented across the entire Gemini 3 family. The language is drier and more formal than in previous versions.
Grok 4.6 (NEW)
$$
EfC 2 · Sp 1 · Th 3
generates images*
A new version that replaced Grok 4.5 in August 2026, 4.6 has increased flagship-level intelligence while maintaining a moderate price. It's up to date on the latest news and social media, and generates images exceptionally well.

Downsides: it hallucinates more than any other flagship, giving a wrong answer to roughly one in three questions that contain a false premise or a trick, and takes a long time to get going before answering. xAI has significantly dialed back content restrictions, and a string of controversies has followed, including deepfake-related incidents. Don't use Grok where audit and compliance requirements are stringent. Whenever the cost of a mistake is high, verify the model's answers independently.
Kimi K3
$$
EfC 3 · Sp 1 · Th 3
Kimi's open flagship-grade model posts the best long-context result among fully tested models: 82.7% on AA-LCR. (Meta Muse Spark 1.2 scores higher on paper (83.3%), but there isn't enough data about it yet for a proper assessment.) However, testers achieved these marks in Max Effort mode, almost certainly not how you'll actually use K3. It reasons like a flagship and costs like a budget model. Use it for large documents — just keep in mind that speed is modest.

Downsides: K3 fails robustness tests in 36% of cases, versus ~8% for Opus 4.8. It keeps pace with the advanced Opus model on simple tasks, but falls sharply behind on complex ones. Users report issues with hosting quality at the provider, and K3 is five times as expensive as its predecessor.
Meta Muse Spark 1.2 (NEW)
EfC 3 · other scores to follow
Meta released Muse Spark 1.2, focused on coding and agentic tasks, on August 5, 2026. With the highest long-context score of any model tested (83.3%), it also offers moderate pricing and a budget tier that's twelve times less expensive than the standard, though you have to give Meta permission to train its models on your data if you want that price.

Downsides: independent speed data isn't yet available, and no outlet has independently confirmed coding benchmarks.

For everyday use

Claude Sonnet 5
$$
EfC 2 · Sp 3 · Th 2
The core Claude model is fast, capable, and affordable. It handles long context well, tackles complex, multi-step tasks with confidence, and writes in a lively, natural voice.
GPT-5.6 Terra
$$
EfC 2 · Sp 3 · Th 2
generates images*
A balanced GPT model at a reasonable price, Terra handles long context well, runs fast, and generates images. Like all OpenAI models, it's more structurally disciplined than those from other companies, making it a solid choice for emails and articles where form matters.
DeepSeek V4 Pro
$
EfC 2 · Sp 2 · Th 2
The standout discovery of our review is DeepSeek V4 Pro, which holds long context at nearly flagship level, runs faster than previously assumed, and costs next to nothing. It's an open model that you can use for analysis and research on a tight budget.

Downsides: censorship is baked into the model and cannot be removed even with self-hosted deployment — previous-generation models (R1) showed ~85% refusal rates on topics the Chinese government dislikes. Consumer service data is stored on servers in China, and regulators in several countries have banned downloading the app or using it for government employees.

Fast and light

Gemini 3.6 Flash
$
EfC 1 · Sp 3 · Th 1
generates images*
The fastest model in the table uses approximately 197 tokens per second by independent measurements. It understands images and has a free tier. Use it for drafts and quick questions.
Gemini 3.7 Flash (NEW)
no scores yet
We have no independent data on Flash's new version yet, but will add scores once the first benchmarks are available.
Claude Haiku 4.5
$$
EfC 1 · Sp 3 · Th 1
Haiku is the fastest and most affordable Claude model. Use it for short responses, summaries, and simple high-volume tasks. There's no independent long-context data available, so the EfC score is preliminary.
GPT-5.6 Luna
$
EfC 1 · Sp 3 · Th 1
generates images*
The most affordable model in the new GPT lineup got even cheaper this summer. Use it for simple tasks: condensing text, sorting, and finding what you need.
GLM 5 Turbo
$
EfC 1 · Sp 3 · Th 1
While fast, cheap, and capable of following instructions precisely, Turbo is modest on the intelligence index. Use it for short, everyday tasks, as it misses details and draws unsupported conclusions on more complex ones. Its quality is inconsistent and highly sensitive to how a prompt is worded, and it only performs well in English and Chinese.

Smart savings

DeepSeek V4 Flash
$
EfC 1 · Sp 2 · Th 1
The cheapest model in the table is still fast. Use it for simple tasks and drafts. The same censorship restrictions as V4 Pro apply, and no independent long-context data is available.
Qwen 3.8 Max
$$
EfC 2 · Sp 0 · Th 3
Alibaba's flagship has strong reasoning and solid long-context retention, per independent benchmarks. It's also modestly priced, and works well for analysis in Asian languages. However, Qwen tends to be verbose and has the lowest generation speed in the table.
GLM 5.2
$
EfC 1 · Sp 3 · Th 2
You can deploy GLM on your own infrastructure, and it scores highest on the independent intelligence index among open models in its price range. It's also fast, so use it if local deployment matters. On the other hand, GLM 5.2 performs much worse in languages other than English and Chinese. There's also no independent long-context data available for it.
MiMo V2.5 Pro
$
EfC 1 · Sp 1 · Th 2
Independent benchmarks confirm middling reasoning and modest speed. The price is as low as it gets. Use it for simple tasks on a tight budget.
Mistral Large 3
$
EfC 1 · Sp 1 · Th 0
A European model, Mistral Large 3 offers broad language support and strict privacy rules from servers in the EU. Its independent intelligence index score is low. Use it when language coverage and privacy are the priority.

Behind the table

Every metric in the table above has a specific source behind it. Artificial Analysis measures intelligence index, AA-LCR long-context retrieval test, speed, and latency. LMArena provides blind comparison: real people pick the best answer without knowing which model produced it. LiveBench offers objective tests designed to be leak-proof, while OpenRouter enables current API price verification.

What didn't fit in the cards but matters

Our review unearthed a few surprising facts. Among fully tested models, Kimi K3 delivers the best long-context performance at a budget price. DeepSeek V4 Pro is significantly underrated, with near-flagship long-context capability for minimal cost. Mistral Large 3 scores low on the independent intelligence index.

On long-context performance, the flagships cluster tightly together. Nearly all score between 74–80% on the AA-LCR test, with gaps within the group comparable to the margin of error. That's why only models with a clear lead earned EfC 3: Kimi K3 and Meta Muse Spark. The rest honestly received EfC 2 — including both top-tier Claude models, despite their first-class reputation in long context.

The AA-LCR test simulates professional document work: finding and assembling an answer from multiple fragments scattered across a large corpus. Each question draws on roughly one hundred thousand tokens of reports, scientific papers, legal texts, and reviews. There are currently no other independent platforms that have measured long-context performance for this generation of models — Fiction.LiveBench, RULER, and LongBench don't cover it. That's a gap worth knowing about.

Results vary by effort mode. At maximum reasoning effort, some models pick up a few points on long-context tasks. However, the scores in the table are based on standard mode, the default setting models run on out of the box.

Why Fable 5 leads with human voters but not on benchmarks

Benchmarks measure specific skills like math, code, science, and instruction-following: narrow tasks with correct answers. LMArena, on the other hand, tracks whether a real person actually likes the model's response. That puts style, structure, tone, the feeling of "intelligence," and readability front and center. Fable 5 is a model Anthropic clearly tuned, in part, for the sheer pleasure of conversation. It writes beautifully, thinks out loud, and lays out its reasoning coherently — it's genuinely easy to talk to.

LMArena also has a documented verbosity bias, as people tend to prefer longer answers even when shorter ones are more accurate. OpenAI's models appear to exploit this actively.

Independent research confirms that the correlation between benchmark ratings and human perception is low. The model you pick as "best" in a blind test may objectively perform worse on benchmarks — it just feels better to talk to. This human-centered approach has real strengths of its own. Which car would you rather drive: the one that gets you from A to B more comfortably and quickly, or the one that posts better lap times at the track?

That doesn't mean benchmarks are bad or LMArena's rankings are wrong. It just means "best model" depends on what you're measuring. That's precisely what this article is about — a single composite score doesn't tell the whole story.

What this means

Which model for which job

For analysis and complex decisions, use Claude Fable 5 or Opus 5.0 if budget matters. For large documents on a tight budget, go with Kimi K3 or DeepSeek V4 Pro. For a steady stream of daily tasks, pick Sonnet 5 or Terra. For drafts and quick turnarounds, you can choose any model from the Fast and light section.

Independent testing upends the conventional wisdom

Some results diverge from what's generally taken for granted. Claude doesn't lead on long context. DeepSeek is stronger than its reputation suggests. Grok hallucinates more than any other flagship.

Limitations worth keeping in mind

We drew on four independent sources. For long context, AA-LCR is the only test available for this year's models — there simply aren't others. Data on niche and newer models is thin, which is reflected directly in the model cards.

The table is a starting point

No table replaces testing on your own workload. Prices shift (GPT and DeepSeek got cheaper this summer, for instance), models get updated, and each has its own quirks. Where you go from here is up to you — your own experience is what counts most.

Data collected in September 2026.

Sources

References cited in this piece. Last verified on the published or revision date.

  1. 01

    Claude Fable 5 Model Review

    www.coderabbit.ai/blog/fable-5-model-review

  2. 02

    My Honest Review of Claude Fable 5

    www.chatprd.ai/how-i-ai/claude-fable-5-review

  3. 03

    Why Claude Fable 5 Struggles With Expert Level Research

    www.newstex.com/blog/why-claude-fable-5-struggles-with-expert-level-research

  4. 04

    Claude Fable 5 Review: What Developers Really Think

    tosea.ai/blog/claude-fable-5-review-developer-reactions

  5. 05

    Claude Fable 5 Review

    techjacksolutions.com/ai-tools/anthropic-claude/claude-fable-5-review

  6. 06

    Claude Fable 5 Review: Is It as Good as They Say?

    nanoskill.ai/blog/claude-fable-5

  7. 07

    OpenAI Fixed GPT-5.6 Sol's Most Frustrating Flaw

    thenewstack.io/sol-usage-limits-reset

  8. 08

    Four Days of GPT-5.6 Sol Failures

    community.openai.com/t/august-23-update-four-days-of-gpt-5-6-sol-failures-new-routing-evidence-long-conversation-changes-safety-guardrail-concerns-and-why-local-workarounds-are-no-longer-an-adequate-answer/1392086

  9. 09

    Mi Experiencia Con 5.6 Sol y Terra

    community.openai.com/t/mi-experiencia-con-5-6-sol-y-terra/1389748

  10. 10

    Gemini 3.1 Pro — Some Users Aren't Impressed

    www.techradar.com/ai-platforms-assistants/gemini/google-just-upgraded-gemini-again-and-3-1-pro-more-than-doubles-its-ai-reasoning-power-but-some-users-arent-impressed

  11. 11

    Gemini 3: Model Card and Safety Framework Report

    thezvi.substack.com/p/gemini-3-model-card-and-safety-framework?r=67wny

  12. 12

    Kimi K3 Benchmarked: Is It Really as Good as the Hype?

    www.mindstudio.ai/blog/kimi-k3-real-world-coding-review

  13. 13

    Kimi K3 Benchmarks: Strong on Paper, Weak on Precision

    semgrep.dev/blog/2026/kimi-k3s-code-security-results-lack-precision

  14. 14

    On Kimi K3: Its Capabilities and Related Discontents

    thezvi.substack.com/p/on-kimi-k3-its-capabilities-and-related

  15. 15

    Kimi K3 Review: Benchmarks, Pricing, and K2 Comparison

    www.buildfastwithai.com/blogs/kimi-k3-review

  16. 16
  17. 17

    xAI Says It Has Fixed Grok 4's Problematic Responses

    techcrunch.com/2025/07/15/xai-says-it-has-fixed-grok-4s-problematic-responses

  18. 18

    Musk Under Fire as Grok's AI Image Tool Sparks Deepfake Nudes Scandal

    www.malaymail.com/news/world/2026/01/10/musk-under-fire-as-groks-ai-image-tool-sparks-deepfake-nudes-scandal-despite-paywall-update/204903

  19. 19

    Grok 4.6 Review: Real Capabilities Assessed (2026)

    www.layer3labs.io/guides/grok-4-6-review

  20. 20

    DeepSeek Is Censored — and Here's What That Actually Means

    www.qwe.edu.pl/tutorial/deepseek-is-censored-what-it-means

  21. 21

    Security and Privacy Risks of DeepSeek (2026)

    techjacksolutions.com/ai-tools/deepseek/deepseek-security-and-privacy-risks

  22. 22

    Factbox: Governments, Regulators Increase Scrutiny of DeepSeek

    finance.yahoo.com/news/factbox-governments-regulators-increase-scrutiny-095326244.html

  23. 23

    Chat Z.AI (GLM-5) Review 2026

    mysummit.school/blog/en/glm5-zai-review-2026

  24. 24

    GLM-5.2's Code Reviews Are Only as Good as Your Prompt

    blog.kilo.ai/p/glm-52s-code-reviews-are-only-as-424

  25. 25

    GLM-5.2 Review 2026

    www.buildfastwithai.com/blogs/glm-5-2-review-2026