AI Model Tier List

Last updated September 16, 2026

A snapshot ranking of the current flagship model from each major lab into rough tiers, based on general capability, reliability, and value — not a single benchmark leaderboard. Each entry names the specific model AInformed's own daily coverage has tracked as that lab's current flagship, not a generic placeholder, so it stays tied to real releases rather than a vague family name.

S Tier

  • OpenAI's current agentic flagship — picked up by third-party coding and automation tools like Devin and Perplexity within weeks of release.

  • Anthropic's most capable model to date, built for complex agentic workflows; the cheaper Opus 5 trades some of Fable's capability for lower cost.

  • Very large context window and the deepest integration across Google's own products, from Search to Workspace.

A Tier

  • Ships with a dedicated terminal-based coding assistant, 'Grok Build' — leans hardest into agentic coding and knowledge work of this group.

  • DeepSeek-V4DeepSeek

    Open-weight with a 1-million-token context window; independent testing has it matching GPT-5-class performance at a fraction of the cost to run.

B Tier

  • Mistral 2.0Mistral AI

    Efficient and cost-effective, though it trails the frontier labs above on complex, multi-step reasoning.

  • Still the most widely fine-tuned open-weight base for self-hosted deployments, but Meta hasn't shipped a frontier-class successor while rivals iterated past it.

C Tier

  • Smaller open-weight models (7B–13B class)various

    Good for narrow, well-defined tasks and cheap self-hosting, but noticeably weaker on multi-step reasoning than the larger models above.

Methodology

Tiers are assigned based on general reasoning ability, instruction-following, context window, and real-world tooling maturity, weighted toward how a typical user or developer would experience the model rather than a single benchmark score. Each lab's entry is updated to whatever model AInformed's own coverage most recently confirmed as that lab's current flagship — when a new one ships, this list is revisited rather than left pointing at a superseded release.

Frequently asked

What does 'S tier' mean on this list?
S tier is today's best-in-class general capability across reasoning, following instructions, and ecosystem maturity — GPT-6 Astra, Claude Fable 5.1, and Gemini 3.5 as of this update — not necessarily the top score on any single benchmark.
How often is this tier list updated?
It's revisited whenever a lab ships a new flagship that AInformed's own coverage confirms — check the "last updated" date at the top of the page for how current it is.
Why is DeepSeek-V4 ranked below GPT-6 Astra if it's open-weight and cheaper?
Independent testing puts DeepSeek-V4 at roughly GPT-5-class performance, a tier below OpenAI's current GPT-6 Astra flagship — very strong for the price, but not yet matching the frontier leaders on hardest-case reasoning.
Is Claude Opus 5 or Claude Fable 5.1 better?
Fable 5.1 is Anthropic's most capable model and the one ranked here; Opus 5 is a real alternative but is explicitly positioned by Anthropic as a cheaper, less capable option than Fable.