The LLM Capability Frontier

Major large-language-model releases plotted by public release date (x) and capability (y). Capability is the Artificial Analysis Intelligence Index, a 0–100 aggregate of hard reasoning/coding/agentic benchmarks. Each point is the releasing company's logo.

⚠︎ Read the y-axis carefully. Scores are the current Intelligence Index v4.1 (July 2026 snapshot). v4.1 is far harder than the index versions in use when older models launched, so a model's number here is well below its launch-day headline (e.g. Grok 4 was ~68 at launch → 33 on v4.1). All solid markers are measured on v4.1. Pre-2025 models predate v4.1 and are shown at estimated positions (dashed ring) so the timeline still starts in 2022. Microsoft's MAI models have no published AA score and are shown at estimated positions too — flagged below and in tooltips.

solid ring = measured (v4.1) dashed ring = estimated
ModelCompany ReleasedIndex BasisNotes / source
Sources & method. Capability = Artificial Analysis Intelligence Index v4.1 (artificialanalysis.ai, snapshot 18 Jul 2026, 0–100 scale). Release dates verified against vendor announcements, Wikipedia, and press (OpenAI, Anthropic, Google, Meta, xAI, Microsoft, Mistral, DeepSeek, Alibaba, Moonshot, Zhipu/Z.ai).  Basis legend: measured = published v4.1 score; estimated = model positioned approximately on the same scale — either pre-v4.1 (older models score low on today's harder benchmark) or lacking a published AA score (Microsoft MAI); directional, not official. The dashed frontier line traces the highest capability reached by any model up to each date. Because the index was re-versioned, this chart shows capability on today's yardstick over time, not each model's launch-day score.