The State of
Artificial Intelligence

27 leading models · 15+ data fields · scenario picks, speed & latency, function-calling, bilingual ability and enterprise features. Sort, filter, and compare up to 4 models side by side. Click a row to open the full profile.

27
Models
15+
Data Fields
6
Sort Fields
4
Compare Max
95.0
Top Score

Who leads the intelligence race?

Switch dimensions to rank models by capability area, use filters and sorting, click a rank badge to add to comparison, and click a row to open the full profile.

Aggregated from public information online and independently weighted · Updated 2026/09/09
Licensing
Region
Price
27/27
#ModelBenchmarksScore
#ModelBenchmarksScore
#ModelBenchmarksScore

Which model is right for you?

Recommendations for different buyer profiles, based on overall score, price, speed and scenario fit.

What the data tells us

★

Claude Fable 5 leads by a wide margin

It tops all three dimensions at an overall 95.0, eight points clear of the runner-up. It is also the priciest ($10/$50) and relatively slow (~30 tok/s), best suited to deep-reasoning workloads where quality outweighs cost.

⚡

DeepSeek, the value champion

DeepSeek-V4-Pro ranks 8th overall at just $0.44/$1.32, roughly 30× the price-performance of Claude Fable 5. Open weights with self-hosting, strong math reasoning and an 82 function-calling score — ideal for budget-constrained production.

◉

Chinese-Mainland multimodal models lead in spots

Doubao Seed 2.1 Pro ranks first globally in video understanding at 89.2, and Qwen3.8-Max offers unmatched multimodal value. Chinese-Mainland models, however, generally trail international models by 15–20 points on English — evaluate carefully for overseas deployments.

Release Timeline
2026.04 — 2026.08 · Gold dots mark board leaders · Scroll sideways for all 22
CN vs Global · Top score by origin
Compared across all three dimensions · Orange = best Chinese-Mainland, deep green = best international
Open vs Closed · Performance gap
Overall dimension · Green = open, black = closed

How we built this index

↻
Cadence
Monthly
Scores and metadata refresh at the start of each month, with immediate additions on major releases
∑
LMArena sample
2M+
Over 2M human blind pairwise votes across 100+ language pairs
▦
AA benchmarks
10+
Artificial Analysis aggregates 10+ standardized benchmarks incl. MMLU, GPQA and HumanEval
Zh
SuperCLUE items
3000+
3000+ items across foundational, professional and Chinese-specific ability dimensions
Score variance: Models within 3 points of each other (e.g. GPT-5.5 and Claude Opus 4.8, both 81) show no statistically significant difference and may trade places on specific tasks. Speed figures (tokens/s, TTFT) are Artificial Analysis measurements under standardized conditions and vary with network, concurrency and prompt length. Function-calling ability is based on BFCL (Berkeley Function Calling Leaderboard). Context effectiveness reflects actual recall from needle-in-a-haystack tests. Enterprise-feature information comes from vendors' official documentation and is ultimately governed by the commercial contract. Models marked with a ⚠ symbol carry specific restrictions; see the expanded row notes. Per-source License and commercial-use status are detailed on the Terms page.
Disclaimer: This is an independent third-party evaluation with no affiliation, sponsorship or partnership with any model vendor (including OpenAI, Anthropic, Google, ByteDance and Alibaba). All scores come from public benchmarks and are independently weighted by Modelspectra; we accept no vendor payment to influence rankings. Scores, prices and performance metrics may change with releases — defer to each vendor's latest official information. The data is provided for reference only and does not constitute investment, procurement or technical-decision advice; decisions based on it are made at the decision-maker's own risk. Model names and trademarks belong to their respective owners and are used here descriptively.
Stay Updated

Monthly model rankings, straight to your inbox

Subscribe to the Modelspectra monthly report: new releases, ranking moves, price changes and selection advice. Free, 1–2 emails a month, unsubscribe anytime.

Please enter a valid email address
By subscribing you agree to our Privacy Policy. We never share your email or send spam.
Compare [0/4]