Artificial Analysis & Benchmarks

Frontier AI Intelligence & Benchmark Matrix

Standardized evaluation metrics tracking frontier foundation models across coding benchmarks, reasoning accuracy, latency throughput, and token economic efficiency.

Artificial Intelligence Matrix

Frontier AI Models Benchmark Leaderboard

Independent evaluation of quality index, coding benchmarks, throughput, and token pricing.

# Model
Quality Index
Coding
Price / 1M
1
Claude Fable 5
Anthropic
98.4
84.8%
$90.00
in: $18.00
2
Claude Opus 5 (max)
Anthropic
97.8
83.5%
$75.00
in: $15.00
3
GPT-5.6 Sol (max)
OpenAI
97.5
84.2%
$48.00
in: $12.00
4
Grok 4.6 (high)
xAI
96.9
81.8%
$20.00
in: $5.00
5
Kimi K3 (max)
Moonshot AI
96.5
80.5%
$16.00
in: $4.00
6
GLM-5.3
Zhipu AIOpen Weights
96.2
79.8%
$2.00
in: $0.50
7
GPT-5.6 Terra (max)
OpenAI
96
82.8%
$24.00
in: $6.00
8
Stealth Ox Alpha
Anonymous
95.8
81.4%
$0.00
in: $0.00
9
DeepSeek V4-Pro
DeepSeekOpen Weights
95.5
78.5%
$1.10
in: $0.28
10
Gemini 3.7 Flash
Google DeepMind
94.2
75.1%
$0.40
in: $0.10
Interactive Token Economics Calculator

Frontier AI Monthly API Cost Simulator

Estimate monthly inference costs across foundation models based on your projected prompt and generation volume.

Monthly Volume
10M tokens/mo
Total Volume (Input + Generation)10 Million
1M ($)25M50M100M tokens
Input vs Output Token Ratio75% In · 25% Out
10% In (Chat)50:50 (Code gen)90% In (RAG/Doc)
Model & ArchitectureEst. Monthly Cost
Stealth Ox AlphaAnonymousLowest Cost
$0.00/mo
In: $0/1M·Out: $0/1M
Score: 95.8·78 t/s
Gemini 3.7 FlashGoogle DeepMind
$1.75/mo
In: $0.1/1M·Out: $0.4/1M
Score: 94.2·175 t/s
DeepSeek V4-ProDeepSeek
$4.85/mo
In: $0.28/1M·Out: $1.1/1M
Score: 95.5·92 t/s
GLM-5.3Zhipu AI
$8.75/mo
In: $0.5/1M·Out: $2/1M
Score: 96.2·82 t/s
Kimi K3 (max)Moonshot AI
$70.00/mo
In: $4/1M·Out: $16/1M
Score: 96.5·76 t/s
Grok 4.6 (high)xAI
$87.50/mo
In: $5/1M·Out: $20/1M
Score: 96.9·85 t/s
GPT-5.6 Terra (max)OpenAI
$105.00/mo
In: $6/1M·Out: $24/1M
Score: 96·90 t/s
GPT-5.6 Sol (max)OpenAI
$210.00/mo
In: $12/1M·Out: $48/1M
Score: 97.5·72 t/s
Claude Opus 5 (max)Anthropic
$300.00/mo
In: $15/1M·Out: $75/1M
Score: 97.8·68 t/s
Claude Fable 5Anthropic
$360.00/mo
In: $18/1M·Out: $90/1M
Score: 98.4·62 t/s

Model Profiles & Analysis

AnthropicScore: 98.4/100

Claude Fable 5

Anthropic's flagship frontier model leading the Artificial Analysis Intelligence Index with adaptive reasoning and Mythos guardrails.

#1 on Artificial Analysis Intelligence Index with adaptive max reasoning
Highest scores on AA-Omniscience knowledge and hallucination benchmarks
Dynamic safety fallback mechanism routing to Claude Opus
Context Window
1000k tokens
Output Price
$90/1M
AnthropicScore: 97.8/100

Claude Opus 5 (max)

The leader in autonomous enterprise knowledge work, ranking #1 on the AA-Briefcase agentic benchmark.

#1 on AA-Briefcase benchmark for multi-hour autonomous knowledge work
500,000 continuous reasoning tokens for repository-level refactoring
Unmatched precision across legal synthesis, medicine, and STEM
Context Window
500k tokens
Output Price
$75/1M
OpenAIScore: 97.5/100

GPT-5.6 Sol (max)

OpenAI's flagship GPT-5.6 frontier system, leading the Artificial Analysis Coding Agent Index with 1M context.

#1 on Artificial Analysis Coding Agent Index across complex repositories
1,000,000 token context window with native test-time compute scaling
Seamless integration with autonomous sandboxed Python environments
Context Window
1000k tokens
Output Price
$48/1M
xAIScore: 96.9/100

Grok 4.6 (high)

xAI's newly released frontier model trained on the expanded Colossus megacluster, tied for #3 on the Intelligence Index.

Tied for #3 on Artificial Analysis Intelligence Index v4.1.1
Ultra-low latency reasoning powered by next-gen Colossus GPU fabric
Native real-time search grounding and multimodal video comprehension
Context Window
500k tokens
Output Price
$20/1M
Moonshot AIScore: 96.5/100

Kimi K3 (max)

Moonshot AI's 2.8-Trillion parameter MoE powerhouse ranking in the top 5 of the Artificial Analysis Intelligence Index.

2.8 Trillion total parameters with advanced reinforcement learning
Top 5 on Artificial Analysis global Intelligence Index
1M context window specialized in long-horizon scientific agents
Context Window
1000k tokens
Output Price
$16/1M
Zhipu AIScore: 96.2/100

GLM-5.3

Zhipu AI's upgraded reasoning powerhouse recognized in the Artificial Analysis Index with score 60 and 1M context.

Scored 60 on Artificial Analysis Intelligence Index v4.1.1
1,048,576 token context window with specialized cybersecurity auditing
Open weights scheduled for public release in late August
Context Window
1049k tokens
Output Price
$2/1M
AnonymousScore: 95.8/100

Stealth Ox Alpha

Anonymous frontier reasoning powerhouse with 1M context and 131K output capacity; fingerprinted as Zhipu GLM-5.3.

Massive 1M token context paired with 131K max continuous output tokens
Achieves 81.4% Pass@1 on DeepSWE software engineering benchmarks
Native multimodal video and high-resolution document processing
Context Window
1049k tokens
Output Price
$0/1M
OpenAIScore: 96/100

GPT-5.6 Terra (max)

The cost-optimized tier in OpenAI's GPT-5.6 series delivering near-Sol intelligence at half the API price.

Positions near GPT-5.6 Sol on coding benchmarks at 50% lower price
Full 1,000,000 token context window for large codebase refactoring
Sub-350ms TTFT for responsive agentic loops
Context Window
1000k tokens
Output Price
$24/1M
DeepSeekScore: 95.5/100

DeepSeek V4-Pro

1.6T parameter MoE architecture utilizing DualPipe v2 communication overlap and fractional token pricing.

1.6T total parameters activating 128B per token
DualPipe v2 eliminates multi-node GPU interconnect bottlenecks
Commercially permissive open weights on Hugging Face
Context Window
1000k tokens
Output Price
$1.1/1M
Google DeepMindScore: 94.2/100

Gemini 3.7 Flash

Google DeepMind's workhorse high-throughput model optimized for sub-200ms real-time audio and live video streams.

Blazing 175 tokens/sec median generation throughput
Sub-150ms TTFT latency for live conversational applications
1M token multimodal context window at $0.10 / $0.40 pricing
Context Window
1000k tokens
Output Price
$0.4/1M