| Model & Lab | AA Index | SWE-bench Pro | Terminal-Bench | CursorBench | HLE / Math | Pricing (Prompt / Out) | Context / Output |
|---|---|---|---|---|---|---|---|
|
Claude Opus 5.5
Anthropic (Sept 22, 2026)
|
58 (#1) | 89.9% | 66.4% | 91.2% | 38.4% (31.2% FMath) | $4.00 / $20.00 ($0.20 cached) | 1M (128k out) |
|
Architecture & Features:
Extended MoE with native multimodal reasoning and recursive multi-layer latent workspace. Adaptive Thinking up to 128k+ deliberation tokens.
Throughput & Latency:
TTFT: ~0.95s (uncached), 0.20s (cached) | Throughput: 60–80 tokens/sec. 40% efficiency boost over Opus 5.
Deployment & Access:
Public API, AWS Bedrock, Google Vertex AI. ASL-4 Responsible Scaling Policy compliance. |
|||||||
|
GPT-6 Astra
OpenAI (Sept 3, 2026)
|
53 (#2) | 85.2% | 59.6% | 87.5% | 34.6% (29.5% FMath) | $10.00 / $50.00 ($1.25 cached) | 1.05M (128k out) |
|
Architecture & Features:
Causal MoE with integrated multimodal computer-operator harness & native OS-level action tokens. Dynamic MCTS test-time search with self-verification.
Throughput & Latency:
TTFT: ~1.50s | Throughput: 35–55 tokens/sec. High deliberation depth for autonomous software engineering.
Safety & Access Gate:
OpenAI Preparedness Framework 'Critical' classification. Public API Tier 5, ChatGPT Pro, Enterprise. |
|||||||
|
Gemini 4 Argon
Google DeepMind (Sept 30, 2026)
|
53 (#2) | 86.1% | 60.1% | 88.0% | 36.1% (33.0% FMath) | $2.00 / $10.00 ($0.20 cached) | 1M (1M Out) |
|
Architecture & Features:
Causal Encoder-Decoder MoE co-designed with TPU v6e. Deep tree-search test-time scaling with continuous differentiable latent verification.
Throughput & Latency:
TTFT: ~1.40s | Throughput: 50–75 tokens/sec. 1M-token output capacity for full-repo generation in a single pass.
Access & Cyber Defense:
Gated under Google Fairwind Program for verified critical defense and enterprise partners. Rank #1 Cyber Defense. |
|||||||
|
Grok 4.7
xAI (Sept 21, 2026)
|
49 | 80.4% | 54.8% | 81.7% | 28.7% (22.6% FMath) | $2.00 / $6.00 ($0.50 cached) | 500k (64k out) |
|
Architecture & Features:
Massive MoE trained on Colossus cluster with integrated real-time search & code execution layers. Configurable reasoning effort.
Throughput & Latency:
TTFT: ~0.85s | Throughput: 70–95 tokens/sec. Aggressive price-to-performance ratio for software engineering.
Deployment & Access:
Public API, xAI Cloud, SuperGrok. |
|||||||
|
Gemini 3.8 Flash
Google DeepMind (Sept 10, 2026)
|
46 | 77.2% | 51.4% | 79.1% | 25.2% (19.4% FMath) | $0.75 / $3.75 ($0.075 cached) | 1.05M (65k out) |
|
Architecture & Features:
Linear-attention / Sparse MoE hybrid optimized for extreme batching throughput and sub-agent dispatching.
Throughput & Latency:
TTFT: ~0.22s | Throughput: 200–260 tokens/sec. Industry speed leader.
Specialized Cyber Variant:
Gemini 3.8 Flash Cyber variant leads AutoPatchBench and CyberSOCEval for real-time triage. |
|||||||
|
DeepSeek V4.1 Flash
DeepSeek (Sept 10, 2026)
|
45 (Pareto) | 75.8% | 49.6% | 77.4% | 24.0% (17.5% FMath) | $0.15 / $0.60 ($0.03 cached) | 1M (64k out) |
|
Architecture & Features:
Compressed attention MoE with dynamic expert pruning during inference. Ultra-low KV cache footprint.
Throughput & Latency:
TTFT: ~0.28s | Throughput: 180–240 tokens/sec. 30x cheaper than closed flagships.
Deployment:
Public API / Open Ecosystem (vLLM, Ollama, Together, Fireworks). |
|||||||
|
Meta Muse Spark 1.3
Meta AI (Sept 2, 2026)
|
44 | 73.6% | 48.2% | 75.8% | 22.8% (16.0% FMath) | $1.25 / $4.25 ($0.15 cached) | 1.05M (65k out) |
|
Architecture & Features:
Multimodal reasoning MoE optimized for messy multi-agent orchestration and dynamic tool routing with Llama Guard 5.
Throughput & Latency:
TTFT: ~0.38s | Throughput: 120–160 tokens/sec.
Deployment:
Public API, Llama Stack Partners, Open Weights. |
|||||||
Architectural & Strategic Trends (September 2026)
The Terminal-Bench & Agentic Pivot
Static benchmark saturation has pushed evaluation toward live terminal execution harnesses (Terminal-Bench 4.0) and multi-file IDE refactoring (CursorBench 4.0). Claude Opus 5.5 leads with 66.4% on Terminal-Bench and 89.9% on SWE-bench Pro.
The Cyber-Defense Frontier
Models like Gemini 4 Argon (Google Fairwind Program) and GPT-6 Astra (Preparedness 'Critical') are co-engineered for autonomous vulnerability discovery and instant zero-day patch synthesis, operating under strict cryptographic gating.
Output Horizon Expansion
Gemini 4 Argon's 1M-token output limit allows whole-repository synthesis in a single generation, while DeepSeek V4.1 Flash delivers extreme efficiency ($0.15 / 1M tokens) for continuous agent memory loops.