The Benchmark Picture Is Messier Than Any Single Leaderboard Admits
Grok 4 is being cited as the SWE-bench leader at 75%, GPT-5.5 is reportedly rolling out to ChatGPT and Codex for Plus, Pro, Business, and Enterprise users right now, and two separate comparison articles disagree on Gemini 3.1 Pro’s GPQA Diamond score by more than two percentage points. If you’re trying to pick a model to ship with this week, here is what the numbers actually say — and where you should distrust them.
Transparency note: The data below is drawn from secondary comparison roundups and one unverified LinkedIn post, not primary announcements from Anthropic, OpenAI, Google, xAI, or Kimi. Where sources conflict, both figures are shown. Treat all claims as provisional until confirmed by official documentation.
What’s New in the Last 48 Hours
GPT-5.5 is the only time-sensitive availability claim in current circulation — and it rests on a LinkedIn post rather than an OpenAI blog or API changelog. The reported rollout covers ChatGPT and Codex for Plus ($20/month), Pro ($200/month), Business, and Enterprise tiers. API access is said to be pending. Until OpenAI publishes a primary announcement, treat this as unconfirmed.
No verified launches from Anthropic, Google, xAI, or Kimi have surfaced in the same window. The comparison activity is high, but it is largely recycling existing model versions.
Benchmark Comparison Table
The table below consolidates claims from multiple roundup sources. Conflicts between sources are flagged explicitly — this is not a clean leaderboard, and that ambiguity is itself the story.
| Model | SWE-bench | GPQA Diamond | ARC-AGI-2 | Source Confidence |
|---|---|---|---|---|
| Grok 4 | 75.0% | — | — | Secondary roundup only |
| GPT-5.4 | 74.9% | — | — | Secondary roundup only |
| Claude Opus 4.6 | >74.0% | — | — | Secondary roundup only |
| Gemini 3.1 Pro | — | 91.9% vs 94.3% (sources conflict) | 31.1% vs 77.1% (sources conflict) | Two roundups disagree materially |
| Kimi K2 Thinking | — | — | — | No fresh benchmark data in window |
What changed vs. the previous generation: The SWE-bench cluster at 74–75% represents a meaningful jump from the 50–60% range that defined frontier coding performance twelve months ago. Gemini’s ARC-AGI-2 discrepancy — 31.1% in one source, 77.1% in another — is large enough to suggest either a version mismatch or a methodology difference that neither article explains. That gap should make any builder pause before treating either figure as a purchasing signal.
Pricing and Access Snapshot
- GPT-5.4 / 5.5: Pro plan ~$200/month (consumer); API pricing for 5.5 not yet confirmed
- Gemini 3.1 Pro: Ultra subscription ~$250/month; Gemini 2.5 Flash API at $0.30/million tokens
- Claude Opus 4.6: Max 5x plan ~$100/month; Claude 4.6 Sonnet API at $3/million tokens
- Grok 4: Entry tier ~$30/month; higher-capability tier reported at ~$300/month
- Kimi K2 Thinking: ~$0.60/million tokens — the most cost-competitive reasoning option cited across sources
Why It Matters for Builders
The practical decision is no longer just model quality — it’s rollout status and API availability. GPT-5.5’s reported absence from the API means teams can demo it in ChatGPT but cannot yet integrate it into production pipelines. Kimi K2 Thinking’s $0.60/million-token price point, if accurate, makes it worth a direct evaluation for cost-sensitive inference workloads, particularly given DeepSeek and Kimi’s consistent appearance as near-frontier challengers in recent comparisons.
For agentic and code workflows, Claude Opus 4.6 continues to appear near the top of qualitative rankings across multiple roundups, even where its raw benchmark numbers trail Grok 4 on SWE-bench by less than one percentage point. For reasoning-heavy tasks, Gemini 3.1 Pro’s GPQA Diamond claim is compelling — but the source conflict is too large to ignore without running your own evals.
The Honest Counter-Argument
It’s worth saying plainly: secondary comparison posts have a structural incentive to declare winners, and benchmark scores on SWE-bench or GPQA Diamond do not reliably predict performance on your specific workload. The most useful thing a builder can do with this tracker is use it to shortlist models for internal evaluation — not to skip that evaluation entirely.
