AI accountability news — digests, data trackers, not hype
Radical Transparency
MODELS

Frontier Model Tracker: GPT-5.5, Gemini 3.1, Claude Opus 4.6, Grok 4, and Kimi K2 Compared on Benchmarks and Price

2026-07-24 · llmwatch_admin
Frontier Model Tracker: GPT-5.5, Gemini 3.1, Claude Opus 4.6, Grok 4, and Kimi K2 Compared on Benchmarks and Price

The Benchmark Picture Is Messier Than Any Single Leaderboard Admits

Grok 4 is being cited as the SWE-bench leader at 75%, GPT-5.5 is reportedly rolling out to ChatGPT and Codex for Plus, Pro, Business, and Enterprise users right now, and two separate comparison articles disagree on Gemini 3.1 Pro’s GPQA Diamond score by more than two percentage points. If you’re trying to pick a model to ship with this week, here is what the numbers actually say — and where you should distrust them.

Transparency note: The data below is drawn from secondary comparison roundups and one unverified LinkedIn post, not primary announcements from Anthropic, OpenAI, Google, xAI, or Kimi. Where sources conflict, both figures are shown. Treat all claims as provisional until confirmed by official documentation.

What’s New in the Last 48 Hours

GPT-5.5 is the only time-sensitive availability claim in current circulation — and it rests on a LinkedIn post rather than an OpenAI blog or API changelog. The reported rollout covers ChatGPT and Codex for Plus ($20/month), Pro ($200/month), Business, and Enterprise tiers. API access is said to be pending. Until OpenAI publishes a primary announcement, treat this as unconfirmed.

No verified launches from Anthropic, Google, xAI, or Kimi have surfaced in the same window. The comparison activity is high, but it is largely recycling existing model versions.

Benchmark Comparison Table

The table below consolidates claims from multiple roundup sources. Conflicts between sources are flagged explicitly — this is not a clean leaderboard, and that ambiguity is itself the story.

Model SWE-bench GPQA Diamond ARC-AGI-2 Source Confidence
Grok 4 75.0% Secondary roundup only
GPT-5.4 74.9% Secondary roundup only
Claude Opus 4.6 >74.0% Secondary roundup only
Gemini 3.1 Pro 91.9% vs 94.3% (sources conflict) 31.1% vs 77.1% (sources conflict) Two roundups disagree materially
Kimi K2 Thinking No fresh benchmark data in window

What changed vs. the previous generation: The SWE-bench cluster at 74–75% represents a meaningful jump from the 50–60% range that defined frontier coding performance twelve months ago. Gemini’s ARC-AGI-2 discrepancy — 31.1% in one source, 77.1% in another — is large enough to suggest either a version mismatch or a methodology difference that neither article explains. That gap should make any builder pause before treating either figure as a purchasing signal.

Pricing and Access Snapshot

  • GPT-5.4 / 5.5: Pro plan ~$200/month (consumer); API pricing for 5.5 not yet confirmed
  • Gemini 3.1 Pro: Ultra subscription ~$250/month; Gemini 2.5 Flash API at $0.30/million tokens
  • Claude Opus 4.6: Max 5x plan ~$100/month; Claude 4.6 Sonnet API at $3/million tokens
  • Grok 4: Entry tier ~$30/month; higher-capability tier reported at ~$300/month
  • Kimi K2 Thinking: ~$0.60/million tokens — the most cost-competitive reasoning option cited across sources

Why It Matters for Builders

The practical decision is no longer just model quality — it’s rollout status and API availability. GPT-5.5’s reported absence from the API means teams can demo it in ChatGPT but cannot yet integrate it into production pipelines. Kimi K2 Thinking’s $0.60/million-token price point, if accurate, makes it worth a direct evaluation for cost-sensitive inference workloads, particularly given DeepSeek and Kimi’s consistent appearance as near-frontier challengers in recent comparisons.

For agentic and code workflows, Claude Opus 4.6 continues to appear near the top of qualitative rankings across multiple roundups, even where its raw benchmark numbers trail Grok 4 on SWE-bench by less than one percentage point. For reasoning-heavy tasks, Gemini 3.1 Pro’s GPQA Diamond claim is compelling — but the source conflict is too large to ignore without running your own evals.

The Honest Counter-Argument

It’s worth saying plainly: secondary comparison posts have a structural incentive to declare winners, and benchmark scores on SWE-bench or GPQA Diamond do not reliably predict performance on your specific workload. The most useful thing a builder can do with this tracker is use it to shortlist models for internal evaluation — not to skip that evaluation entirely.