AI accountability news — digests, data trackers, not hype
Radical Transparency

Frontier Model Elo Board

Frontier Model Elo Board — featured image

Model Release Tracker + Elo Board: every new Claude, GPT, Gemini, Grok, or Kimi drop gets a comparative benchmark table — not just an announcement post. The Elo Board is the open-method ranking layer underneath.

What a release card includes

  • What shipped (modalities, context, pricing if public).
  • Benchmark table vs prior generation and peers (when numbers are published).
  • What changed for builders: API, rate limits, eval culture, cost.
  • What remains vendor-claimed vs third-party verified.

Honesty rule

We label vendor-reported scores separately from independent leaderboards. Missing cells stay empty — we do not invent Elo.

Start with the latest pieces in Models and practical fallout in Builder Notes.


Live tracker

Frontier Model Elo Board

Updated: 2026-09-21T06:41:59Z

Top-tier Elo is a tight cluster (1500-1525) led by Claude Fable 5, Claude Opus 5, GPT-5.6 Sol and Kimi K3OSS as of Sep 2026, but absolute scores swing by 100+ points depending on which aggregator or arena variant is cited. Treat cross-source comparisons (main board vs. chatbot-arena summary vs. governance arena) as directional, not equivalent.

Model Lab Key score / Elo Price signal Builder note
Claude Fable 5Anthropic1525 Elo (public leaderboard, Sep 2026)n/dReleased Jun 2026; current #1 on this aggregation, lead over #2 is marginal.
Claude Opus 5Anthropic1522 Elo (public leaderboard, Sep 2026)n/dJul 2026 release, statistically near-tied with Fable 5.
GPT-5.6 SolOpenAI1514 Elo (public leaderboard, Sep 2026)n/dJul 2026; no vendor-published benchmark page found, figure is leaderboard-only.
Kimi K3OSSMoonshot AI1500 Elo (public leaderboard, Sep 2026)n/dOpen-weight model cracking the 1500 line, notable given lab tier.
Grok 4.5xAI1499 Elo (public leaderboard, second Sep 2026 snapshot)n/dScore differs from other Grok entries below by ~100-200 pts across sources; methodology-sensitive.
Gemini 3.2 ProGoogle~1448 Elo (chatbot arena summary, Jun 2026)n/dTrails the top cluster by 60-80 pts in this cut; older snapshot than main board.
GPT-5.5 ProOpenAI1374 (governance arena live ranking, Aug 2026)n/dDifferent arena/methodology than the main board; not directly comparable to GPT-5.6 Sol figures above.
Grok 4xAI1304-1385 depending on source (governance arena vs. chatbot arena summary)n/dWidest spread of any model tracked here, flags high sensitivity to ranking methodology.
Sources
  • Public leaderboard, LMArena-style aggregation, Sep 2026 snapshot [1]
  • Public leaderboard, chatbot arena summary, Jun 2026 [2]
  • Public leaderboard, LMArena-style aggregation, second Sep 2026 snapshot [3]
  • Governance arena live ELO ranking, Aug 2026 [4]