Model Release Tracker + Elo Board: every new Claude, GPT, Gemini, Grok, or Kimi drop gets a comparative benchmark table — not just an announcement post. The Elo Board is the open-method ranking layer underneath.
What a release card includes
- What shipped (modalities, context, pricing if public).
- Benchmark table vs prior generation and peers (when numbers are published).
- What changed for builders: API, rate limits, eval culture, cost.
- What remains vendor-claimed vs third-party verified.
Honesty rule
We label vendor-reported scores separately from independent leaderboards. Missing cells stay empty — we do not invent Elo.
Start with the latest pieces in Models and practical fallout in Builder Notes.
Live tracker Updated: 2026-09-21T06:41:59Z Top-tier Elo is a tight cluster (1500-1525) led by Claude Fable 5, Claude Opus 5, GPT-5.6 Sol and Kimi K3OSS as of Sep 2026, but absolute scores swing by 100+ points depending on which aggregator or arena variant is cited. Treat cross-source comparisons (main board vs. chatbot-arena summary vs. governance arena) as directional, not equivalent.
Frontier Model Elo Board
Model
Lab
Key score / Elo
Price signal
Builder note
Claude Fable 5 Anthropic 1525 Elo (public leaderboard, Sep 2026) n/d Released Jun 2026; current #1 on this aggregation, lead over #2 is marginal.
Claude Opus 5 Anthropic 1522 Elo (public leaderboard, Sep 2026) n/d Jul 2026 release, statistically near-tied with Fable 5.
GPT-5.6 Sol OpenAI 1514 Elo (public leaderboard, Sep 2026) n/d Jul 2026; no vendor-published benchmark page found, figure is leaderboard-only.
Kimi K3OSS Moonshot AI 1500 Elo (public leaderboard, Sep 2026) n/d Open-weight model cracking the 1500 line, notable given lab tier.
Grok 4.5 xAI 1499 Elo (public leaderboard, second Sep 2026 snapshot) n/d Score differs from other Grok entries below by ~100-200 pts across sources; methodology-sensitive.
Gemini 3.2 Pro Google ~1448 Elo (chatbot arena summary, Jun 2026) n/d Trails the top cluster by 60-80 pts in this cut; older snapshot than main board.
GPT-5.5 Pro OpenAI 1374 (governance arena live ranking, Aug 2026) n/d Different arena/methodology than the main board; not directly comparable to GPT-5.6 Sol figures above.
Grok 4 xAI 1304-1385 depending on source (governance arena vs. chatbot arena summary) n/d Widest spread of any model tracked here, flags high sensitivity to ranking methodology.
