AI accountability news — digests, data trackers, not hype
Radical Transparency

Frontier Model Elo Board

Frontier Model Elo Board — featured image

Frontier Model Release Tracker & Elo Board

Model Release Tracker: every new Claude, GPT, Gemini, Grok, or Kimi drop gets a comparative benchmark table — not just an announcement post. The Elo Board sits underneath as the open-method ranking layer.

What a release card includes

  • What shipped (model family, access tier, modalities).
  • Axes we score separately: reasoning, coding, tool use, refusal quality — never one mystical number alone.
  • What is still unknown (eval leakage, unpublished system prompts, rate limits).

Ranking principles

  • Publish the prompt packs and scoring rubric.
  • Allow reader votes as a secondary signal, never as the only signal.
  • Label illustrative methodology snapshots clearly — we do not invent precise Elo scores as facts.

Update cycle

Board refresh weekly. Mid-cycle model drops get a provisional release card until the next full run. Disputes and rubric changes are logged in public changelogs.