Frontier Model Release Tracker & Elo Board
Model Release Tracker: every new Claude, GPT, Gemini, Grok, or Kimi drop gets a comparative benchmark table — not just an announcement post. The Elo Board sits underneath as the open-method ranking layer.
What a release card includes
- What shipped (model family, access tier, modalities).
- Axes we score separately: reasoning, coding, tool use, refusal quality — never one mystical number alone.
- What is still unknown (eval leakage, unpublished system prompts, rate limits).
Ranking principles
- Publish the prompt packs and scoring rubric.
- Allow reader votes as a secondary signal, never as the only signal.
- Label illustrative methodology snapshots clearly — we do not invent precise Elo scores as facts.
Update cycle
Board refresh weekly. Mid-cycle model drops get a provisional release card until the next full run. Disputes and rubric changes are logged in public changelogs.
