AI accountability news — digests, data trackers, not hype
Radical Transparency
MODELS

Frontier Elo Snapshot

2026-07-23 ·
Frontier Elo Snapshot — featured image

Key line: Model comparisons without a public rubric amount to marketing. The Release Tracker publishes evaluation axes before any aggregate score.

Why it matters: Pitch decks often highlight single scores. Builders need multi-axis cards linked to the Frontier Model Release Tracker & Elo Board.

Methodology preview (illustrative — not live scores)

The table below shows the structure of a release card. Cells reflect qualitative axes for preview only — no precise Elo numbers are presented as facts.

Family (illustrative)ReasoningCodingTool useRefusal qualityNotes
Claude-classStrong long-formCompetitiveSolidConservative defaultProvisional until rubric run
GPT-classBroadStrongStrongMixed by modeWatch rate limits / tiers
Gemini-classMultimodal leanCompetitiveImprovingPolicy-sensitiveContext window claims need tests
Grok-classFast / chattyVariableBrowsing leanPermissive leanSeparate eval for refusal
Kimi-classLong-context leanEmergingEmergingTBDLabel unknowns explicitly

Open methodology

Prompt packs, sampling defaults, and grader notes will ship with the board. Rubric changes will be logged in a changelog. Reader votes will appear beside the board, clearly labeled as popularity signals — not a substitute for structured evals.