Key line: Model comparisons without a public rubric amount to marketing. The Release Tracker publishes evaluation axes before any aggregate score.
Why it matters: Pitch decks often highlight single scores. Builders need multi-axis cards linked to the Frontier Model Release Tracker & Elo Board.
Methodology preview (illustrative — not live scores)
The table below shows the structure of a release card. Cells reflect qualitative axes for preview only — no precise Elo numbers are presented as facts.
| Family (illustrative) | Reasoning | Coding | Tool use | Refusal quality | Notes |
|---|---|---|---|---|---|
| Claude-class | Strong long-form | Competitive | Solid | Conservative default | Provisional until rubric run |
| GPT-class | Broad | Strong | Strong | Mixed by mode | Watch rate limits / tiers |
| Gemini-class | Multimodal lean | Competitive | Improving | Policy-sensitive | Context window claims need tests |
| Grok-class | Fast / chatty | Variable | Browsing lean | Permissive lean | Separate eval for refusal |
| Kimi-class | Long-context lean | Emerging | Emerging | TBD | Label unknowns explicitly |
Open methodology
Prompt packs, sampling defaults, and grader notes will ship with the board. Rubric changes will be logged in a changelog. Reader votes will appear beside the board, clearly labeled as popularity signals — not a substitute for structured evals.
