Which model actually earns its keep?
We run every model through the work Envoy does every day — answering from policy documents, building quizzes, writing reports, converting messy PDFs, calling tools, closing the books. Deterministic grading where the answer can be checked by code; where it can't, three separate AI models mark every answer blind and we publish how much they disagreed. The task set is frozen as v1; new models join the same board as they ship, so scores stay comparable.
Performance by scenario
One eval, four weightings. Each scenario reweights the nine categories to match a real workload; the cards show the strongest results on this eval under each — not a verdict on the models. Click a card to score the whole page that way.
Leaderboard
| # | Model | Weights | Score | Categories | Standout | Run cost | Median latency | Reliability |
|---|
The pretty graphs
How much does cost matter to you?
The leaderboard assumes cost is no object and the cheapest-first view assumes quality isn't. Neither is true for a real budget. Drag the slider to say how much cost should count, and the board re-ranks between the two.
Does the winner depend on the weighting?
Each model's best and worst rank across the four scenarios. A short bar is a model that holds its place whatever you ask of it; a long bar is a specialist whose position depends entirely on the job. The filled dot is where it sits under the weighting you have selected.
Every score on the board
Nine categories, every model, darker is better. Hatched cells are n/a — a category the model could not attempt, dropped from its score rather than counted as a failure. Sorted by the weighting you have selected.
swipe the grid sideways for the remaining categories →By category
Top three per category, independent of any scenario weighting. This is the raw material the scenarios are built from.
Open vs proprietary
What a year bought you
A caution on reading the performance line as pure capability: two of the nine categories reward knowing what has happened recently. A model released last year has an older knowledge cutoff and loses ground on the report task for that reason alone, so part of the slope is recency rather than reasoning. The cost line has no such problem — that is a clean measure of what the same quality costs over time.
Checking the graders for bias
The five open-ended items — two quizzes, a report, a summary and a recommendation — are about a third of the composite, and they are the part a model could be flattered on. So they are marked blind by four different models, under shuffled anonymous labels, and the labels are only mapped back after every score is locked. All four graders have a model on the board, and that is the point of the audit: it lets us measure whether a grader is kind to itself or to its own family.
The audit says no. Each grader marks its own answers within +0.07 of what the other three give them, on a 0–5 scale; the two Claude graders show no lean toward Claude Fable 5.1; the one family lean in the matrix — GPT-5.6 Sol marking GPT-6 Astra +0.24 — is cancelled by Qwen marking the same answers −0.24. Averaging four graders keeps any single leaning to a quarter of its weight. The full grader-by-model matrix is below for anyone who wants to check; it is an audit of the marking, not a ranking of the models — a model that leads on the checkable categories can sit well down it.
Show the full grader-by-model matrix
Latency
Median per-item response time across the whole run — documents, images and tool calls included, so this is workload latency, not time-to-first-token marketing.
How the eval works
Release dates. Each model shows when it was released, taken from its first appearance in the OpenRouter catalogue. It is there so you can weigh a score against how old the model is — a year-old model holding its own against last month's is a different result from the same score fresh out.
Versioned, not rounds. The 37-item task set is frozen as v1. New models are added to this board as they ship and stay directly comparable with everything already on it. When the tasks themselves change, that becomes v2 — a fresh board, not a reshuffle of this one.
Nine categories, five weightings. Every model gets one score per category. The scenarios are different weightings over the same nine numbers, published so you can check them or make your own. Per item is the derived reference: no editorial judgement at all, every task counting equally, so a category is worth exactly as many items as it holds.
The published weighting follows discrimination, and changed on 5 August 2026. Weight now sits where models actually differ. Grounded QA has twelve items but 53 of 56 models score a perfect 1.00 on it — it is a gate everyone passes, and weighting it heavily measured very little. Programming, problem solving and report writing spread the field by 0.63 to 0.75 and were carrying less weight than they revealed. So qa fell from 18% to 10% and retrieval from 14% to 10%, while problem solving rose to 16% and programming to 14%. The board separates models better as a result — the spread across the field widened from 0.33 to 0.40.
What that change did, stated plainly. It moved 46 of 56 models, and it moved Claude Fable 5 from third to first. We are telling you that because a weighting revised after seeing results, in a direction that happens to restore a previous leader, is exactly the thing you should be suspicious of. The rule was fixed before the board was recomputed, it is stated above, and the per-item view is published alongside so you can see the board with no editorial weighting at all. It also made the eval less precise per item, not more: with weight on categories holding fewer tasks, dropping any single item now moves a score by ±0.019 rather than ±0.014.
Deterministic first. Everything numeric, exact-match, table F1, the Ruby test suite and forecast RMSE are graded by code. Confabulation traps — questions the source can't answer — are scored on whether the model says so. Code has format assumptions, and they get found: on 5 September 2026 the quiz checker was found to count an answer key's numbered lines as extra questions, to miss options written as Markdown bullets, and to miss explanations given inline after each keyed answer. Fixed, it raised quiz scores on 46 items across 41 models, by up to 0.375 on a single item, and moved 39 composites — all upward, none by more than 0.037. The graders' rubric marks were unaffected, since they read the quizzes rather than parsing them.
Blind where it matters. Open-ended items are reviewed under shuffled anonymous labels by graders who have never seen the answer key, scored 0–5 on fixed rubric dimensions, and only unblinded after scores are locked.
Four graders, not one. Every open-ended item is scored independently by four models — Claude Fable 5, Claude Opus 5, GPT-5.6 Sol and Qwen 3.8 Max — each on the same shuffled pack, and the published rubric score is the mean of the four. On the main pack every pair of graders agrees between 0.71 and 0.88; on the September pack, between 0.75 and 0.89. The two Claude graders are closest to each other both times. Adding the fourth grader moved seventeen of the fifty-six models then on the board — eleven by a single place, six by more — and changed the average score by 0.001. Which model marks the work is not what determines the ranking.
Eight models were added in September 2026 and graded against anchors. Re-marking all sixty-four candidates for every addition would cost more than it tells, so the eight new models were reviewed as their own blind pack: their answers plus ten anchor candidates the panel had already scored in August, chosen to span the board from Llama 4 Scout to Claude Opus 5 and reshuffled under fresh labels. The graders were not told which was which. Comparing each grader's two scores for the same anchor text measures how the panel's scale moved between packs: it marked the September pack 0.22 points stricter on the 0–5 scale — Fable 5 by 0.09, Sol by 0.17, Opus by 0.27, Qwen by 0.35 — while agreeing with its August self at r = 0.82–0.88. The new models are published as marked, not corrected. Adding the shift back would raise their composites by 0.003 to 0.008, all inside the eval's resolution of ±0.019; it would move Gemini 3.8 Flash up two places, Grok 4.6, GLM 5.3 Flash and Muse Spark 1.3 up one, and change nothing at the top. The anchors' August scores stand; their September marks were used only for this comparison.
We checked the graders for self-interest. All four graders have a model on this board, so we measured whether a grader scores its own model's answers above what the others give them. Across the twenty rubric scores each gives itself, the gap averages +0.04 on a 0–5 scale — Opus +0.05, Sol +0.06, Qwen +0.07, and Claude Fable 5 scores itself exactly what the others do. Family is the sharper test now that the graders' successors are on the board. GPT-5.6 Sol marks GPT-6 Astra +0.24 above the panel, the largest lean in the matrix, while Qwen marks the same answers −0.24 below it; the two cancel in the average. The Claude graders show nothing of the kind toward Claude Fable 5.1: Fable 5 marks it −0.07, Opus +0.07. Averaging four graders keeps any one of those leanings to a quarter of its weight, and it is why we publish the panel rather than a single judge.
One grader marks through the API, not the repository. Qwen 3.8 Max reads each pack as a single request with no access to this repository, so it could not reach the answer key or the label map even had it tried — the blinding is structural rather than promised. It is also the strictest of the four, marking well below the panel on average. At fifty-six candidates it could no longer hold a whole pack in one pass on the quiz item, so that pack was graded in groups of eight against the same rubric; a grader comparing eight answers at a time sees a narrower field than one comparing fifty-six, which is a real difference from the other three and worth knowing when reading the spread. The eighteen-candidate September pack it read whole.
Honest n/a. Text-only models aren't punished for vision items; endpoints without function calling aren't punished for tool items. Those categories are dropped and the score reweighted — refusals and errors, though, score zero.
Re-runs, and why. The first pass had two faults of our own making. A token cap set too low let reasoning models spend their whole budget thinking and return nothing visible, which scored zero on 30 items across 15 models; and a provider guardrail truncated some responses mid-sentence. Both are properties of the harness and the route, not of the models, so the affected items were re-run with a higher cap and the scores below reflect the second attempt. Every re-run item is counted in the table, and the raw first-pass results are kept.
Where a route failed entirely. A handful of items kept returning truncated output on every provider that serves that model. Rather than score a model zero for its router, those answers were produced through a direct session. They are not like-for-like with the provider runs — different plumbing, no document attachment path — and the models carrying them are listed below.
Three models were answered in agent sessions, not through the API. Claude Fable 5.1, GPT-6 Astra and Gemini 3.8 Flash were run in September 2026 as Claude Code, Codex and Antigravity sessions respectively. Each session was handed a pack exported from the harness — the same system prompt, the same user message, the same attached documents, data inlined exactly where the API path inlines it — and wrote one reply per item under rules forbidding web search, code execution and any file outside the pack. The three tool-use items were answered through a call shim exposing the same four functions the API models receive, with every call logged. What differs is the wrapper: an agent harness around the model instead of a bare API call, no metered usage, and no timing. So those rows carry a session run label, their run cost is taken from the tokens the model's own predecessor spent on each item through the API, priced at the new model's list price, so the tokeniser and the handling of an attached PDF are the provider's own, and they sit out of the latency chart. Treat their scores as the same tasks under a different harness, and read the gap to their API-run siblings with that in mind.
When a model said nothing, we found out why. Three items came back empty: the request succeeded, the whole completion budget went on hidden reasoning, and no answer appeared. Raising the budget did not help — it only bought more thinking. Capping the reasoning instead, so the model had room left to reply, produced an answer on all three within a few thousand tokens. Those caps are recorded per model in the registry and applied to every item that model runs, not just the one that failed. It is worth knowing that a reasoning model can be unable to answer at any budget and answer easily at a smaller one. The September additions needed the same treatment — six empty items across GLM 5.3, GLM 5.3 Flash and Qwen 3.8 Flash, all answered once reasoning was capped, and two of those routes ignored an effort level and honoured only a token budget. One took a fifth attempt: GLM 5.3 on the inline forecast item returned nothing at 32,000 tokens on four attempts across three providers, all of which ignored the cap. Reasoning cannot be switched off on that model at all — the endpoints reject the request — but a fourth provider honours the budget and the model answered in under 5,000 tokens. That route is recorded on the item and the answer is scored like any other.