Landing / Model scorecard
Which AI model is best at what?
Six skills, scored 0–100 every morning from public benchmarks. Every number shows its sources and dates. Pick a task and get a ranked list.
Loading the latest scores…
The tool
What do you need?
Choose a task, or mix the skills yourself. Add a budget or keep to open weights.
Set how much each skill matters, from 0 (ignore) to 5 (essential).
Thin scores rest on a single benchmark and are marked. Self-reported scores are the provider's own numbers, left out unless you include them.
Head to head
Compare two or three models.
Choose up to three. The shape shows where each one is strong.
Change log
New this week.
New models, moved scores and sources that need attention, from the latest daily run.
Method
How the scores work.
Crawled from public pages and datasets every day at about 04:00. No source is published until its licence and crawl permission are checked.
Four tiers of evidence
- 1Structured public files
Benchmark datasets published as files, with a clear licence.
- 2Official leaderboards
Each benchmark's own results page, read within its robots.txt.
- 3Provider model cards
Self-reported. Marked provisional until a tier 1 or tier 2 source confirms it.
- 4AERIS field results
How often AERIS work lanes pass, per model and task. Opt-in, aggregated, no content.
From benchmark to score
- Percentile per benchmark. A model's score is its percentile among every model measured on that benchmark in the last twelve months. 100 is the best measured, 50 the middle.
- Mean per skill. A skill score is the weighted mean of its benchmarks. With only one benchmark it is shown, but flagged as thin.
- Evidence-weighted order. Rankings are ordered by an evidence-weighted score: each benchmark behind a skill adds weight, and a single benchmark pulls the score toward 50. The number shown is the plain score.
- Your ranking. The tool weights the skills you choose. A skill a model has not been measured on counts as zero, and the row says so. Thin scores count but are marked; provisional scores are left out unless you include self-reported numbers, and a skill with fewer than five measured models waits until there is enough public data.
Confidence
- High Two or more independent benchmarks.
- Medium Two or more, with older or less consistent evidence.
- Thin One benchmark only. Counted, marked, and pulled toward 50 when ranking.
- Provisional The provider's own numbers, not yet confirmed. Shown when you include self-reported.
Sources and attribution
Every source the current scores draw on, with its licence and the date it was last read.
Inside AERIS OS
How AERIS uses this.
Rolling out in AERIS OS
AERIS agents hand work to other models. This feed is how they choose: coding goes to a strong coding model, computer use to a strong computer-use model, from the models you already have.
Your preference always wins. Pin a model for a kind of task and AERIS uses it, whatever the scores say. Scores are advice, never an override.
About AERIS OS