E:MERGENCELanding

Landing / Model scorecard

Which AI model is best at what?

Six skills, scored 0–100 every morning from public benchmarks. Every number shows its sources and dates. Pick a task and get a ranked list.

Loading the latest scores…

The tool

What do you need?

Choose a task, or mix the skills yourself. Add a budget or keep to open weights.

Head to head

Compare two or three models.

Choose up to three. The shape shows where each one is strong.

Change log

New this week.

New models, moved scores and sources that need attention, from the latest daily run.

    Method

    How the scores work.

    Crawled from public pages and datasets every day at about 04:00. No source is published until its licence and crawl permission are checked.

    Four tiers of evidence

    1. 1
      Structured public files

      Benchmark datasets published as files, with a clear licence.

    2. 2
      Official leaderboards

      Each benchmark's own results page, read within its robots.txt.

    3. 3
      Provider model cards

      Self-reported. Marked provisional until a tier 1 or tier 2 source confirms it.

    4. 4
      AERIS field results

      How often AERIS work lanes pass, per model and task. Opt-in, aggregated, no content.

    From benchmark to score

    Five models on one benchmark. The best scores 100, the middle one 50, the lowest 0. 0255075100 lowestbest measured
    1. Percentile per benchmark. A model's score is its percentile among every model measured on that benchmark in the last twelve months. 100 is the best measured, 50 the middle.
    2. Mean per skill. A skill score is the weighted mean of its benchmarks. With only one benchmark it is shown, but flagged as thin.
    3. Evidence-weighted order. Rankings are ordered by an evidence-weighted score: each benchmark behind a skill adds weight, and a single benchmark pulls the score toward 50. The number shown is the plain score.
    4. Your ranking. The tool weights the skills you choose. A skill a model has not been measured on counts as zero, and the row says so. Thin scores count but are marked; provisional scores are left out unless you include self-reported numbers, and a skill with fewer than five measured models waits until there is enough public data.

    Confidence

    • High Two or more independent benchmarks.
    • Medium Two or more, with older or less consistent evidence.
    • Thin One benchmark only. Counted, marked, and pulled toward 50 when ranking.
    • Provisional The provider's own numbers, not yet confirmed. Shown when you include self-reported.

    Sources and attribution

    Every source the current scores draw on, with its licence and the date it was last read.

      Inside AERIS OS

      How AERIS uses this.

      Rolling out in AERIS OS

      AERIS agents hand work to other models. This feed is how they choose: coding goes to a strong coding model, computer use to a strong computer-use model, from the models you already have.

      Your preference always wins. Pin a model for a kind of task and AERIS uses it, whatever the scores say. Scores are advice, never an override.

      About AERIS OS
      Illustration · with the current data