MCP serverai.theaggregate/the-aggregate
Fused LLM rankings: one IRT/Elo scale across ~5,000 public benchmark leaderboards, updated daily.
Overview
Score?
UNRATED 0.748
of what a free look can see, on 31 looks
Looks
36
last 11 hr ago
Tools
8
changed 22 days ago
More info
URL
theaggregate.ai/mcp
streamable-http
Says it is
the-aggregate 1.0.1
protocol 2025-06-18
In the record since
32 days ago
Among servers18,413 with a card
0median 0.606 · this server 0.748 · highest on record 0.8561
Toolsfrom sha256:542dabbc29…bfb96f · +0 −0 22 days ago
| Tool | Schema |
|---|---|
| about_the_aggregate What this data is: how the IRT fusion works, current coverage counts, update cadence, and how to cite it. |
input · no output |
| compare_models Head-to-head between 2-4 models: aggregate ranks, Elo gap with a significance note based on the standard errors, and notable benchmarks they share. |
input · no output |
| get_benchmark One benchmark in depth: what it measures, the original source leaderboard URL, IRT stats (difficulty, noise, model coverage), skill weights, and the current top models on it. |
input · no output |
| get_leaderboard Top of the cross-benchmark aggregate ranking: every model placed on one Elo scale by an IRT model fit over public benchmark leaderboards (call about_the_aggregate for the current c |
input · no output |
| get_model One model in depth: aggregate rank, Elo with standard error, provider, what it is, cost per task where known, and its most notable benchmark results (with percentiles). |
input · no output |
| get_prediction_duel Guesswork — the public prediction duel: every day frontier LLMs and The Aggregate's own IRT model predict newly scraped benchmark scores before seeing them, and the errors are scor |
input · no output |
| search_benchmarks Find benchmarks in the aggregate by (partial) name. Returns model coverage, difficulty on the Elo scale, and the benchmark page URL. |
input · no output |
| search_models Find ranked models by (partial) name or provider. Returns rank, Elo and the model page URL. One row per model by default, fused across reasoning-effort settings. |
input · no output |
Verify it yourself
npx teppi-check https://theaggregate.ai/mcpcurl -s https://api.teppi.xyz/v1/trust/mcp/mcs_01M1FZ235RQFJZEZ8C8XPV2H4H