Endpoints: 28,729MCP servers: 18,413Payout addresses: 2,071Paid calls: 1,537Letters: 14Defects: 1,323counted 4 min ago
teppi

MCP serverai.theaggregate/the-aggregate

Fused LLM rankings: one IRT/Elo scale across ~5,000 public benchmark leaderboards, updated daily.
UNRATEDActivestreamable-httptheaggregate.ai

Overview

Score?
UNRATED 0.748
of what a free look can see, on 31 looks
Looks
36
last 11 hr ago
Tools
8
changed 22 days ago

More info

URL
theaggregate.ai/mcp
streamable-http
Says it is
the-aggregate 1.0.1
protocol 2025-06-18
In the record since
32 days ago

Among servers18,413 with a card

0median 0.606 · this server 0.748 · highest on record 0.8561

Toolsfrom sha256:542dabbc29…bfb96f · +0 −0 22 days ago

The tools this server lists, read out of the definition it returned
ToolSchema
about_the_aggregate
What this data is: how the IRT fusion works, current coverage counts, update cadence, and how to cite it.
input · no output
compare_models
Head-to-head between 2-4 models: aggregate ranks, Elo gap with a significance note based on the standard errors, and notable benchmarks they share.
input · no output
get_benchmark
One benchmark in depth: what it measures, the original source leaderboard URL, IRT stats (difficulty, noise, model coverage), skill weights, and the current top models on it.
input · no output
get_leaderboard
Top of the cross-benchmark aggregate ranking: every model placed on one Elo scale by an IRT model fit over public benchmark leaderboards (call about_the_aggregate for the current c
input · no output
get_model
One model in depth: aggregate rank, Elo with standard error, provider, what it is, cost per task where known, and its most notable benchmark results (with percentiles).
input · no output
get_prediction_duel
Guesswork — the public prediction duel: every day frontier LLMs and The Aggregate's own IRT model predict newly scraped benchmark scores before seeing them, and the errors are scor
input · no output
search_benchmarks
Find benchmarks in the aggregate by (partial) name. Returns model coverage, difficulty on the Elo scale, and the benchmark page URL.
input · no output
search_models
Find ranked models by (partial) name or provider. Returns rank, Elo and the model page URL. One row per model by default, fused across reasoning-effort settings.
input · no output
Verify it yourselfnpx teppi-check https://theaggregate.ai/mcpcurl -s https://api.teppi.xyz/v1/trust/mcp/mcs_01M1FZ235RQFJZEZ8C8XPV2H4H