MCP serverio.github.mauricekleine/nonobench
An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.
Overview
Score?
UNRATED 0.040
of what a free look can see, on 5 looks
Looks
6
last 8 hr ago
Tools
11
More info
URL
www.nonobench.com/mcp
streamable-http
Says it is
nonobench 1.1.0
protocol 2025-06-18
In the record since
4 days ago
Among servers18,413 with a card
0median 0.606 · this server 0.040 · highest on record 0.8561
Toolsfrom sha256:d105161310…c1661e
| Tool | Schema |
|---|---|
| check_solution Check a grid against a puzzle's clues, using the same rule as the benchmark grader. Reports which rows and columns do not match. |
input · output |
| compare_models Side-by-side core overall and per-size accuracy, cost, latency and token results for model or family names. |
input · output |
| get_leaderboard Models ranked by accuracy; defaults to all effort levels for compatibility. |
input · output |
| get_model_puzzles Which puzzles one model solved, missed, timed out on, or has not run. |
input · output |
| get_model_results Accuracy, cost, latency and token use for one model, broken down by grid size. |
input · output |
| get_puzzle One puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions. |
input · output |
| get_puzzle_results Per-model outcomes for one puzzle. Answers are omitted unless requested. |
input · output |
| list_families Model families, available efforts and best variants. |
input · output |
| list_providers Provider ids, names, families and variant counts. |
input · output |
| list_puzzles The benchmark puzzles with their ids and row/column clues. |
input · output |
| list_runs Individual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output. |
input · output |
Verify it yourself
npx teppi-check https://www.nonobench.com/mcpcurl -s https://api.teppi.xyz/v1/trust/mcp/mcs_01M3RGDVP8KP94ZDQ41EHJ8V5B