MCP servercom.completionkit/evals
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Overview
Score?
UNRATED 0.681
of what a free look can see, on 32 looks
Looks
36
last 5 hr ago
Tools
54
More info
URL
completionkit.com/mcp
streamable-http
Says it is
CompletionKit 0.28.43
protocol 2025-03-26
In the record since
32 days ago
Among servers18,413 with a card
0median 0.606 · this server 0.681 · highest on record 0.8561
Toolsfrom sha256:bee1dcc64c…a09600
| Tool | Schema |
|---|---|
| agreements_create Upsert an agreement for (run, response, metric, created_by). Verdict is one of agree, disagree, borderline. corrected_score (1..5) is required when verdict is 'disagree'. |
input · no output |
| agreements_list List agreements. Filter by run_id, response_id, metric_id, or created_by. |
input · no output |
| datasets_create Create a dataset with CSV data. First row is the header. Two column names are recognized specially: "expected_output" is each row's answer key (ground truth) given to the judge and |
input · no output |
| datasets_create_from_url Create a dataset by downloading CSV from a URL instead of inlining it. Use this for large datasets: pass a public http(s) URL and the server fetches the CSV directly, so the data n |
input · no output |
| datasets_delete Delete a dataset |
input · no output |
| datasets_get Get a dataset by ID |
input · no output |
| datasets_list List all datasets |
input · no output |
| datasets_update Update a dataset |
input · no output |
| judges_compare Compare two versions of one metric's agreement stats side by side. Requires metric_id, metric_version_a_id, and metric_version_b_id (both versions must belong to that metric). Unav |
input · no output |
| judges_replay Create a scoring run for the current judge over a dataset's existing outputs (wraps runs_create with prompt_id omitted and output_column supplied). This only sets up the run; call |
input · no output |
| metric_groups_create Create a metric group |
input · no output |
| metric_groups_delete Delete a metric group |
input · no output |
| metric_groups_get Get a metric group by ID |
input · no output |
| metric_groups_list List all metric groups |
input · no output |
| metric_groups_update Update a metric group |
input · no output |
| metric_versions_dismiss Destroy a draft MetricVersion (use for either source: 'edit' or source: 'suggestion'). Published versions are refused — to demote a published version, publish a different one as cu |
input · no output |
| metric_versions_list List every MetricVersion (drafts + published) for a metric, newest first. Each row carries version_number, state, source, current flag, and timestamps. |
input · no output |
| metric_versions_publish Publish a MetricVersion as the live version of its metric. Works for both 'draft → published' and 'revert to an older published version → current'. Transactionally flips current, d |
input · no output |
| metrics_create Create a metric with evaluation criteria. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value (contains/not_contains/equals), pattern |
input · no output |
| metrics_delete Delete a metric |
input · no output |
| metrics_get Get a metric by ID |
input · no output |
| metrics_list List all metrics |
input · no output |
| metrics_suggest_variants Ask the model to rewrite the metric's judge instruction in N variants targeted at the recent disagreements. Each variant is saved as a draft MetricVersion with source="suggestion". |
input · no output |
| metrics_update Update a metric. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value (contains/not_contains/equals), pattern (regex), json_path+expect |
input · no output |
| promptfoo_import Import a promptfooconfig.yaml. Creates a prompt, a dataset from the test vars, and metrics from the assert blocks (llm-rubric/g-eval become judge metrics; contains/equals/regex/is- |
input · no output |
| prompts_create Create a prompt |
input · no output |
| prompts_delete Delete a prompt |
input · no output |
| prompts_get Get a prompt by ID |
input · no output |
| prompts_list List all prompts |
input · no output |
| prompts_publish Publish a prompt version, making it the current version |
input · no output |
| prompts_suggest_improvement Suggest an improved version of a prompt, grounded in a run's test results and judge feedback. Analyzes the run's responses, scores, and reviews, then returns reasoning plus a rewri |
input · no output |
| prompts_update Update a prompt. If the prompt already has runs, this creates a new DRAFT version (current=false) rather than editing in place or publishing — promote it with prompts_publish — so |
input · no output |
| provider_credentials_create Create a provider credential |
input · no output |
| provider_credentials_delete Delete a provider credential |
input · no output |
| provider_credentials_get Get a provider credential by ID (API key is not exposed) |
input · no output |
| provider_credentials_list List all provider credentials (API keys are not exposed) |
input · no output |
| provider_credentials_update Update a provider credential |
input · no output |
| responses_get Get a specific response |
input · no output |
| responses_list List responses for a run, in row order. Returns {total, limit, offset, returned, responses}. Defaults to 50 rows because full payloads are large: use "fields" to drop the bodies, " |
input · no output |
| runs_create Create a run. Omit prompt_id and provide output_column to score existing outputs by grading a pre-existing dataset column instead of generating new ones. |
input · no output |
| runs_delete Delete a run |
input · no output |
| runs_generate Start a run. Required for every run, including score-only runs (no prompt): generates responses with the prompt when there is one, otherwise copies the graded dataset column and gr |
input · no output |
| runs_get Get a run by ID, including "metric_averages": a per-metric breakdown with each metric's average score (or pass rate for checks), how many rows it graded, and how many scored low. U |
input · no output |
| runs_list List all runs |
input · no output |
| runs_regrade Re-grade a run's existing responses with its currently attached metrics, without regenerating. Use after attaching or editing metrics on an already-generated run. |
input · no output |
| runs_rerun Create and start a fresh copy of a run with the same prompt, dataset, metrics, and settings. Use when the judge changed and you want a clean run instead of mixing versions. |
input · no output |
| runs_retry_failures Re-run only the failed responses of a run, optionally limited to specific response ids via "only". |
input · no output |
| runs_update Update a run |
input · no output |
| tags_create Create a tag. Color is auto-assigned. |
input · no output |
| tags_delete Delete a tag. Removes the tag from every linked metric, prompt, run, and dataset. |
input · no output |
| tags_get Get a tag by ID |
input · no output |
| tags_list List all tags |
input · no output |
| tags_update Rename a tag. |
input · no output |
| usage_get Get this organization's plan usage and limits for the current billing period: runs and prompt fetches used, their limits, how many remain, and when the period resets. Call this to |
input · no output |
Verify it yourself
npx teppi-check https://completionkit.com/mcpcurl -s https://api.teppi.xyz/v1/trust/mcp/mcs_01M1FZ25QFSF1VQATC6G22CBSY