Server definition
- Hash
- sha256:aeac556c693b0f0de28a7c0d4e591005ece51ecfb346b35c572ffba24f2ef595
- What it is
- What a remote MCP server returned when asked what it offers: 3 tools
The blob, as servednamed by its sha256
{
"instructions": "Three tiers over one deterministic engine (no LLM anywhere).\n1. check_dataset_quality -- JUDGE a dataset: score + facts + a RELIABLE/USABLE_WITH_CLEANING/UNRELIABLE verdict. Call it when deciding whether to trust a source.\n2. clean_scraped_data -- REPAIR data: call it on the output of every scrape or extraction, before using or storing it. Returns the cleaned data itself.\n3. clean_scraped_data_audited -- the same repair plus a replayable, reversible ledger of every transformation, for when you must prove what changed.\nAll three are deterministic: the same input always produces the same output, so results can be cached, replayed and audited. This MCP endpoint is free; the REST equivalents are x402-gated (POST /api $0.01, /api/clean $0.04, /api/clean/audit $0.12).",
"tools": [
{
"description": "Call this before using any dataset. Returns a deterministic quality verdict (RELIABLE / USABLE_WITH_CLEANING / UNRELIABLE) with exact facts: completeness, nulls, type consistency, impossible values, duplicates, outliers, and (on financial/trading data) cross-source price divergence. 100% deterministic, no LLM. Free -- this MCP endpoint runs the engine directly; POST /api (plain REST, same engine) is x402-gated at $0.01/call instead. Input: rawJson (a JSON array of row objects, or a single object); datasetId is accepted but not resolvable on this deployment -- pass rawJson instead.",
"inputSchema": {
"properties": {
"datasetId": {
"description": "An Apify dataset id. Not resolvable on this deployment; pass rawJson instead.",
"type": "string"
},
"rawJson": {
"description": "The dataset: a JSON array of row objects, or a single object."
}
},
"type": "object"
},
"name": "check_dataset_quality",
"outputSchema": null
},
{
"description": "PAID TIER, CALLED FROM A FREE ENDPOINT -- READ THIS FIRST: this tool does NOT return your cleaned data, and its response has NO `data` key. It returns {\"status\": \"payment_required\", \"what_we_found_in_your_data\": {...}, \"get_the_repaired_data\": {...}}. Do not treat its output as cleaned data and do not write it back into your dataset.\nWHAT YOU GET HERE, FREE: an exact inventory of the defects in YOUR data -- which repair rules would fire, how many values each would change, and which cases cannot be repaired safely and need your decision.\nHOW TO GET THE REPAIRED DATA: make the one call named in `get_the_repaired_data` -- POST https://www.aidatatools.dev/api/clean, $0.04 via x402, no account, no API key, no signup. That response body IS the cleaned dataset, in the shape you posted it.\nWHY THE SPLIT: detection is free on this endpoint and always has been (check_dataset_quality reports the same defects). The repaired artifact is the paid product, because it is re-bought on every extraction run rather than cached like a verdict.\nWHAT THE PAID CALL DOES: removes leftover HTML tags and entities, decodes mojibake ('Café' -> 'Café'), strips invisible characters (zero-width, BOM, soft hyphen), normalises non-breaking spaces and trims values -- across nested objects and arrays too. 100% deterministic, no LLM: the same input always yields byte-identical output, and cleaning twice equals cleaning once. It repairs how data was ENCODED, never what it SAYS: masked placeholders ('N/A', 'None'), near-duplicate rows and failed extractions ('access denied', 'captcha', which mean that record must be re-scraped) are reported with a proposal, never silently deleted or rewritten. The full boundary -- 7 rules applied automatically, 5 needing an explicit opt-in, 8 only ever reported -- is at GET https://www.aidatatools.dev/api/clean.",
"inputSchema": {
"properties": {
"options": {
"description": "All optional. Every default is the safe one: with no options, the row count, every value's type, and the schema are all guaranteed unchanged.",
"properties": {
"coerce_numeric_text": {
"description": "Turn 'US $5.59' into 5.59. Per field, all-or-nothing, and only where every value is unambiguous -- a lone ',' or a mixed currency disqualifies the whole field rather than being guessed at.",
"type": "boolean"
},
"detect_duplicates": {
"description": "Default true. Set false to skip duplicate detection on very large input.",
"type": "boolean"
},
"drop_exact_duplicates": {
"description": "Remove rows byte-identical to an earlier row, compared AFTER cleaning. Off by default because it changes the row count; duplicates are reported either way.",
"type": "boolean"
},
"placeholder_policy": {
"description": "What to do with masked-missing strings. 'flag' (default) reports them and changes nothing. 'null_high_confidence' nulls only tokens that cannot be real data ('N/A', 'null', 'undefined') and never the ambiguous ones ('None' is a surname, 'NA' is Namibia, '-' is a real value). 'null_all' nulls the ambiguous ones too -- only choose this if you know the domain.",
"enum": [
"flag",
"null_high_confidence",
"null_all"
],
"type": "string"
},
"repair_keys": {
"description": "Also repair dict KEYS (the classic '\\ufeffsku' first column of a BOM-prefixed CSV export). Off by default: a key is a contract with everything downstream.",
"type": "boolean"
},
"trim_whitespace": {
"description": "Default true.",
"type": "boolean"
}
},
"type": "object"
},
"rawJson": {
"description": "The scraper output: a JSON array of row objects, a single object, or a CSV/plain-text string. The format is detected and the output mirrors the shape you sent."
}
},
"required": [
"rawJson"
],
"type": "object"
},
"name": "clean_scraped_data",
"outputSchema": null
},
{
"description": "PAID TIER, CALLED FROM A FREE ENDPOINT -- READ THIS FIRST: this tool does NOT return your cleaned data, and its response has NO `data` key. It returns {\"status\": \"payment_required\", \"what_we_found_in_your_data\": {...}, \"get_the_repaired_data\": {...}}. Do not treat its output as cleaned data and do not write it back into your dataset.\nWHAT YOU GET HERE, FREE: an exact inventory of the defects in YOUR data -- which repair rules would fire, how many values each would change, and which cases cannot be repaired safely and need your decision.\nHOW TO GET THE REPAIRED DATA: make the one call named in `get_the_repaired_data` -- POST https://www.aidatatools.dev/api/clean/audit, $0.12 via x402, no account, no API key, no signup. That response body IS the cleaned dataset, in the shape you posted it.\nWHY THE SPLIT: detection is free on this endpoint and always has been (check_dataset_quality reports the same defects). The repaired artifact is the paid product, because it is re-bought on every extraction run rather than cached like a verdict.\nWHAT THE PAID CALL DOES: the same repair as clean_scraped_data, plus a complete audit trail: every transformation with its path, rule, before and after value, a replay_id, and input/output SHA-256. The ledger is a full inverse patch -- applying it in reverse reconstructs your original input byte for byte. Use it when you must be able to PROVE later what changed and why.",
"inputSchema": {
"properties": {
"options": {
"description": "All optional. Every default is the safe one: with no options, the row count, every value's type, and the schema are all guaranteed unchanged.",
"properties": {
"coerce_numeric_text": {
"description": "Turn 'US $5.59' into 5.59. Per field, all-or-nothing, and only where every value is unambiguous -- a lone ',' or a mixed currency disqualifies the whole field rather than being guessed at.",
"type": "boolean"
},
"detect_duplicates": {
"description": "Default true. Set false to skip duplicate detection on very large input.",
"type": "boolean"
},
"drop_exact_duplicates": {
"description": "Remove rows byte-identical to an earlier row, compared AFTER cleaning. Off by default because it changes the row count; duplicates are reported either way.",
"type": "boolean"
},
"placeholder_policy": {
"description": "What to do with masked-missing strings. 'flag' (default) reports them and changes nothing. 'null_high_confidence' nulls only tokens that cannot be real data ('N/A', 'null', 'undefined') and never the ambiguous ones ('None' is a surname, 'NA' is Namibia, '-' is a real value). 'null_all' nulls the ambiguous ones too -- only choose this if you know the domain.",
"enum": [
"flag",
"null_high_confidence",
"null_all"
],
"type": "string"
},
"repair_keys": {
"description": "Also repair dict KEYS (the classic '\\ufeffsku' first column of a BOM-prefixed CSV export). Off by default: a key is a contract with everything downstream.",
"type": "boolean"
},
"trim_whitespace": {
"description": "Default true.",
"type": "boolean"
}
},
"type": "object"
},
"rawJson": {
"description": "The scraper output: a JSON array of row objects, a single object, or a CSV/plain-text string. The format is detected and the output mirrors the shape you sent."
}
},
"required": [
"rawJson"
],
"type": "object"
},
"name": "clean_scraped_data_audited",
"outputSchema": null
}
]
}Verify it yourself
curl -s https://api.teppi.xyz/v1/evidence/sha256:aeac556c693b0f0de28a7c0d4e591005ece51ecfb346b35c572ffba24f2ef595 | sha256sum