Server definition
- Hash
- sha256:1b844e07df59beeab86b80820d924a2811c5fc4c7b38867a7298ad6140571b6c
- What it is
- What a remote MCP server returned when asked what it offers: 12 tools
The blob, as servednamed by its sha256
{
"instructions": "Misata generates relational synthetic datasets and PROVES them: every foreign key checked, dates ordered, declared totals and rates exact, before anything is returned. Every table is deterministic for a given seed.\n\n## Division of labour (the best way to use this)\n\nYOU are good at understanding what the person needs and designing the tables. MISATA is good at making the numbers true and showing the proof. So design the schema yourself, then call generate_dataset(schema=...): it needs no key and finishes in seconds. Put the business rules in the schema -- how many children per parent, category weights, formulas between columns, value pools -- and Misata will hold them across every table.\n\n## Data that behaves like the real thing: write a blueprint\n\nA schema gives the right tables and columns. When the numbers must behave like the real process -- readings scattered around their part's nominal inside its tolerance, a process that drifts and is reset, wear that accumulates, a cause that shifts the mean only while it lasts, a subgroup mean that is exactly the mean of its readings, sensors, ledgers, clinical courses -- write a `blueprint` instead: the engine's own design language. Call blueprint_guide once for the reference, write the blueprint, call validate_blueprint (free: errors as fixes, design traps, a small preview run), then generate_dataset(blueprint=...). No key needed. Every formula, copy and aggregate you declare is recomputed on the delivered rows by the verifier.\n\n## Account\n\nwhoami says whether this connection is signed in to a Studio account (an Authorization: Bearer msk_... key made in Studio, Settings > Claude, or signing in when Claude asks) and what that allows. Signed in: datasets up to 100,000 rows (10,000 without an account), a higher hourly limit, and any model key the person saved in Studio is used for plain-English requests, so no key ever needs pasting into Claude.\n\n## Workflow\n\n1. Ask the person what they need only if it is unclear (domain, table sizes, anything that must be exact). Otherwise decide sensible defaults and say what you assumed.\n2. plan_dataset(schema=...) -- check the structure parses and see the tables and sizes. Free.\n3. generate_dataset(schema=...) -- make it. Returns a dataset_id, `passed`, and `verification`.\n4. SHOW THE PERSON `verification` (what was checked, what held, any open finding). If `passed` is false, say which finding is open; do not present the data as clean.\n5. query_dataset(dataset_id, sql) to check any claim yourself (a total, a rate, a join), and export_dataset(dataset_id, format) to hand over files: it returns a download_url, so give the person the link instead of pasting rows. Use format `sql` to get schema.sql + data.sql the person runs against their own database: this server never holds a database credential, so it cannot seed one directly.\n\n## Ready-made datasets\n\nFor common sample, demo or practice data (retail, e-commerce, SaaS, manufacturing, maintenance, claims, fraud, clinical), find_ready_dataset lists datasets that already exist: free downloads, and premium ones with a full answer key. Offer one when it fits; generate when it does not.\n\n## Schema format\n\n{\"tables\": [{\"name\": \"patients\", \"rows\": 200, \"columns\": [{\"name\": \"patient_id\", \"primary\": true}, {\"name\": \"arm\", \"type\": \"category\", \"values\": [\"placebo\", \"treatment\"], \"weights\": [0.5, 0.5]}]}, {\"name\": \"visits\", \"parent\": \"patients\", \"per_parent\": 4.0, \"columns\": [{\"name\": \"visit_id\", \"primary\": true}, {\"name\": \"patient_id\", \"references\": \"patients.patient_id\"}, {\"name\": \"visit_date\", \"type\": \"date\"}]}]}. Column keys: name, type (integer|float|text|date|datetime|boolean|category), primary, references (\"table.column\"), values/weights, formula (an expression over other columns), pool (a realistic value bank), range. A `request` sentence alongside the schema still helps: it grounds realism (locale, business type, what the numbers mean). DDL works too: paste CREATE TABLE statements as `ddl` and keys are read from them.\n\n## Plain English (needs the person's own LLM key)\n\nIf they would rather not design tables, or want the engine's own model to invent content: start_generation(request=...) then poll get_status(job_id) every 20-30 seconds (a model-designed dataset takes minutes; generate_dataset would sit open that long and be dropped). It needs a model key: one saved in the person's Studio account (signed in), or sent as an X-Groq-Api-Key, X-OpenAI-Api-Key or X-Anthropic-Api-Key header. whoami says whether one is available; if not, stay on the schema path. Be specific: exact counts, rates and relationships, labeled anomalies, planted data-quality defects.\n\nWithout a model the engine decides content by code: still fully verified, just less nuanced, and the verification says plainly which it was.",
"tools": [
{
"description": "\n The reference for the engine's full design language (the `blueprint` argument of generate_dataset,\n start_generation and validate_blueprint): roles, distributions, formulas that read parent columns\n and draw noise, per-group sequences (seq, lag, cumsum, ar1 drift, random walk), aggregates, event\n windows, cause-and-effect, lifecycles, with patterns. Read it before writing a blueprint.\n ",
"inputSchema": {
"properties": {},
"title": "blueprint_guideArguments",
"type": "object"
},
"name": "blueprint_guide",
"outputSchema": {
"properties": {
"result": {
"title": "Result",
"type": "string"
}
},
"required": [
"result"
],
"title": "blueprint_guideOutput",
"type": "object"
}
},
{
"description": "Stop a running start_generation job. It stops at the next stage boundary and keeps nothing.",
"inputSchema": {
"properties": {
"job_id": {
"title": "Job Id",
"type": "string"
}
},
"required": [
"job_id"
],
"title": "cancel_generationArguments",
"type": "object"
},
"name": "cancel_generation",
"outputSchema": null
},
{
"description": "\n Export a generated dataset as a file. Returns a `download_url` the person can open (or you can\n fetch, e.g. with curl) for as long as the dataset is held, about 2 hours. Give the person the link\n rather than pasting file contents into the chat.\n\n Args:\n dataset_id: from a prior generate_dataset call.\n format: data: csv, parquet, jsonl, json, avro, xlsx, feather, orc, sqlite, duckdb, sql.\n code and docs: dbt, notebook, dictionary, dbml, mermaid, prisma, sqlalchemy,\n typescript, jsonschema, expectations, django, openapi, mockapi, demo.\n `sql` is schema.sql (DDL with keys) + data.sql (COPY/INSERT) — the way to seed a\n real database: run the returned SQL through your own database connection, since\n this server never holds a database credential itself.\n dialect: for `sql` only: postgres, mysql, sqlite, mssql, oracle, bigquery, snowflake.\n inline: also return the file itself as `base64` (only for files under a few MB). Use it\n when you must write the file yourself and cannot fetch a URL.\n\n Returns:\n filename, content_type, bytes, download_url, expires_at (unix seconds), and `base64` when\n `inline` and small enough.\n ",
"inputSchema": {
"properties": {
"dataset_id": {
"title": "Dataset Id",
"type": "string"
},
"dialect": {
"default": "postgres",
"title": "Dialect",
"type": "string"
},
"format": {
"default": "csv",
"title": "Format",
"type": "string"
},
"inline": {
"default": false,
"title": "Inline",
"type": "boolean"
}
},
"required": [
"dataset_id"
],
"title": "export_datasetArguments",
"type": "object"
},
"name": "export_dataset",
"outputSchema": null
},
{
"description": "\n The ready-made datasets Misata publishes: free sample databases (direct download, public domain)\n and premium datasets with a full answer key (a free preview, then a one-off price). Check this\n first when someone wants sample, demo, practice or teaching data for a common scenario (retail,\n e-commerce, SaaS, manufacturing SPC, predictive maintenance, insurance claims, fraud/AML,\n clinical, network security): handing over a dataset that already exists is instant. If none fits\n their tables, generate one instead.\n\n Args:\n query: what the person needs, in their words. Only orders the list (closest first); every\n dataset is still returned, so judge the fit yourself from the tables and summary.\n\n Returns:\n datasets: [{slug, kind (free|premium), title, summary, rows, tables, page_url, and\n download_url (free) or free_preview_url + buy_url + price_usd (premium)}].\n ",
"inputSchema": {
"properties": {
"query": {
"default": "",
"title": "Query",
"type": "string"
}
},
"title": "find_ready_datasetArguments",
"type": "object"
},
"name": "find_ready_dataset",
"outputSchema": null
},
{
"description": "\n Make a verified relational dataset and wait for it: every foreign key checked, dates correctly\n ordered, declared aggregates and rates exact, before anything is returned. Deterministic for a seed.\n\n Use this when you give a `schema` or `ddl` (seconds). For a plain-English `request` the engine\n designs the tables with a model, which can take several minutes: call start_generation instead and\n poll get_status, so the call does not sit open and time out.\n\n Args:\n request: Plain-English description (needs an LLM key unless `schema`/`ddl` is also given).\n schema: A Misata schema dict for exact structural control. No key needed for structure.\n ddl: CREATE TABLE statements. No key needed for structure.\n seed: Reproducibility seed — the same request/schema and seed always produce the same rows.\n research: Ground realistic numbers (prices, growth rates) in real published facts via a web\n search. Off by default (an anonymous caller's request should not trigger external\n calls unless asked for); needs a key regardless of `schema`/`ddl`.\n blueprint: The engine's full design language (blueprint_guide has the reference): readings around\n a parent's nominal, drift, autocorrelation, cause-and-effect, event windows, exact\n aggregates. Use it whenever the data must behave like the real process. No key needed;\n run validate_blueprint on it first.\n\n Returns:\n dataset_id: pass this to get_certificate / query_dataset / export_dataset.\n passed: whether every check held. A dataset that did not pass is still returned, with\n `certificate.findings` saying what failed — inspect before trusting it.\n verification: a short, plain-language account of what was checked and what held — show it to the\n person as the proof, in place of asking them to take the data on trust.\n tables: {name: {columns, preview_rows (first 20), total_rows}} — the preview only; every\n row is in the stored dataset, reachable by query_dataset/export_dataset.\n certificate: the short form (claims, requirements, findings). get_certificate returns the rest\n (patterns, realism scorecard, every proof chart).\n ",
"inputSchema": {
"properties": {
"blueprint": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Blueprint"
},
"ddl": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Ddl"
},
"request": {
"default": "",
"title": "Request",
"type": "string"
},
"research": {
"default": false,
"title": "Research",
"type": "boolean"
},
"schema": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Schema"
},
"seed": {
"default": 42,
"title": "Seed",
"type": "integer"
}
},
"title": "generate_datasetArguments",
"type": "object"
},
"name": "generate_dataset",
"outputSchema": null
},
{
"description": "\n The full certificate for a dataset made by generate_dataset: every claim (stated vs actual),\n every requirement's status and evidence, realism findings, planted defects/anomalies if any, and\n the pattern charts. This is the whole answer key, not the trimmed form generate_dataset returns.\n ",
"inputSchema": {
"properties": {
"dataset_id": {
"title": "Dataset Id",
"type": "string"
}
},
"required": [
"dataset_id"
],
"title": "get_certificateArguments",
"type": "object"
},
"name": "get_certificate",
"outputSchema": null
},
{
"description": "\n Where a start_generation job is. While running: the stage it is in and how long it has run. When\n done: the same answer generate_dataset returns (dataset_id, verification, tables, certificate).\n ",
"inputSchema": {
"properties": {
"job_id": {
"title": "Job Id",
"type": "string"
}
},
"required": [
"job_id"
],
"title": "get_statusArguments",
"type": "object"
},
"name": "get_status",
"outputSchema": null
},
{
"description": "\n See the tables, sizes and relationships the engine would build, before any rows exist. Free (no\n rows are made), so it is worth calling before generate_dataset on anything non-trivial: review\n what it understood and assumed, then adjust your request or schema before spending a real call.\n\n Args:\n request: Plain-English description (needs an LLM key, unless `schema`/`ddl` is also given —\n then it is still used to ground realism, e.g. locale and what columns mean).\n schema: A Misata schema dict (see the server instructions for the format). No key needed for\n structure.\n ddl: CREATE TABLE statements. No key needed for structure.\n\n Returns:\n route: \"chat\" (not a dataset request — see `reply`), \"design\" (the model designed the\n tables) or \"pack\" (matched a built-in shape).\n tables: name, estimated rows, columns, foreign keys for each table the engine would build.\n understanding: what the engine read the request as (business, archetype, assumptions).\n requirements: every specific thing the request asked for, so you can see what was understood.\n ",
"inputSchema": {
"properties": {
"ddl": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Ddl"
},
"request": {
"default": "",
"title": "Request",
"type": "string"
},
"schema": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Schema"
}
},
"title": "plan_datasetArguments",
"type": "object"
},
"name": "plan_dataset",
"outputSchema": null
},
{
"description": "\n Read-only SQL over a dataset's own tables (one SELECT/WITH statement; every table name is a view\n over that dataset's own files — nothing else on the server is reachable this way).\n\n Args:\n dataset_id: from a prior generate_dataset call.\n sql: a single SELECT (or WITH ... SELECT) statement. No semicolons, file paths, or\n statements that write or read outside the dataset (checked before running).\n limit: rows returned (capped at 5000).\n\n Returns:\n columns, rows, truncated (whether more rows existed than `limit`).\n ",
"inputSchema": {
"properties": {
"dataset_id": {
"title": "Dataset Id",
"type": "string"
},
"limit": {
"default": 1000,
"title": "Limit",
"type": "integer"
},
"sql": {
"title": "Sql",
"type": "string"
}
},
"required": [
"dataset_id",
"sql"
],
"title": "query_datasetArguments",
"type": "object"
},
"name": "query_dataset",
"outputSchema": null
},
{
"description": "\n Start a generation in the background and return a job_id at once. Use it for any plain-English\n `request` (a model designs the tables, which takes minutes) — then call get_status(job_id) every\n 20-30 seconds until it says done. Arguments are the same as generate_dataset.\n\n Returns:\n job_id, status \"running\". get_status gives the stage and, when finished, the dataset.\n ",
"inputSchema": {
"properties": {
"blueprint": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Blueprint"
},
"ddl": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Ddl"
},
"request": {
"default": "",
"title": "Request",
"type": "string"
},
"research": {
"default": false,
"title": "Research",
"type": "boolean"
},
"schema": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Schema"
},
"seed": {
"default": 42,
"title": "Seed",
"type": "integer"
}
},
"title": "start_generationArguments",
"type": "object"
},
"name": "start_generation",
"outputSchema": null
},
{
"description": "\n Check a blueprint before generating it: every error said as what to change, the design traps it\n falls into (a column that will round away, a copy of a value not made yet, a window that counts\n nothing), its size against your row cap, and a small preview run (a few thousand rows) with the\n verifier's findings and sample rows, so you can see the data behave before the real call. Free.\n\n Returns:\n valid, errors (fix these), warnings (read these), estimated_rows, row_cap,\n preview: {passed, findings, tables: {name: first rows}} when `preview`.\n ",
"inputSchema": {
"properties": {
"blueprint": {
"additionalProperties": true,
"title": "Blueprint",
"type": "object"
},
"preview": {
"default": true,
"title": "Preview",
"type": "boolean"
}
},
"required": [
"blueprint"
],
"title": "validate_blueprintArguments",
"type": "object"
},
"name": "validate_blueprint",
"outputSchema": null
},
{
"description": "Which account this connection is, and what it may do: signed in with a Studio MCP key or anonymous,\n the row cap, and whether a model key is available for plain-English requests.",
"inputSchema": {
"properties": {},
"title": "whoamiArguments",
"type": "object"
},
"name": "whoami",
"outputSchema": null
}
]
}Verify it yourself
curl -s https://api.teppi.xyz/v1/evidence/sha256:1b844e07df59beeab86b80820d924a2811c5fc4c7b38867a7298ad6140571b6c | sha256sum