Server definition
- Hash
- sha256:55e2002af030b029538fb993bbc6f285f16bf91f46da49dfc46eed7f15ccb273
- What it is
- What a remote MCP server returned when asked what it offers: 3 tools
The blob, as servednamed by its sha256
{
"instructions": "The Mozilla Data Collective is a catalog of ethically sourced AI training datasets — speech, text and other modalities — contributed by community organizations, each setting its own license and access terms.\n\nThis server is a read-only discovery surface over the public catalog. Use it to find datasets and describe what exists. It cannot download data, read private datasets, or make purchases.\n\nRecommended workflow:\n\n1. Call `search` with the user's need phrased as natural language. The backend runs a hybrid lexical and semantic pipeline, so a descriptive phrase (\"conversational Swahili audio for ASR fine-tuning\") retrieves better than a bare keyword.\n2. To narrow results by task, language, license, format, price or sample availability, call `list_filters` first and use the exact values it returns. Filter values are matched exactly and are case-sensitive, so never guess or invent them.\n3. Call `fetch` with an id from the search results to get the full description, license, size and pricing for one dataset.\n4. Give the user the dataset's `url`. Accepting the contributor's terms, payment and downloading all happen on that page. For programmatic access, point them to the REST API and Python SDK documented at /api-reference.\n\nGround every answer in what the tools returned: only name datasets that appeared in a result, and quote license and pricing from `fetch` rather than inferring them. If a search returns nothing, say so plainly — the catalog does not cover every language or task, and a report of \"no matches\" is more useful than a near miss presented as a match.",
"tools": [
{
"description": "Fetch the full public details of one Mozilla Data Collective dataset by id or slug: description, organization, task, locale, license, format, size, pricing, and its page URL.",
"inputSchema": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"properties": {
"id": {
"description": "Dataset id or slug, as returned in the id field of search results.",
"maxLength": 500,
"minLength": 1,
"type": "string"
}
},
"required": [
"id"
],
"type": "object"
},
"name": "fetch",
"outputSchema": null
},
{
"description": "List every value the search tool's filters accept: the tasks, locales, licenses and formats present in the catalog, plus the sort and date-range options. Task and license values are abbreviations, so taskLabels and licenseLabels spell them out. Filter values are matched exactly, so call this before filtering a search rather than guessing values. Takes no arguments.",
"inputSchema": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"properties": {},
"type": "object"
},
"name": "list_filters",
"outputSchema": null
},
{
"description": "Search the Mozilla Data Collective catalog of AI training datasets by natural-language query, optionally narrowed by task, language, license, format, price, sample availability or publish date. Returns matching datasets as {id, title, url}; pass an id to the fetch tool for full details. Call list_filters first if you intend to filter — filter values must match the catalog exactly.",
"inputSchema": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"properties": {
"format": {
"description": "Restrict to these file formats, e.g. ['WAV', 'MP3']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
"items": {
"minLength": 1,
"type": "string"
},
"maxItems": 20,
"minItems": 1,
"type": "array"
},
"hasSample": {
"description": "true returns only datasets that publish a downloadable sample, useful when the user wants to try data before committing. false behaves the same as omitting it.",
"type": "boolean"
},
"isPaid": {
"description": "true returns only paid datasets, false only free ones. Omit to include both.",
"type": "boolean"
},
"license": {
"description": "Restrict to these license abbreviations, e.g. ['CC0-1.0', 'CC-BY-4.0']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
"items": {
"minLength": 1,
"type": "string"
},
"maxItems": 20,
"minItems": 1,
"type": "array"
},
"limit": {
"default": 10,
"description": "Maximum number of results to return (1-25).",
"maximum": 25,
"minimum": 1,
"type": "integer"
},
"locale": {
"description": "Restrict to these language/locale codes, e.g. ['sw', 'pt-BR']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
"items": {
"minLength": 1,
"type": "string"
},
"maxItems": 20,
"minItems": 1,
"type": "array"
},
"query": {
"description": "Natural-language search query describing the datasets you are looking for, e.g. 'Spanish speech recordings for TTS training'. Descriptive phrases retrieve better than single keywords.",
"maxLength": 500,
"minLength": 1,
"type": "string"
},
"sort": {
"description": "Result ordering. Defaults to 'relevance'; use 'newest' or 'size' only when the user asks for it.",
"enum": [
"relevance",
"newest",
"size"
],
"type": "string"
},
"sortDirection": {
"description": "Direction for the sort field. Only meaningful alongside sort='newest' or sort='size'.",
"enum": [
"asc",
"desc"
],
"type": "string"
},
"task": {
"description": "Restrict to these machine-learning tasks, e.g. ['ASR', 'TTS'].",
"items": {
"enum": [
"N/A",
"NLP",
"ASR",
"LID",
"TTS",
"MT",
"LM",
"LLM",
"NLU",
"NLG",
"CALL",
"RAG",
"CV",
"ML",
"OTH"
],
"type": "string"
},
"maxItems": 20,
"minItems": 1,
"type": "array"
},
"uploadDate": {
"description": "Restrict to datasets published within this recent window.",
"enum": [
"today",
"thisWeek",
"thisMonth",
"thisYear"
],
"type": "string"
}
},
"required": [
"query"
],
"type": "object"
},
"name": "search",
"outputSchema": null
}
]
}Verify it yourself
curl -s https://api.teppi.xyz/v1/evidence/sha256:55e2002af030b029538fb993bbc6f285f16bf91f46da49dfc46eed7f15ccb273 | sha256sum