Server definition
- Hash
- sha256:358189d226bdce18c7a97fc93d75369b07eaae6fdfc2e183fc5419fd7ed3768f
- What it is
- What a remote MCP server returned when asked what it offers: 27 tools
The blob, as servednamed by its sha256
{
"instructions": null,
"tools": [
{
"description": "Audit whether AI assistants (ChatGPT, Claude, Perplexity, Google AI Overview, Bing Copilot) can read and cite a page, and optionally ask them. On-page pass (always): the live robots.txt resolved for 24 AI crawlers per RFC 9309 with the deciding line, Content-Signal, a fetch that identifies as GPTBot to catch WAFs filtering on user-agent, noindex/nosnippet/noai/data-nosnippet, text present without JavaScript, JSON-LD types and resolvable Organization/Person entities, heading outline, question-shaped headings, answer-first paragraph, lists/tables, numeric facts and quotes, chunk-sized sections, dateModified with age, author, outbound sources. Also readability grade, paragraph length, definitional openers, named-entity density, keyword stuffing, first-hand content, images/video, paywall and retired robots tokens. Retrievability first: where Google ranks the page for its own H1 question and whether it is indexed (2 SERPs) — a page that is not retrievable is not cited whatever its on-page score. Google AI Overview and Bing Copilot report brand MENTIONS only: their no-JS SERP exposes no sources. Returns a 0-100 score per pillar (retrievability, access, readability, structure, answerability, trust, plus offsite when requested), blockers that cap the score, every check with evidence and fix, and topFixes. Citation panel (when `queries` is set): asks each engine, reports cited / mentioned / rank per (query × engine), share of voice across all cited domains, and the domains winning the questions where the page is absent. Use this instead of seo_audit when the question is AI answers rather than Google rankings.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"brand": {
"description": "Brand name to look for in the answer text ('mentioned' even when not cited). Defaults to the page's og:site_name / Organization name.",
"maxLength": 80,
"type": "string"
},
"competitors": {
"description": "Competitor domains to flag in the share of voice, e.g. ['brightdata.com']",
"items": {
"maxLength": 253,
"minLength": 1,
"type": "string"
},
"maxItems": 20,
"type": "array"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"country": {
"description": "ISO country code for the proxy exit, e.g. 'us' — also the locale of the AI Overview / Copilot SERP",
"maxLength": 2,
"minLength": 2,
"type": "string"
},
"engines": {
"description": "Engines to ask (default: all). aio = Google AI Overview read from a live SERP, copilot = Bing's generative answer, openai/anthropic = the vendors' APIs with web search (an approximation of ChatGPT/Claude search), deepseek = our own Google top-10 handed to DeepSeek to answer and cite (cheapest; measures whether a model picks your page from the same results).",
"items": {
"enum": [
"perplexity",
"openai",
"anthropic",
"aio",
"copilot",
"deepseek"
],
"type": "string"
},
"type": "array"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"no_bot_fetch": {
"description": "Skip the extra request that identifies itself as GPTBot",
"type": "boolean"
},
"no_render": {
"description": "Skip the rendered pass (cheaper — the two JS-parity checks are reported as skipped)",
"type": "boolean"
},
"no_retrieval": {
"description": "Skip the retrievability probe (2 SERPs: Google rank of the page for its own H1 question, and whether it is indexed). On by default — it is the strongest single predictor of citation and a blocker when the page is not indexed.",
"type": "boolean"
},
"offsite": {
"description": "Also measure the brand OFF the page with five searches (\"brand\" site:youtube.com / reddit.com / wikipedia.org / linkedin.com / review sites) — the signals studies rank above anything on-page for whether a brand gets named. Adds an `offsite` pillar; billed as 5 SERP calls.",
"type": "boolean"
},
"queries": {
"description": "Questions to ask the AI engines (max 10). Omit for the on-page audit only — each (query × engine) pair is a billed engine call.",
"items": {
"maxLength": 300,
"minLength": 1,
"type": "string"
},
"maxItems": 10,
"type": "array"
},
"url": {
"description": "The page URL to audit",
"format": "uri",
"type": "string"
}
},
"required": [
"url",
"context",
"llm_model"
],
"type": "object"
},
"name": "ai_visibility",
"outputSchema": null
},
{
"description": "Scrape many URLs asynchronously with shared options. Returns a job id — poll with batch_status. For SEO/status audits over many pages set mode 'summary': items carry metadata only (title, description, canonical, contentLength) instead of full page content.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"content_mode": {
"description": "Per-URL content scope: smart (default) | article | full",
"enum": [
"smart",
"article",
"full"
],
"type": "string"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"country": {
"description": "ISO country code for the proxy exit",
"maxLength": 2,
"minLength": 2,
"type": "string"
},
"engine": {
"description": "Fetch engine (default auto)",
"enum": [
"auto",
"tls",
"fetch",
"render"
],
"type": "string"
},
"format": {
"description": "Output format (default markdown)",
"enum": [
"markdown",
"html",
"text"
],
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"mode": {
"description": "summary: per-URL metadata only, no page content — the light mode for audits",
"enum": [
"summary"
],
"type": "string"
},
"urls": {
"description": "URLs to scrape",
"items": {
"format": "uri",
"type": "string"
},
"maxItems": 5000,
"minItems": 1,
"type": "array"
}
},
"required": [
"urls",
"context",
"llm_model"
],
"type": "object"
},
"name": "batch",
"outputSchema": null
},
{
"description": "Poll a batch job for progress and per-URL results. Polls are incremental: pass the previous response's `nextCursor` as `since` to receive only the items completed after your last poll. Items omit page content by default — set include_content true only when you actually need the text.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"include_content": {
"description": "Include each item's full page content (default false — metadata only)",
"type": "boolean"
},
"jobId": {
"description": "The batch job id returned by batch",
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"since": {
"description": "Item cursor from the previous poll's `nextCursor` — returns only newer items",
"minimum": 0,
"type": "integer"
}
},
"required": [
"jobId",
"context",
"llm_model"
],
"type": "object"
},
"name": "batch_status",
"outputSchema": null
},
{
"description": "Fetch a Collector run by run_id: status (queued|running|done|failed), result count, cost, partial flag and the result rows. Use after run_collector returned 202/async. Pass format 'csv' to get the rows as CSV text.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"format": {
"description": "Return rows as JSON (default) or CSV text",
"enum": [
"json",
"csv"
],
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"run_id": {
"description": "The run id returned by run_collector",
"minLength": 1,
"type": "string"
}
},
"required": [
"run_id",
"context",
"llm_model"
],
"type": "object"
},
"name": "collector_run_status",
"outputSchema": null
},
{
"description": "Start an asynchronous BFS crawl of a site from a seed URL, converting each page to Markdown. Returns a job id — poll with crawl_status.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"content_mode": {
"description": "Per-page content scope: smart (default) | article | full",
"enum": [
"smart",
"article",
"full"
],
"type": "string"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"country": {
"description": "ISO country code for the proxy exit",
"maxLength": 2,
"minLength": 2,
"type": "string"
},
"depth": {
"description": "Max link depth (default 3)",
"maximum": 10,
"minimum": 0,
"type": "integer"
},
"exclude": {
"description": "URL substrings/globs to exclude",
"items": {
"type": "string"
},
"type": "array"
},
"include": {
"description": "URL substrings/globs to include",
"items": {
"type": "string"
},
"type": "array"
},
"limit": {
"description": "Max pages (default 50)",
"maximum": 500,
"minimum": 1,
"type": "integer"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"url": {
"description": "Seed URL",
"format": "uri",
"type": "string"
}
},
"required": [
"url",
"context",
"llm_model"
],
"type": "object"
},
"name": "crawl",
"outputSchema": null
},
{
"description": "Poll a crawl job for progress and the pages crawled so far. Polls are incremental: pass the previous response's `nextCursor` as `since` to receive only the pages crawled since your last poll. Pages omit their content by default — set include_content true only when you actually need the text (a large crawl's full content can be hundreds of KB).",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"include_content": {
"description": "Include each page's full content (default false — metadata only)",
"type": "boolean"
},
"jobId": {
"description": "The crawl job id returned by crawl",
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"since": {
"description": "Page cursor from the previous poll's `nextCursor` — returns only newer pages",
"minimum": 0,
"type": "integer"
}
},
"required": [
"jobId",
"context",
"llm_model"
],
"type": "object"
},
"name": "crawl_status",
"outputSchema": null
},
{
"description": "Build a structured dataset from a plain-language prompt. Quantic AI plans the search queries, searches Google/Bing/DuckDuckGo, maps the sites it finds and scrapes them into validated rows (CSV/JSON). Returns a job id — poll with dataset_status. Billed per delivered, validated record (email/phone fields cost extra, only when found); the run never exceeds limits.max_cost_usd, and the unspent budget is refunded.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"columns": {
"description": "Columns to extract; omit to let the planner infer them",
"items": {
"additionalProperties": false,
"properties": {
"description": {
"type": "string"
},
"name": {
"type": "string"
},
"type": {
"description": "email/phone/deep are premium fields, billed only when found",
"enum": [
"string",
"number",
"email",
"phone",
"url",
"boolean",
"deep"
],
"type": "string"
}
},
"required": [
"name"
],
"type": "object"
},
"type": "array"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"country": {
"description": "ISO country code for the proxy exit geo",
"maxLength": 2,
"minLength": 2,
"type": "string"
},
"limits": {
"additionalProperties": false,
"properties": {
"max_cost_usd": {
"description": "Budget cap for the run (default 5)",
"minimum": 0.05,
"type": "number"
},
"max_pages": {
"minimum": 1,
"type": "integer"
},
"max_rows": {
"minimum": 1,
"type": "integer"
}
},
"type": "object"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"prompt": {
"description": "What dataset you want, in plain language (e.g. 'coffee roasters in Portland with email and phone')",
"type": "string"
},
"sources": {
"additionalProperties": false,
"description": "Domain allow/deny lists",
"properties": {
"exclude": {
"items": {
"type": "string"
},
"type": "array"
},
"include": {
"items": {
"type": "string"
},
"type": "array"
}
},
"type": "object"
},
"webhook": {
"description": "Public URL to POST the finished dataset to",
"format": "uri",
"type": "string"
}
},
"required": [
"prompt",
"context",
"llm_model"
],
"type": "object"
},
"name": "create_dataset",
"outputSchema": null
},
{
"description": "Poll a dataset job for progress, the collection trace (steps) and the rows so far. Polls are incremental: pass the previous response's `nextCursor` as `since` to receive only rows delivered after your last poll. Set mode 'summary' to omit rows and get only progress + steps (light poll). When status is completed, the response includes signed CSV/JSON download URLs.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"jobId": {
"description": "The dataset job id returned by create_dataset",
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"mode": {
"description": "summary: progress + steps only, no rows",
"enum": [
"summary"
],
"type": "string"
},
"since": {
"description": "Row cursor from the previous poll's `nextCursor` — returns only newer rows",
"minimum": 0,
"type": "integer"
}
},
"required": [
"jobId",
"context",
"llm_model"
],
"type": "object"
},
"name": "dataset_status",
"outputSchema": null
},
{
"description": "Look at a page ONCE with an LLM and get back CSS selectors that extract the fields you asked for. Pass the returned `parser` as the `extract` argument on every later scrape of that same layout and no AI runs again — it becomes a plain, free, deterministic extraction. Use this instead of ai_prompt whenever you will scrape more than a couple of pages of the same shape. Every selector is run against the page before being returned, so `report`/`coverage` tell you which fields are actually reliable.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"country": {
"description": "ISO country code for the proxy exit",
"maxLength": 2,
"minLength": 2,
"type": "string"
},
"fields": {
"additionalProperties": {
"type": "string"
},
"description": "What to extract, as { field_name: \"plain-English description\" } — e.g. { \"price\": \"the product price\", \"specs\": \"every spec bullet, as a list\" }. Max 25.",
"type": "object"
},
"html": {
"description": "Markup you already have, instead of fetching a URL (no proxy bandwidth used)",
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"prompt": {
"description": "Free-text alternative to `fields` — the model picks and names the fields itself",
"type": "string"
},
"render": {
"description": "Learn from the browser-rendered DOM instead of the raw HTML (needed for SPA pages)",
"type": "boolean"
},
"url": {
"description": "The page to learn the layout from",
"format": "uri",
"type": "string"
}
},
"required": [
"context",
"llm_model"
],
"type": "object"
},
"name": "generate_parser",
"outputSchema": null
},
{
"description": "Generate ready-to-use proxy endpoint strings (credentials included) from one of the account's active proxy services — any type: residential, mobile, datacenter, ISP, IPv6. Supports geo targeting (country/state/city, ISP or ASN where the plan allows it), rotating or sticky sessions, HTTP or SOCKS5, and several output formats. Use list_proxies first to get the orderId, and proxy_locations for valid targeting codes. The returned strings plug straight into any HTTP client, e.g. curl -x.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"asn": {
"description": "ASN for Residential/Datacenter Basic targeting, e.g. 'AS12345'",
"type": "string"
},
"city": {
"description": "City (slug from proxy_locations where applicable; 'all' for any)",
"type": "string"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"country": {
"description": "Country code for geo targeting, lowercase, e.g. 'us'",
"maxLength": 10,
"type": "string"
},
"filter": {
"description": "Residential Premium / Mobile V2 pool filter (omit for the full pool)",
"enum": [
"speed",
"speed-quality",
"quality"
],
"type": "string"
},
"format": {
"description": "Output string format (default user:pass@host:port)",
"enum": [
"user:pass@host:port",
"host:port:user:pass",
"http://user:pass@host:port",
"socks5://user:pass@host:port"
],
"type": "string"
},
"gateway": {
"description": "Mobile V2 region gateway (default ww)",
"enum": [
"ww",
"us",
"eu",
"as"
],
"type": "string"
},
"ip": {
"description": "Mobile V2 only: a whitelisted IP (see whitelist_ip) to fetch the IP-auth proxy list instead of user:pass proxies",
"type": "string"
},
"isp": {
"description": "ISP code for Residential Premium / Mobile V2 targeting (from proxy_locations tree, e.g. 'tmobile')",
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"orderId": {
"description": "The proxy service's orderId (from list_proxies)",
"type": "string"
},
"protocol": {
"description": "Proxy protocol (default http)",
"enum": [
"http",
"socks5"
],
"type": "string"
},
"quantity": {
"description": "Number of proxy strings (default 10)",
"maximum": 10000,
"minimum": 1,
"type": "integer"
},
"rotation": {
"description": "rotating (default): new IP per request. sticky: keep the IP for sessionTime. static: IPv6 only, fixed session with no TTL.",
"enum": [
"rotating",
"sticky",
"static"
],
"type": "string"
},
"sessionTime": {
"description": "Sticky session duration in minutes (default 10; Residential Basic/Datacenter minimum 3)",
"maximum": 1440,
"minimum": 1,
"type": "integer"
},
"state": {
"description": "State/region (Residential Premium & Mobile V2: use the slug from proxy_locations; 'all' for any)",
"type": "string"
},
"strict": {
"description": "Residential/Datacenter Basic: true allows fallback to nearby locations when the exact target has no IPs",
"type": "boolean"
}
},
"required": [
"orderId",
"context",
"llm_model"
],
"type": "object"
},
"name": "generate_proxies",
"outputSchema": null
},
{
"description": "Regenerate a preset's selectors now (the manual trigger for the automatic repair). Refetches the source page and adopts new selectors ONLY if they extract more than the current ones — a heal that finds nothing better leaves the preset untouched and is not billed.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"force": {
"description": "Bypass the cooldown between heals",
"type": "boolean"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"preset_id": {
"description": "The preset id",
"type": "string"
}
},
"required": [
"preset_id",
"context",
"llm_model"
],
"type": "object"
},
"name": "heal_parser_preset",
"outputSchema": null
},
{
"description": "List the ready-made Collectors: paid, versioned scrapers you run with a semantic input (keyword + location, place id, product id, domain…) instead of URLs — e.g. web_search, search_images, search_videos, keyword_ideas, amazon_search, amazon_product, ebay_search, aliexpress_search, linkedin_jobs, indeed_jobs, reddit_posts, youtube_search, youtube_channel, instagram_profile, tiktok_profile, tiktok_video, linkedin_profile, linkedin_company, zillow_search, zillow_property, app_store_apps, app_store_reviews, google_play_apps, google_maps_places, place_reviews, google_jobs, google_news, google_shopping, product_offers, hotels, google_flights, google_events, google_trends, google_autocomplete, google_lens, youtube_video, ebay_product, flipkart_search, idealista_search, kleinanzeigen_search, autotrader_search, github_repos, hacker_news, coingecko_coins, wikipedia_articles, yahoo_finance, stackoverflow, steam, npm_packages, sec_filings, defillama, wayback_machine, clinical_trials, certificate_transparency, wikidata, nvd_cve, openfda, openalex, pypi_packages, exchange_rates, gleif_lei, docker_hub, crates_io, world_bank, openlibrary_books, arxiv_papers, weather_forecast, whois_domain, dns_records, itunes_search, local_business_leads, site_contacts, company_profile, business_directory. Returns each collector's slug, input/output schema, example input, price per delivered result and current health. Billing is pay-per-success: only delivered rows are charged.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"category": {
"description": "Optional category filter (e.g. 'local', 'ecommerce', 'jobs', 'news', 'travel', 'leads', 'finance', 'dev', 'gaming', 'osint', 'research', 'classifieds', 'knowledge')",
"maxLength": 40,
"type": "string"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
}
},
"required": [
"context",
"llm_model"
],
"type": "object"
},
"name": "list_collectors",
"outputSchema": null
},
{
"description": "List your stored parser presets with their version, health stats and changelog.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
}
},
"required": [
"context",
"llm_model"
],
"type": "object"
},
"name": "list_parser_presets",
"outputSchema": null
},
{
"description": "List the account's proxy services of every type — Residential Basic/Premium/Private, Mobile, Mobile V2, Datacenter (static or traffic-based), ISP, IPv6 — with plan type, bandwidth left, expiry, whitelisted IPs and the orderId to pass to generate_proxies. Call this first to see which proxy plans are available.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"active": {
"description": "true: only non-expired services (recommended). false: only expired. Omit for all.",
"type": "boolean"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"limit": {
"description": "Max services returned (default 50)",
"maximum": 100,
"minimum": 1,
"type": "integer"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"offset": {
"description": "Pagination offset (default 0)",
"minimum": 0,
"type": "integer"
},
"planType": {
"description": "Only services of this plan type",
"enum": [
"residentialbasic",
"residentialpremium",
"resiprivate",
"isp",
"datacenter",
"datacentertraffic",
"ipv6",
"mobile",
"mobile_v2"
],
"type": "string"
}
},
"required": [
"context",
"llm_model"
],
"type": "object"
},
"name": "list_proxies",
"outputSchema": null
},
{
"description": "Discover a site's URLs fast (robots.txt sitemaps + /sitemap.xml + homepage links) without a full crawl. Returns up to `limit` URLs (default 100) plus the site-wide `total` and a per-section `summary` (e.g. '/blog': 1988) so you see the site's shape without the full list. Narrow with `search` (substring filter — the primary way to find specific pages) or set group_by 'path' for the path tree with counts instead of URLs.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"group_by": {
"description": "path: return the path tree with per-prefix counts instead of the flat URL list",
"enum": [
"path"
],
"type": "string"
},
"includeSubdomains": {
"description": "Include subdomains of the seed host",
"type": "boolean"
},
"limit": {
"description": "Max URLs returned (default 100). `total`/`summary` always cover the whole site.",
"maximum": 5000,
"minimum": 1,
"type": "integer"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"search": {
"description": "Only return URLs containing this substring — use this to narrow before raising limit",
"type": "string"
},
"url": {
"description": "The site URL to map",
"format": "uri",
"type": "string"
}
},
"required": [
"url",
"context",
"llm_model"
],
"type": "object"
},
"name": "map",
"outputSchema": null
},
{
"description": "How well a stored parser is still working: success rate per field, mean coverage over the recent runs, and whether it now counts as decayed (i.e. the site probably changed).",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"preset_id": {
"description": "The preset id returned by save_parser_preset",
"type": "string"
}
},
"required": [
"preset_id",
"context",
"llm_model"
],
"type": "object"
},
"name": "parser_preset_stats",
"outputSchema": null
},
{
"description": "Discover valid geo-targeting values for a proxy plan type before calling generate_proxies: countries, states, cities, ASNs, or the full location tree (countries → regions → cities → ISPs). Use level 'tree' for Residential Premium / Mobile V2 slugs and ISP codes, or for the static datacenter gateway list; note the tree can be large.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"country": {
"description": "Country code, required for states/cities, optional filter for asns",
"maxLength": 10,
"type": "string"
},
"level": {
"description": "countries (default) | states (needs country) | cities (needs country) | asns | tree (full location tree: residentialpremium, mobile/mobile_v2, datacenter)",
"enum": [
"countries",
"states",
"cities",
"asns",
"tree"
],
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"planType": {
"description": "The plan type to look up (same value as list_proxies planType)",
"enum": [
"residentialbasic",
"residentialpremium",
"resiprivate",
"isp",
"datacenter",
"datacentertraffic",
"ipv6",
"mobile",
"mobile_v2"
],
"type": "string"
},
"state": {
"description": "Cities only: filter by state",
"type": "string"
}
},
"required": [
"planType",
"context",
"llm_model"
],
"type": "object"
},
"name": "proxy_locations",
"outputSchema": null
},
{
"description": "Send feedback to the QuanticData team: a bug (a result that is wrong or empty after a retry, a broken parser, a blocked page that should work), a missing tool, site, option or collector, or a question about behaviour or pricing. Works WITHOUT an API key. Call it when a tool clearly failed at its job, when the user needs something no tool here covers, or when the user asks you to report something. Do NOT call it for a missing, rejected or unset API key, for the no-key trial or the balance being used up, for a network timeout, or for your own wrong input: those are fixed by the user, not by the team, and the tool that failed already told you how. One report per issue, never one per retry. Every report is read by the people who build the product.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"actual": {
"description": "What actually came back (error text, empty payload, wrong fields…). Trim page content; a short excerpt is enough.",
"maxLength": 2000,
"type": "string"
},
"contact": {
"description": "Optional email or handle to follow up on — ask the user before sending it; never send it unasked.",
"maxLength": 200,
"type": "string"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"expected": {
"description": "What the user needed to get back",
"maxLength": 2000,
"type": "string"
},
"kind": {
"description": "bug: something returned wrong/empty/errored. feature: a missing tool, option, site or collector. question: unclear behaviour or pricing. other: anything else.",
"enum": [
"bug",
"feature",
"question",
"other"
],
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"message": {
"description": "What happened or what is missing, in plain words. Include the URL/query/collector involved when there is one.",
"maxLength": 4000,
"minLength": 10,
"type": "string"
},
"tool": {
"description": "Name of the tool involved, e.g. 'scrape' or 'run_collector' (omit for general feedback)",
"maxLength": 64,
"type": "string"
},
"tool_input": {
"additionalProperties": {},
"description": "The arguments passed to the failing tool, so the team can reproduce it. Leave out cookies, credentials and anything private.",
"type": "object"
}
},
"required": [
"kind",
"message",
"context",
"llm_model"
],
"type": "object"
},
"name": "report",
"outputSchema": null
},
{
"description": "Run a Collector by slug with a semantic input (see list_collectors for each collector's inputSchema and example). Short runs return the rows inline; long runs return 202 with a run_id + statusUrl — poll with collector_run_status. Results are billed per delivered row (never for failures). Set `async` true to force background execution.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"async": {
"description": "Force background execution and return a run_id to poll",
"type": "boolean"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"input": {
"additionalProperties": {},
"description": "Input fields matching the collector's inputSchema (e.g. { keyword: 'dentist', location: 'Austin, TX', max_results: 20 })",
"type": "object"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"slug": {
"description": "Collector slug from list_collectors, e.g. 'google_maps_places'",
"minLength": 1,
"type": "string"
}
},
"required": [
"slug",
"input",
"context",
"llm_model"
],
"type": "object"
},
"name": "run_collector",
"outputSchema": null
},
{
"description": "Store a generated parser under a name so it can be reused by id. Scrape later with scrape's `preset_id` instead of repeating the selectors, and every run is scored per field — when the recent success rate decays (the site redesigned), the preset regenerates itself from `source_url` and bumps a version. Give it a source_url whenever you can: without one it can never self-heal.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"auto_heal": {
"description": "Regenerate automatically on decay (default true when source_url is set)",
"type": "boolean"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"fields": {
"additionalProperties": {
"type": "string"
},
"description": "The original field descriptions, so a self-heal regenerates the same shape",
"type": "object"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"name": {
"description": "A name you'll recognise, e.g. 'amazon product page'",
"maxLength": 120,
"type": "string"
},
"parser": {
"additionalProperties": {},
"description": "The parser to store — normally the `parser` object returned by generate_parser",
"type": "object"
},
"render": {
"description": "The page needs a browser render to show its content",
"type": "boolean"
},
"source_url": {
"description": "Page to relearn from when the parser decays — required for self-healing",
"format": "uri",
"type": "string"
}
},
"required": [
"name",
"parser",
"context",
"llm_model"
],
"type": "object"
},
"name": "save_parser_preset",
"outputSchema": null
},
{
"description": "Scrape a single web page through a residential proxy and return it as clean Markdown (or HTML/text). Uses a real Chrome TLS fingerprint by default and only spins up a headless browser if the page is bot-challenged. Optionally run structured extraction (CSS selectors) or AI extraction (natural-language prompt). Markdown keeps the complete page by default (content_mode 'smart': everything except nav/footer/cookie chrome, with GFM tables and absolutized links); to inspect a page's raw no-JS/SEO fallback use format 'html'.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"actions": {
"description": "Ordered browser interactions before capture (forces a render). Each is one object: {\"click\":\"#sel\"}, {\"clickText\":\"Accept\"} (click by visible text — dismiss a consent wall without knowing its CSS), {\"type\":{\"selector\":\"#q\",\"text\":\"shoes\"}}, {\"scroll\":\"bottom\"}, {\"wait\":1000}, {\"waitForSelector\":\".results\"}. Add \"optional\":true to skip a miss, or \"timeoutMs\":N to bound one action.",
"items": {
"additionalProperties": {},
"type": "object"
},
"maxItems": 20,
"type": "array"
},
"ai_prompt": {
"description": "Natural-language instruction — the LLM turns the page into structured JSON",
"type": "string"
},
"ai_schema": {
"additionalProperties": {},
"description": "JSON Schema for deterministic AI extraction; returned under payload.ai.data",
"type": "object"
},
"app_state": {
"anyOf": [
{
"type": "boolean"
},
{
"enum": [
"auto",
"raw"
],
"type": "string"
}
],
"description": "Mine the page's own hydration state (Next.js __NEXT_DATA__, Nuxt, embedded JSON islands) into payload.metadata.appState. This is where SPAs keep the real data — prices behind a picker, stock, download counts, listings — even when the DOM shows only a shell, so it often answers the question without a browser render. true/'auto': pruned to the informative parts (recommended). 'raw': the complete blobs, up to 512KB."
},
"chunk": {
"additionalProperties": false,
"description": "Segment the output into payload.chunks[] for RAG/vector-DB ingestion — each chunk carries its heading path and token count. Fences and tables are never split.",
"properties": {
"by": {
"enum": [
"heading",
"sentence",
"tokens"
],
"type": "string"
},
"overlap": {
"maximum": 100000,
"minimum": 0,
"type": "integer"
},
"size": {
"maximum": 100000,
"minimum": 1,
"type": "integer"
}
},
"type": "object"
},
"content_mode": {
"description": "smart (default): whole page minus nav/footer/cookie chrome. article: Readability main article only (news/blogs). full: entire body as-is.",
"enum": [
"smart",
"article",
"full"
],
"type": "string"
},
"content_modes": {
"description": "Return several content scopes from ONE fetch under payload.contents (e.g. compare smart vs full)",
"items": {
"enum": [
"smart",
"article",
"full"
],
"type": "string"
},
"maxItems": 3,
"type": "array"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"cookies": {
"additionalProperties": {
"type": "string"
},
"description": "Cookies to send as name→value — the simple way to scrape behind a login",
"type": "object"
},
"country": {
"description": "ISO country code for the proxy exit, e.g. 'us'",
"maxLength": 2,
"minLength": 2,
"type": "string"
},
"engine": {
"description": "auto (default): TLS tier, escalate to browser on block. tls: never escalate — exactly what a pure HTTP bot (no JS) sees, right for SEO checks. render: force browser.",
"enum": [
"auto",
"tls",
"fetch",
"render"
],
"type": "string"
},
"extract": {
"additionalProperties": {},
"description": "Structured-extraction schema: { field: \"css selector\" | { selector, attr, all, fns } }. `fns` is a transform pipeline run on the value — e.g. { \"price\": { \"selector\": \".price\", \"fns\": [\"amount_from_string\"] } } returns a number, not text. Functions: amount_from_string, amount_range_from_string, convert_to_float/int/str, trim, lower, upper, {regex_search|regex_find_all: \"pat\"}, {replace:{from,to}}, {join:\",\"}, {select_nth:0}, length, unique, max, min, average, product.",
"type": "object"
},
"fetch_resource": {
"description": "Regex matched against the page's network requests: the first matching response's BODY becomes the result instead of the page HTML (e.g. '/api/products' to get an SPA's JSON directly). Forces a render. Fails with 504 if nothing matches.",
"maxLength": 500,
"type": "string"
},
"format": {
"description": "Output format (default markdown)",
"enum": [
"markdown",
"html",
"text"
],
"type": "string"
},
"formats": {
"description": "Additional formats to return together in payload.formats, e.g. ['markdown','text']",
"items": {
"enum": [
"markdown",
"html",
"text"
],
"type": "string"
},
"maxItems": 3,
"type": "array"
},
"frontmatter": {
"description": "Prepend YAML front-matter (title, url, canonical, description, author, date) so the markdown is self-contained for RAG/Obsidian pipelines",
"type": "boolean"
},
"highlights": {
"description": "With `query`: also return the N most relevant passages in payload.highlights",
"maximum": 20,
"minimum": 1,
"type": "integer"
},
"html": {
"description": "Convert HTML you already have instead of fetching: no proxy bandwidth is used, and the full parser pipeline still applies. Pass `url` too if you want relative links absolutized.",
"type": "string"
},
"images_mode": {
"description": "inline (default) keeps ; 'alt' keeps only alt text; 'strip' removes images",
"enum": [
"inline",
"alt",
"strip"
],
"type": "string"
},
"include_links": {
"description": "Return all de-duplicated absolute page links in payload.links",
"type": "boolean"
},
"links_mode": {
"description": "Link rendering. inline (default): [text](url). footnote: URLs moved to a numbered reference list at the end. strip: keep only the link text — cuts 30-48% of the tokens on link-dense pages when you only need the prose.",
"enum": [
"inline",
"footnote",
"strip"
],
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"max_tokens": {
"description": "Cap the markdown at ~this many tokens, cutting at a section boundary (never inside a table or code block) and noting how much was omitted",
"maximum": 2000000,
"minimum": 200,
"type": "integer"
},
"mode": {
"description": "summary: return only metadata (title, description, canonical, contentLength, status, engine, bytes) with no page content — use this when auditing pages instead of reading them",
"enum": [
"summary"
],
"type": "string"
},
"parser": {
"additionalProperties": false,
"description": "Your own parsing rules, as CSS selector lists — use these when you know the page and don't want to rely on heuristics. include: keep ONLY these subtrees (targeted extraction, e.g. ['article.post']). exclude: delete site-specific chrome we kept. keep: protect a section (sidebar, dialog, form) that smart mode would strip.",
"properties": {
"exclude": {
"items": {
"type": "string"
},
"maxItems": 25,
"type": "array"
},
"include": {
"items": {
"type": "string"
},
"maxItems": 25,
"type": "array"
},
"keep": {
"items": {
"type": "string"
},
"maxItems": 25,
"type": "array"
}
},
"type": "object"
},
"preset_id": {
"description": "Run a stored parser preset (see save_parser_preset) instead of passing `extract` selectors. Results land in payload.data exactly the same way, and the run is scored so the preset can detect decay and self-heal.",
"type": "string"
},
"query": {
"description": "What you are looking for on the page. Keeps only the relevant sections (BM25 scoring over blocks, headings preserved) — the way to read one fact off a huge page without spending its whole token budget.",
"maxLength": 512,
"type": "string"
},
"render": {
"description": "Force the headless browser (JS execution)",
"type": "boolean"
},
"reveal_hidden": {
"description": "Render tier only: before capturing, open <details>/accordions and click through every tab, appending each revealed panel to the page. Use it for tabbed code samples or spec accordions where a plain render captures only the visible variant.",
"type": "boolean"
},
"summary_sections": {
"description": "Append 'Links on this page' / 'Images on this page' sections — handy when deciding the next hop",
"type": "boolean"
},
"toc": {
"description": "Prepend a table of contents built from the page headings",
"type": "boolean"
},
"url": {
"description": "The page URL to scrape (optional only when you pass `html` to convert)",
"format": "uri",
"type": "string"
},
"xhr": {
"description": "Record the page's XHR/fetch traffic (URL, method, status, response body) into payload.xhr. Forces a browser render. An SPA's own JSON API is usually far cleaner than its DOM — use this to DISCOVER the API, then fetch_resource to return it directly.",
"type": "boolean"
}
},
"required": [
"context",
"llm_model"
],
"type": "object"
},
"name": "scrape",
"outputSchema": null
},
{
"description": "Run structured Google, Bing or DuckDuckGo searches through a residential proxy. Bing supports web, shopping, images, news, videos, places/maps and autocomplete over HTTP, including Copilot AI answers and citations when Bing returns them. Google web search also parses rich blocks directly from its HTTP response.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"accommodation_type": {
"description": "Hotels: property kind (default hotels)",
"enum": [
"hotels",
"vacation_rentals"
],
"type": "string"
},
"adults": {
"description": "Hotels: number of adults",
"maximum": 10,
"minimum": 1,
"type": "integer"
},
"arrival_id": {
"description": "Flights: arrival airport IATA code, e.g. 'LAX'",
"type": "string"
},
"browser": {
"description": "TLS/browser identity for the fetch path",
"enum": [
"chrome",
"firefox",
"safari"
],
"type": "string"
},
"check_in_date": {
"description": "Hotels: check-in date YYYY-MM-DD",
"type": "string"
},
"check_out_date": {
"description": "Hotels: check-out date YYYY-MM-DD",
"type": "string"
},
"children_ages": {
"description": "Hotels: children's ages, e.g. [5, 7]",
"items": {
"maximum": 17,
"minimum": 0,
"type": "integer"
},
"type": "array"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"country": {
"description": "ISO country code, e.g. 'us'",
"maxLength": 2,
"minLength": 2,
"type": "string"
},
"currency": {
"description": "Hotels/Flights: price currency, e.g. 'EUR'",
"maxLength": 3,
"minLength": 3,
"type": "string"
},
"data_id": {
"description": "Maps data id, hex fid '0x…:0x…' (from maps/place_details results) — required for reviews",
"type": "string"
},
"departure_id": {
"description": "Flights: departure airport IATA code, e.g. 'JFK'",
"type": "string"
},
"device": {
"description": "SERP device shape (default desktop)",
"enum": [
"desktop",
"mobile"
],
"type": "string"
},
"engine": {
"description": "Search engine (default google)",
"enum": [
"google",
"bing",
"duckduckgo"
],
"type": "string"
},
"exact_matches": {
"description": "Lens: return the exact-matches tab (pages using this exact image) instead of visual matches",
"type": "boolean"
},
"filter": {
"description": "Reviews: only reviews whose text contains this keyword",
"type": "string"
},
"free_cancellation": {
"description": "Hotels: only offers with free cancellation",
"type": "boolean"
},
"google_params": {
"additionalProperties": {
"type": [
"string",
"number"
]
},
"description": "Additional Google query parameters not modeled above",
"type": "object"
},
"gps_coordinates": {
"description": "Maps: center the search on 'lat,lon' or 'lat,lon,zoom' (zoom 3-21)",
"type": "string"
},
"image_url": {
"description": "Lens: publicly reachable image URL to reverse-search",
"type": "string"
},
"lang": {
"description": "Search UI language, e.g. 'en' or 'it'",
"maxLength": 10,
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"location": {
"description": "Search from this location, e.g. 'Milan, Italy' (encoded to Google's uule server-side)",
"type": "string"
},
"next_page_token": {
"description": "Reviews: continuation token from the previous response's serpapi_pagination",
"type": "string"
},
"nfpr": {
"description": "Disable Google spelling correction",
"type": "boolean"
},
"num": {
"description": "How many organic results to aim for (default 10, max 100). Google serves ~10 per page, so a larger num is satisfied by fetching consecutive pages and merging them — it is NOT ignored. `search_metadata.search_url` is necessarily the first page's URL and therefore shows num=<page size>; `search_metadata.paging` reports what was actually requested, the page size, and how many pages were fetched. Getting fewer results than requested means Google ran out, not that num was dropped. Use `page` to address one specific page, or search_bulk for many queries.",
"maximum": 100,
"minimum": 1,
"type": "integer"
},
"outbound_date": {
"description": "Flights: outbound date YYYY-MM-DD",
"type": "string"
},
"page": {
"description": "Result page, 1-based (default 1). The response's pagination.available_pages lists which pages exist; use search_bulk to fetch many pages at once.",
"minimum": 1,
"type": "integer"
},
"place_id": {
"description": "Google Maps place id for place_details — the hex '0x…:0x…' fid from maps/places results, a 'ChIJ…' place id or a numeric cid all work; served from Maps' place card over HTTP (name, address, phone, website, rating, reviews, category, weekly hours, open state) in about a second",
"type": "string"
},
"product_id": {
"description": "Google Shopping product id (the product_id of a shopping result — the rendered grid carries it) for search_type=product: the seller list with price, old price, discount, stock and delivery per merchant. Without it, pass a query and the first product is opened.",
"type": "string"
},
"product_ids": {
"description": "shopping only: render the Shopping grid so each result carries product_id/offer_id (the input of search_type=product). Costs a render; the default no-JS shopping page has no ids.",
"type": "boolean"
},
"query": {
"description": "The search query (optional for place_details/product/flights/lens/reviews, which are ID/URL-addressed)",
"type": "string"
},
"render": {
"description": "Force browser rendering where supported; Google/Bing web search rich blocks are parsed over HTTP",
"type": "boolean"
},
"return_date": {
"description": "Flights: return date YYYY-MM-DD (omit for one-way)",
"type": "string"
},
"safe": {
"description": "Google SafeSearch setting",
"enum": [
"active",
"off"
],
"type": "string"
},
"search_type": {
"description": "Vertical (default search). Bing supports shopping/images/news/videos/places/maps/autocomplete. Google additionally supports scholar/jobs/place_details/hotels/flights/events/product/lens/reviews; maps accepts gps_coordinates, place_details uses place_id, and reviews uses data_id.",
"enum": [
"search",
"shopping",
"images",
"news",
"places",
"maps",
"videos",
"scholar",
"jobs",
"autocomplete",
"place_details",
"hotels",
"flights",
"events",
"product",
"lens",
"reviews",
"trends"
],
"type": "string"
},
"sort_by": {
"description": "Reviews: sort order (default relevance)",
"enum": [
"relevance",
"newest",
"highest_rating",
"lowest_rating"
],
"type": "string"
},
"start": {
"description": "Result offset alias (0, 10, 20…)",
"minimum": 0,
"type": "integer"
},
"timeframe": {
"description": "Trends only: Google timeframe token — 'today 12-m' (default), 'now 7-d', or an explicit 'YYYY-MM-DD YYYY-MM-DD' range",
"type": "string"
},
"uule": {
"description": "Geo token: encoded uule, or raw coordinates 'lat,lon' / 'lat,lon,radius_m' (encoded server-side)",
"type": "string"
},
"wait_for": {
"description": "Rendered path: wait for this CSS selector before parsing late panels",
"maxLength": 512,
"type": "string"
}
},
"required": [
"context",
"llm_model"
],
"type": "object"
},
"name": "search",
"outputSchema": null
},
{
"description": "Search the live web, fetch the top organic pages as clean Markdown, and return citation-ready numbered sources plus one token-bounded `context` string ready for an AI prompt. Use this when the goal is answering/researching, and use `search` when raw SERP structure or a specialized vertical is needed.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"country": {
"description": "ISO country code for search and proxy geo",
"maxLength": 2,
"minLength": 2,
"type": "string"
},
"engine": {
"description": "Search engine (default google)",
"enum": [
"google",
"bing",
"duckduckgo"
],
"type": "string"
},
"fetch_content": {
"description": "False returns snippet-only context without fetching result pages",
"type": "boolean"
},
"lang": {
"description": "Search UI language, e.g. 'en' or 'it'",
"maxLength": 10,
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"max_tokens": {
"description": "Maximum estimated tokens in the assembled context (default 8000)",
"maximum": 50000,
"minimum": 500,
"type": "integer"
},
"query": {
"description": "The research/search query",
"minLength": 1,
"type": "string"
},
"top_n": {
"description": "Top organic pages to fetch (default 3, max 5)",
"maximum": 5,
"minimum": 1,
"type": "integer"
}
},
"required": [
"query",
"context",
"llm_model"
],
"type": "object"
},
"name": "search_and_read",
"outputSchema": null
},
{
"description": "Paginate ONE search query asynchronously and merge deduplicated organic results. Page-one AI Overview/PAA/Knowledge Graph/answer enrichments are retained; set render:true to request those Google JS blocks. Billed per page actually fetched, with unavailable pages refunded.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"browser": {
"description": "Fetch-path browser identity",
"enum": [
"chrome",
"firefox",
"safari"
],
"type": "string"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"country": {
"description": "ISO country code, e.g. 'us'",
"maxLength": 2,
"minLength": 2,
"type": "string"
},
"device": {
"description": "SERP device shape",
"enum": [
"desktop",
"mobile"
],
"type": "string"
},
"engine": {
"description": "Search engine (default google)",
"enum": [
"google",
"bing",
"duckduckgo"
],
"type": "string"
},
"google_params": {
"additionalProperties": {
"type": [
"string",
"number"
]
},
"description": "Additional Google query parameters",
"type": "object"
},
"lang": {
"description": "UI language, e.g. 'en'",
"maxLength": 10,
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"location": {
"description": "Search location, e.g. 'Milan, Italy'",
"maxLength": 256,
"type": "string"
},
"max_pages": {
"description": "Max pages to fetch (1-10, default 5). Stops early when Google has no more pages.",
"maximum": 10,
"minimum": 1,
"type": "integer"
},
"nfpr": {
"description": "Disable Google spelling correction",
"type": "boolean"
},
"query": {
"description": "The search query to paginate",
"minLength": 1,
"type": "string"
},
"render": {
"description": "Force rendering to capture page-one Google JS enrichments",
"type": "boolean"
},
"safe": {
"description": "Google SafeSearch setting",
"enum": [
"active",
"off"
],
"type": "string"
},
"search_type": {
"description": "Vertical to paginate (default search)",
"enum": [
"search",
"news",
"videos",
"images",
"shopping"
],
"type": "string"
},
"uule": {
"description": "Encoded geo token or raw coordinates",
"maxLength": 512,
"type": "string"
},
"wait_for": {
"description": "Rendered path CSS selector for late panels",
"maxLength": 512,
"type": "string"
},
"webhook": {
"description": "Public URL to POST the finished job to",
"format": "uri",
"type": "string"
}
},
"required": [
"query",
"context",
"llm_model"
],
"type": "object"
},
"name": "search_bulk",
"outputSchema": null
},
{
"description": "Poll a bulk search job for progress and merged organic results. Polls are incremental: pass the previous response's `nextCursor` as `since` to receive only the organic results gathered after your last poll.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"jobId": {
"description": "The bulk search job id returned by search_bulk",
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"since": {
"description": "Organic cursor from the previous poll's `nextCursor` — returns only newer results",
"minimum": 0,
"type": "integer"
}
},
"required": [
"jobId",
"context",
"llm_model"
],
"type": "object"
},
"name": "search_bulk_status",
"outputSchema": null
},
{
"description": "Audit a URL's SEO in one call: fetches it twice — as a pure HTTP bot (no JS) and fully rendered — and returns both views (title, description, canonical, h1, word count) plus the diff (JS-only content, changed title/description, canonical missing without JS) and bot-facing meta (robots, Open Graph, JSON-LD types). Use this instead of scraping manually when checking how a page indexes.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"country": {
"description": "ISO country code for the proxy exit, e.g. 'us'",
"maxLength": 2,
"minLength": 2,
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"no_render": {
"description": "Skip the rendered pass (cheaper — returns the no-JS view only, no diff)",
"type": "boolean"
},
"url": {
"description": "The page URL to audit",
"format": "uri",
"type": "string"
}
},
"required": [
"url",
"context",
"llm_model"
],
"type": "object"
},
"name": "seo_audit",
"outputSchema": null
},
{
"description": "Manage IP-auth whitelisting on a proxy service (Residential Basic, Datacenter, ISP, IPv6, Mobile): add or remove an IP, or list the current entries. A whitelisted machine uses the proxies without username/password — required for the Mobile V2 IP-auth proxy list. Residential Premium/Private use user:pass auth and don't need this.",
"inputSchema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"action": {
"description": "What to do with the order's whitelist",
"enum": [
"add",
"list",
"remove"
],
"type": "string"
},
"city": {
"description": "Mobile add: city slug",
"type": "string"
},
"context": {
"description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
"type": "string"
},
"country": {
"description": "Mobile add: geo targeting for the ports, e.g. 'us'",
"maxLength": 10,
"type": "string"
},
"ip": {
"description": "The IP to add/remove (required for add and remove)",
"type": "string"
},
"isp": {
"description": "Mobile add: ISP code, e.g. 'tmobile'",
"type": "string"
},
"llm_model": {
"description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
"type": "string"
},
"orderId": {
"description": "The proxy service's orderId (from list_proxies)",
"type": "string"
},
"ports_count": {
"description": "Mobile add: number of ports to allocate",
"maximum": 1000,
"minimum": 1,
"type": "integer"
},
"protocol": {
"description": "Mobile add: protocol for the allocated ports",
"enum": [
"HTTP",
"SOCKS5"
],
"type": "string"
},
"region": {
"description": "Mobile add: region slug",
"type": "string"
},
"sticky": {
"description": "Mobile add: keep the same IP per port",
"type": "boolean"
},
"ttl": {
"description": "Mobile add: sticky session TTL in seconds",
"minimum": 1,
"type": "integer"
}
},
"required": [
"action",
"orderId",
"context",
"llm_model"
],
"type": "object"
},
"name": "whitelist_ip",
"outputSchema": null
}
]
}Verify it yourself
curl -s https://api.teppi.xyz/v1/evidence/sha256:358189d226bdce18c7a97fc93d75369b07eaae6fdfc2e183fc5419fd7ed3768f | sha256sum