MCP serverio.github.theoddden/terradev
Cross-cloud GPU orchestration CLI.
Overview
Score?
UNRATED 0.589
of what a free look can see, on 25 looks
Looks
27
last 5 hr ago
Tools
277
changed 5 days ago
More info
URL
terradev-mcp.terradev.cloud/sse
sse
Says it is
terradev-mcp 1.30.0
protocol 2025-06-18
In the record since
25 days ago
Among servers18,413 with a card
0median 0.606 · this server 0.589 · highest on record 0.8561
Toolsfrom sha256:77aac112b5…896d84 · +26 −26 5 days ago
| Tool | Schema |
|---|---|
| active_context Get current Terradev state: running training jobs, active instances, spend-to-date, alerts. Call this on session start to resume context from previous sessions. |
input · no output |
| agent_agentic_serving_configure Configure agentic inference serving settings. |
input · no output |
| agent_agentic_serving_helm_values Print Helm values for agentic inference deployment. |
input · no output |
| agent_agentic_serving_k8s Print K8s deployment manifests for agentic inference. |
input · no output |
| agent_agentic_serving_launch_args Print engine launch arguments for copy-paste. |
input · no output |
| agent_agentic_serving_lmcache_env Print LMCache environment variables. |
input · no output |
| agent_agentic_serving_show_config Show current agentic serving configuration. |
input · no output |
| agent_cost Show real-time cost breakdown for a fleet by tier. |
input · no output |
| agent_deploy Provision a heterogeneous agent fleet across all tiers simultaneously. |
input · no output |
| agent_langchain_create_langgraph Create a LangGraph workflow. |
input · no output |
| agent_langchain_create_pipeline Create an SGLang pipeline. |
input · no output |
| agent_langchain_create_workflow Create a LangChain workflow. |
input · no output |
| agent_langchain_test Test connection to LangChain service. |
input · no output |
| agent_langgraph_create_workflow Create a LangGraph workflow. |
input · no output |
| agent_langgraph_deploy Deploy a workflow. |
input · no output |
| agent_langgraph_status Get workflow status. |
input · no output |
| agent_langgraph_test Test connection to LangGraph service. |
input · no output |
| agent_letta_chat Send a message to a Letta agent. |
input · no output |
| agent_letta_create Create a new stateful Letta agent. |
input · no output |
| agent_letta_delete Delete a Letta agent. |
input · no output |
| agent_letta_list List Letta agents. |
input · no output |
| agent_letta_remember Teach a Letta agent a durable fact. |
input · no output |
| agent_letta_status Show the state of a Letta agent. |
input · no output |
| agent_list List all known agent fleets. |
input · no output |
| agent_mem0_add Store a memory in Mem0 for an agent or user. |
input · no output |
| agent_mem0_configure Configure Mem0 credentials and defaults. |
input · no output |
| agent_mem0_delete Delete a memory by ID. |
input · no output |
| agent_mem0_forget Delete all memories matching an entity scope. |
input · no output |
| agent_mem0_get Get a single memory by ID. |
input · no output |
| agent_mem0_list List memories for an entity scope. |
input · no output |
| agent_mem0_search Search agent/user memories. |
input · no output |
| agent_mem0_test Test connection to Mem0. |
input · no output |
| agent_mem0_update Update a memory by ID. |
input · no output |
| agent_plan Plan a heterogeneous agent fleet without provisioning. |
input · no output |
| agent_scale Scale a single fleet tier up or down without affecting other tiers. |
input · no output |
| agent_skill_attach Attach a skill.md to a Letta agent as a durable memory block. |
input · no output |
| agent_skill_init Create a skill.md template for an agent. |
input · no output |
| agent_status Show live status of a fleet — tier health, KV hit rate, queue depth, cost. |
input · no output |
| agent_teardown Terminate all fleet instances and remove fleet state. |
input · no output |
| agent_vector_db_down Teardown a vector database provisioned for an agent fleet. |
input · no output |
| agent_vector_db_up Provision a vector database for an agent fleet. |
input · no output |
| analytics Get cost analytics |
input · no output |
| checkpoint_delete Delete a checkpoint. |
input · no output |
| checkpoint_list List all checkpoints for a training job. |
input · no output |
| checkpoint_promote Promote a checkpoint to a final model path for serving. |
input · no output |
| checkpoint_restore Restore a specific checkpoint for a training job. |
input · no output |
| checkpoint_save Manually trigger a checkpoint save for a running training job. |
input · no output |
| configure_provider Configure provider credentials |
input · no output |
| cost_analyze Deep cost analysis of current GPU infrastructure: per-provider breakdown, utilization efficiency, waste identification, and optimization potential. |
input · no output |
| cost_optimize_recommend Generate actionable cost optimization recommendations: spot migration, GPU right-sizing, provider arbitrage, idle shutdown, and density packing. |
input · no output |
| cost_simulate Simulate cost optimization scenarios with ROI projections. Compare current vs optimized infrastructure costs. |
input · no output |
| create_postgresql_connection Create a PostgreSQL database connection with auto-table creation. Returns a connection ID for subsequent operations. |
input · no output |
| create_sqlite_connection Create a SQLite database connection with auto-table creation. Returns a connection ID for subsequent operations. |
input · no output |
| database_weaviate_create_collection Create a Weaviate collection. |
input · no output |
| database_weaviate_delete_collection Delete a Weaviate collection. |
input · no output |
| database_weaviate_hybrid_search Hybrid vector + BM25 search in a Weaviate collection. |
input · no output |
| database_weaviate_insert Insert objects into a Weaviate collection. |
input · no output |
| database_weaviate_list_collections List Weaviate collections. |
input · no output |
| database_weaviate_query Vector similarity search in a Weaviate collection. |
input · no output |
| database_weaviate_up Initialize a Weaviate connection. |
input · no output |
| deepeval_evaluate Evaluate a single LLM output with a DeepEval metric (AnswerRelevancy, Faithfulness, Hallucination, etc.). |
input · no output |
| deepeval_init Generate a starter DeepEval test file. |
input · no output |
| deepeval_metrics List available DeepEval metrics for LLM evaluation. |
input · no output |
| deepeval_run Run a DeepEval test suite from a Python test file. |
input · no output |
| dvc_diff Show DVC diff between two revisions (e.g. training checkpoints). Shows added, modified, deleted files. |
input · no output |
| dvc_push Push DVC-tracked data to the configured remote storage. |
input · no output |
| dvc_stage_checkpoint Atomic checkpoint staging: DVC add + push + git commit in one operation. Promotes a training checkpoint to versioned storage. |
input · no output |
| dvc_status Get DVC repository status: tracked files, remotes, and changes since last commit. |
input · no output |
| egress_cheapest_route Find the cheapest egress route between cloud providers/regions for model weights or dataset transfer. Supports multi-hop routing. |
input · no output |
| egress_optimize_staging Optimize dataset or model staging across regions by finding the cheapest transfer plan. Integrates with the dataset stager for parallel uploads. |
input · no output |
| environments_approve added Approve a pending environment promotion. |
input · no output |
| environments_create added Create a named deployment environment (dev/staging/prod) with config. |
input · no output |
| environments_history added Show promotion/change history for an environment. |
input · no output |
| environments_list added List deployment environments. |
input · no output |
| environments_promote added Promote a workload or config between environments. |
input · no output |
| get_database_connection Get information about a database connection including type, status, and configuration. |
input · no output |
| governance_compliance_report Generate comprehensive compliance report: consent stats, policy evaluations, data movements, violations. For GDPR/SOC2/HIPAA audits. |
input · no output |
| governance_evaluate_opa Evaluate OPA (Open Policy Agent) policies for data access. Checks region restrictions, classification rules, and compliance requirements. |
input · no output |
| governance_move_data Move data with full governance audit trail. Requires prior consent and OPA policy approval. Tracks integrity, encryption, and compliance. |
input · no output |
| governance_movement_history Get data movement audit log. Filter by user, dataset, or time range. |
input · no output |
| governance_record_consent Record a consent response (granted or denied) for a pending consent request. |
input · no output |
| governance_request_consent Request user consent for data movement across cloud regions. GDPR/SOC2 compliant consent tracking with audit trail. |
input · no output |
| gpu_topology GPU NUMA topology report with intra-GPU XCD (Accelerated Compute Die) awareness. Models MI300X (8 XCDs, 192GB HBM3), MI300A (6 XCDs, 128GB), H200 (unified 141GB HBM3e), H100 (80GB) |
input · no output |
| guardrails_chat Send a message through NeMo Guardrails and return the safety-filtered response. Applies topical, jailbreak, PII, and factcheck rails. |
input · no output |
| guardrails_generate_config Generate default Colang 2.x guardrails configuration files (topical, jailbreak, PII, factcheck rails). |
input · no output |
| guardrails_k8s Generate Kubernetes deployment manifest for NeMo Guardrails server (standalone or sidecar mode). |
input · no output |
| guardrails_test Test connection to NeMo Guardrails server. |
input · no output |
| helm_generate Generate Helm charts from workload specifications. |
input · no output |
| hf_create_endpoint Create a HuggingFace Inference Endpoint (paid GPU endpoint). Supports custom GPU types, regions, and scaling. |
input · no output |
| hf_delete_endpoint Delete a HuggingFace Inference Endpoint. |
input · no output |
| hf_endpoint_infer Run inference on a HuggingFace Inference Endpoint. Supports text generation, embeddings, and custom inputs. |
input · no output |
| hf_endpoint_info Get detailed info about a specific HuggingFace Inference Endpoint: status, URL, scaling config, cost. |
input · no output |
| hf_hardware_compare Compare all hardware options for a HuggingFace model. Returns side-by-side cost, performance, and compatibility analysis. |
input · no output |
| hf_hardware_recommend Get hardware recommendation with cost breakdown for any HuggingFace model. Returns optimal GPU type, estimated cost, and performance score. |
input · no output |
| hf_list_datasets Search and browse HuggingFace Hub datasets. Filter by author and search query. |
input · no output |
| hf_list_endpoints List all active HuggingFace Inference Endpoints with status, URL, and cost. |
input · no output |
| hf_list_models Search and browse HuggingFace Hub models. Filter by author, task, library. Returns model ID, downloads, likes, and tags. |
input · no output |
| hf_model_info Get detailed model info: architecture, size, downloads, license, tags, pipeline_tag, and model card. |
input · no output |
| hf_smart_template Auto-generate an optimized deployment template for any HuggingFace model. Analyzes model size, architecture, and quantization to select optimal hardware and generate ready-to-deplo |
input · no output |
| hf_space_deploy Deploy model to HuggingFace Spaces |
input · no output |
| hf_space_status Get HuggingFace Space deployment status. |
input · no output |
| infer_failover Run health checks and auto-failover for inference endpoints. If a primary endpoint is unhealthy and has a backup configured, traffic automatically shifts to the backup provider. |
input · no output |
| infer_route Semantic-aware inference routing. Analyzes query content across 6 signal dimensions (modality, complexity, domain, language, safety, keywords), applies NUMA-aware endpoint scoring, |
input · no output |
| infer_route_disagg Disaggregated Prefill/Decode routing (DistServe architecture). Splits LLM inference into compute-bound prefill phase (routed to FLOPS-optimized GPUs like H100 SXM) and memory-bound |
input · no output |
| inferx_configure Configure InferX serverless platform credentials. |
input · no output |
| inferx_delete Delete an InferX model deployment. |
input · no output |
| inferx_deploy Deploy model to InferX serverless platform |
input · no output |
| inferx_list List deployed InferX models |
input · no output |
| inferx_optimize Get cost analysis for inference endpoints |
input · no output |
| inferx_quote Get InferX pricing quotes for a GPU type. |
input · no output |
| inferx_status Check InferX endpoint status |
input · no output |
| inferx_usage Get InferX account usage statistics: requests, cost, GPU hours, latency. |
input · no output |
| k8s_create Create Kubernetes cluster with GPU nodes for optimal multi-cloud deployment |
input · no output |
| k8s_destroy Destroy a Kubernetes cluster |
input · no output |
| k8s_device_plugin Configure Kubernetes GPU device plugin settings: time-slicing, MIG strategy, and resource naming. |
input · no output |
| k8s_gpu_operator_install Install NVIDIA GPU Operator on a Kubernetes cluster. Configures driver containers, device plugin, DCGM exporter, and GPU Feature Discovery. |
input · no output |
| k8s_info Get information about a specific cluster |
input · no output |
| k8s_list List Kubernetes clusters |
input · no output |
| k8s_mig_configure Configure Multi-Instance GPU (MIG) partitioning on A100/H100 GPUs. Splits a single GPU into isolated instances for multi-tenant workloads. |
input · no output |
| k8s_time_slicing Configure GPU time-slicing for Kubernetes. Allows multiple pods to share a single GPU with configurable oversubscription. |
input · no output |
| karpenter_create_nodepool added Create a Karpenter GPU NodePool. |
input · no output |
| karpenter_delete_nodepool added Delete a Karpenter NodePool. |
input · no output |
| karpenter_events added Show recent Karpenter scaling events. |
input · no output |
| karpenter_gpu_nodes added List GPU nodes provisioned by Karpenter. |
input · no output |
| karpenter_install added Install Karpenter on an EKS cluster for GPU node autoscaling. |
input · no output |
| karpenter_logs added Show Karpenter controller logs. |
input · no output |
| karpenter_nodepools added List Karpenter NodePools. |
input · no output |
| karpenter_resources added Show Karpenter resource usage and capacity. |
input · no output |
| karpenter_status added Show Karpenter installation/controller status. |
input · no output |
| kserve_generate_yaml Generate a GPU-aware KServe InferenceService YAML manifest with NUMA pinning, resource limits derived from model size and VRAM, and topology hints. |
input · no output |
| kserve_list List KServe InferenceServices in a Kubernetes namespace. |
input · no output |
| kserve_status Get detailed status of a KServe InferenceService including readiness, traffic split, and URL. |
input · no output |
| kv_cache_efficiency added Analyze KV cache efficiency across serving endpoints. |
input · no output |
| langchain_create_sglang_pipeline Create an SGLang model-serving pipeline via LangChain. Connects LangChain agents to SGLang inference endpoints. |
input · no output |
| langchain_create_workflow Create a LangChain workflow. |
input · no output |
| langfuse_configure Configure Langfuse credentials (public key, secret key, host URL). |
input · no output |
| langfuse_datasets List Langfuse datasets for evaluation and fine-tuning. |
input · no output |
| langfuse_export_training_data Export Langfuse traces as instruction/response pairs for LoRA fine-tuning. Filters by quality score. |
input · no output |
| langfuse_k8s Generate Kubernetes deployment manifest for self-hosted Langfuse. |
input · no output |
| langfuse_otel_env Print OTEL environment variables for instrumenting LLM apps to send traces to Langfuse. |
input · no output |
| langfuse_quality Get aggregated quality metrics from Langfuse scores for drift detection. |
input · no output |
| langfuse_score Create an evaluation score for a Langfuse trace (e.g. quality, accuracy, relevance). |
input · no output |
| langfuse_scores List evaluation scores from Langfuse, optionally filtered by trace or score name. |
input · no output |
| langfuse_test Test Langfuse connectivity and list accessible projects. |
input · no output |
| langfuse_trace Get a single Langfuse trace with all observations/spans. |
input · no output |
| langfuse_traces List recent LLM traces from Langfuse. |
input · no output |
| langgraph_create_workflow Create a LangGraph stateful workflow with monitoring. Supports agent graphs, tool calling, and state persistence. |
input · no output |
| langgraph_evaluation_workflow Create an evaluator-optimizer workflow in LangGraph. Generates outputs, evaluates quality, and iteratively improves. |
input · no output |
| langgraph_orchestrator_worker Create an orchestrator-worker pattern workflow in LangGraph. The orchestrator delegates tasks to specialized worker agents. |
input · no output |
| langgraph_workflow_status Get the status and metrics of a LangGraph workflow execution. |
input · no output |
| lineage_checkpoint added Record a data/model lineage checkpoint. |
input · no output |
| lineage_diff added Diff lineage between two artifacts or checkpoints. |
input · no output |
| lineage_export added Export lineage data (JSON/graph) for an artifact. |
input · no output |
| lineage_graph added Render the lineage graph for an artifact. |
input · no output |
| lineage_trace added Trace the lineage of a dataset or model artifact. |
input · no output |
| local_scan Scan local machine and network for available GPU devices. Returns total VRAM pool for local-first provisioning. |
input · no output |
| lora_add Hot-load a LoRA adapter onto a running vLLM endpoint. The adapter becomes immediately available as a model name for inference requests. Uses vLLM's fused_moe_lora kernel for 454% h |
input · no output |
| lora_list List LoRA adapters loaded on a running vLLM endpoint. Shows base models and hot-loaded fine-tuned adapters. |
input · no output |
| lora_remove Hot-unload a LoRA adapter from a running vLLM endpoint. Frees GPU memory for other adapters. |
input · no output |
| manage_instance Manage GPU instances (stop/start/terminate) |
input · no output |
| manifests List cached manifests and versions for jobs. |
input · no output |
| migrate_compare added Compare migration target providers by cost/compatibility. |
input · no output |
| migrate_execute added Execute a previously-planned workload migration. |
input · no output |
| migrate_plan added Plan a workload migration between providers with cost analysis. |
input · no output |
| migrate_rollback added Roll back a migration to the source provider. |
input · no output |
| migrate_status added Show status of running/completed migrations. |
input · no output |
| ml_vllm_lora_link Load the active registry version of an adapter onto a vLLM server. |
input · no output |
| ml_vllm_lora_list List LoRA adapters currently loaded on a vLLM server. |
input · no output |
| ml_vllm_lora_load Hot-load a LoRA adapter onto a running vLLM server. |
input · no output |
| ml_vllm_lora_sync Synchronize an adapter from the registry across multiple vLLM replicas. |
input · no output |
| ml_vllm_lora_unload Hot-unload a LoRA adapter from a running vLLM server. |
input · no output |
| mlflow_list_experiments List MLflow experiments on the configured tracking server. |
input · no output |
| mlflow_log_run Log a Terradev training run to MLflow with auto-injected GPU type, provider, cost/hr, and duration as params. |
input · no output |
| mlflow_register_model Register a trained model in the MLflow model registry with Terradev provenance tags. |
input · no output |
| moe_deploy Deploy Mixture-of-Experts models with production-ready cluster templates. Auto-applies vLLM cost optimizations (KV cache offloading for up to 9x throughput, MTP speculative decodin |
input · no output |
| ollama_chat Chat with an Ollama model using the chat/completions API. |
input · no output |
| ollama_generate Generate text using an Ollama model (non-chat completions). |
input · no output |
| ollama_list List models available on an Ollama server. |
input · no output |
| ollama_model_info Get detailed information about an Ollama model (parameters, template, license). |
input · no output |
| ollama_ps List currently running Ollama models. |
input · no output |
| ollama_pull Pull a model to an Ollama server on a remote instance. |
input · no output |
| optimize Find cheaper alternatives for running instances |
input · no output |
| orchestrator_evict Evict a model from GPU memory. |
input · no output |
| orchestrator_infer Test inference with a model via the orchestrator. |
input · no output |
| orchestrator_load Load a model into GPU memory. |
input · no output |
| orchestrator_register Register a model with the orchestrator. |
input · no output |
| orchestrator_start Start the model orchestrator for multi-model GPU sharing with eviction policies. |
input · no output |
| orchestrator_status Get orchestrator and model status including GPU memory utilization. |
input · no output |
| phoenix_k8s Generate Kubernetes deployment manifest for self-hosted Arize Phoenix server. |
input · no output |
| phoenix_otel_env Generate OpenTelemetry environment variables for instrumenting serving pods with Phoenix tracing. |
input · no output |
| phoenix_projects List Phoenix projects (trace namespaces). |
input · no output |
| phoenix_snippet Generate Python instrumentation snippet for adding Phoenix tracing to LLM applications. |
input · no output |
| phoenix_spans List recent spans for a Phoenix project. Supports SpanQuery DSL filters like "span_kind == 'RETRIEVER'" or "status_code == 'ERROR'". |
input · no output |
| phoenix_test Test connection to Arize Phoenix server. Returns collector endpoint and project count. |
input · no output |
| phoenix_trace View full execution tree for a specific trace ID. Shows span hierarchy, latencies, and token counts. |
input · no output |
| preflight Pre-training validation: GPU availability, NCCL, RDMA, drivers across all nodes. |
input · no output |
| preflight_gpu_check GPU-specific preflight validation: NVIDIA drivers, CUDA version, GPU count, NCCL, NVLink topology, NCU stall-signature profiling, and adversarial config verification (V1-V3). |
input · no output |
| preflight_network_check Network-specific preflight validation: RDMA availability, InfiniBand status, inter-node bandwidth, latency matrix, firewall rules. |
input · no output |
| preflight_provision added Run preflight validation before provisioning GPU instances. |
input · no output |
| preflight_report Generate full preflight validation report with pass/warn/fail per check. Covers GPU drivers, CUDA, NCCL, RDMA, network, disk, and Docker. |
input · no output |
GPU price intelligence with quantitative analytics. Computes delta (rate of change), gamma (acceleration), and annualized realized volatility on GPU spot/on-demand prices across 21 |
— |
Spot instance risk assessment per provider. Returns interruption probability, mean time to interruption, and recommended mitigation. |
— |
Get GPU price trend analysis with delta (rate of change), gamma (acceleration), and annualized volatility. Identifies cheapest time windows. |
— |
Provision GPU instances for optimal parallel efficiency |
— |
List all Qdrant vector collections with their point counts and configurations. |
— |
Count points (vectors) in a Qdrant collection. |
— |
Create a Qdrant vector collection. Auto-configures vector dimensions from embedding model name. |
— |
Get detailed info and stats for a Qdrant collection. |
— |
Generate Kubernetes StatefulSet manifest for self-hosted Qdrant vector database. |
— |
Test connection to Qdrant vector database. Returns cluster info and collection count. |
— |
Execute a SELECT query on a database connection. Returns query results as a list of dictionaries. |
— |
Generate a Ray Serve LLM disaggregated Prefill/Decode deployment. Splits inference into compute-bound prefill and memory-bound decode phases with KV cache transfer via NIXL. |
— |
List all running Ray jobs and tasks. |
— |
Compute optimal TP/DP/EP parallelism strategy for a given MoE model and GPU count. Returns recommended configuration with rationale. |
— |
Start a Ray cluster (head node or worker). For distributed ML training and inference. |
— |
Get Ray cluster status including node count, resources, memory, and running jobs. |
— |
Stop the Ray cluster on the current node. |
— |
Submit a job script to the Ray cluster for distributed execution. |
— |
Generate a Ray Serve LLM Wide-EP (Expert Parallel) deployment for MoE models. Returns Python script and config for distributed MoE serving with EPLB and DeepEP. |
— |
Explicit versioned rollback. Format: job@version (e.g., llama3@v3). |
— |
Run a declarative YAML workflow that chains multiple Terradev commands (provision → preflight → train → monitor → checkpoint). Returns step-by-step execution status with cost estim |
— |
Print environment-style export lines for a provider. By default values are masked. |
— |
Retrieve a stored secret. By default the value is masked. |
— |
List stored provider and key names. Values are never shown. |
— |
Remove a provider or a single key from the secret store. |
— |
Run a shell command with secrets injected into the environment. |
— |
Verify it yourself
npx teppi-check https://terradev-mcp.terradev.cloud/ssecurl -s https://api.teppi.xyz/v1/trust/mcp/mcs_01M22BJ00XR3B3P37ZFB8PAGQD