MCP serverio.nodegrove/vram-mcp
Can this LLM run on my GPU?
Read more
Can this LLM run on my GPU? VRAM, speed ceiling and what fits instead, for any model and GPU.Overview
Score?
UNRATED 0.143
of what a free look can see, on 1 looks
Looks
1
last 10 hr ago
Tools
6
More info
URL
mcp.nodegrove.io/mcp
streamable-http
Says it is
nodegrove-vram 1.0.0
protocol 2025-06-18
In the record since
10 hr ago
Among servers18,413 with a card
0median 0.606 · this server 0.143 · highest on record 0.8561
Toolsfrom sha256:10d0b90738…6f7b0a
| Tool | Schema |
|---|---|
| can_i_run Can this GPU run this open-weight LLM? Returns fits, tight or no, the memory split (weights, KV cache, overhead), a decode-speed ceiling, the longest context that fits and, on a no |
input · no output |
| estimate_from_hf_repo Reads any Hugging Face model repo's config.json and parameter count and estimates its memory: the attention layout found (standard, sliding-window, hybrid or latent), how much each |
input · no output |
| estimate_vram How much memory an LLM needs: weights + KV cache + overhead at each quantisation (or one), at a given context, and the smallest common card class that holds each. Model: a name or |
input · no output |
| list_gpus The GPUs and machines nodegrove.io covers: memory, the memory a runtime can use and bandwidth, from the makers' specs, with each one's page. |
input · no output |
| list_models The open-weight LLMs nodegrove.io has verified against their config.json (data version 2026-09-25): id, size, attention design, native context, licence, memory at Q4 with 8k contex |
input · no output |
| what_fits Which open-weight LLMs fit this GPU: every model in list_models checked at one quantisation and context, with a recommended everyday model (the biggest class that fits with room fo |
input · no output |
Verify it yourself
npx teppi-check https://mcp.nodegrove.io/mcpcurl -s https://api.teppi.xyz/v1/trust/mcp/mcs_01M42QZS8NKKFRPBGYHXD1WBFW