{"name":"run.fitllm/fitllm","slug":"fitllm","title":"FitLLM","description":"Will this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only.","url":"https://mcp.market/server/fitllm","rating":null,"grade":"A","score":89,"certified":false,"status":"active","category":"ai","tags":["ai"],"presence":{"score":32,"stars":8,"forks":2,"downloads_week":null,"last_push_at":"2026-09-14T08:38:18.000Z","license":"MIT"},"uptime":{"percent":100,"checks":5,"ok":5,"last_checked_at":"2026-09-20T19:01:17.098Z","last_ok_at":"2026-09-20T19:01:17.098Z","latency_ms":331},"claimed":false,"transport":"remote","callable_via_gateway":true,"default_price_micros":0,"repository":"https://github.com/click6067-ship-it/fitllm-engine","website":"https://fitllm.run","version":"1.1.0","remotes":[{"type":"streamable-http","url":"https://fitllm.run/api/mcp"}],"packages":[],"tools":[{"name":"check_llm_fit","description":"Check whether a specific local LLM fits in the memory of a specific GPU or Apple Silicon Mac. Returns fits/tight/won't-fit verdict with the memory breakdown (weights, KV cache, linear-attention state when present, runtime overhead, reserve), max context, and a concrete fix if it doesn't fit. Use this whenever a user asks anything like \"can I run <model> on my <GPU/Mac>?\", \"will <model> fit in <N>GB?\", or \"what do I need to run <model>?\". Estimates using curated, config-derived architecture fields (MLA, sliding-window, hybrid attention, MoE modeled).","write_action":false,"price_micros":0,"input_schema":{"type":"object","properties":{"model":{"type":"string","description":"LLM name, fuzzy — e.g. \"GLM-4.7-Flash\", \"gpt-oss-20b\", \"gemma 31b\""},"gpu":{"type":"string","description":"GPU name, fuzzy — e.g. \"RTX 4090\", \"RX 7900 XTX\", \"A100 80GB\". Multi-GPU rigs: join with + — e.g. \"RTX 5090 + RTX 3090\" (VRAM pools across cards). Provide gpu OR mac_ram_gb."},"gpu_count":{"type":"integer","minimum":1,"maximum":8,"description":"Number of identical copies of the gpu (e.g. gpu=\"RTX 3090\", gpu_count=2 for a 2×3090 rig). Default 1."},"mac_ram_gb":{"type":"integer","minimum":8,"maximum":2048,"description":"Apple Silicon unified memory in GB — e.g. 16, 64, 512. Provide gpu OR mac_ram_gb."},"quant":{"type":"string","description":"Weight quantization. GPU: Q4_K_M(default)/Q5_K_M/Q6_K/Q8_0/FP16. Mac: 4/8(default)/16 (bits)."},"context_tokens":{"type":"integer","minimum":1024,"description":"Context length in tokens (default 8192). Alias: ctx (same field as the REST API)."},"ctx":{"type":"integer","minimum":1024,"description":"Alias of context_tokens — accepted because the REST API uses this name. Do not pass both with different values."},"kv_bits":{"type":"number","enum":[16,8,4],"description":"KV-cache quantization bits (default 16 = F16)"}},"required":["model"],"additionalProperties":false,"$schema":"http://json-schema.org/draft-07/schema#"}},{"name":"list_supported","description":"List the built-in model names and hardware names this fit-checker knows (for mapping user wording to exact names). Standard text-only HuggingFace transformer configs can also be checked via fitllm.run; unsupported architectures are rejected.","write_action":false,"price_micros":0,"input_schema":{"type":"object","properties":{},"$schema":"http://json-schema.org/draft-07/schema#"}},{"name":"what_fits_on_hardware","description":"Rank which popular local LLMs fit on a given GPU or Apple Silicon Mac (at ~4-bit quantization, 8K context) — models that fit come first, biggest first, with max context each. Use when a user asks \"what can I run on my <GPU/Mac/N GB>?\", \"best local model for my machine?\", or gives hardware without naming a model.","write_action":false,"price_micros":0,"input_schema":{"type":"object","properties":{"gpu":{"type":"string","description":"GPU name, fuzzy. Multi-GPU rigs: join with + (e.g. \"RTX 5090 + RTX 3090\"). Provide gpu OR mac_ram_gb."},"gpu_count":{"type":"integer","minimum":1,"maximum":8,"description":"Number of identical copies of the gpu. Default 1."},"mac_ram_gb":{"type":"integer","minimum":8,"maximum":2048,"description":"Apple Silicon unified memory GB. Provide gpu OR mac_ram_gb."}},"additionalProperties":false,"$schema":"http://json-schema.org/draft-07/schema#"}}],"scan":{"score":89,"grade":"A","scanned_at":"2026-09-20T16:20:02.063Z","report":{"scannerVersion":"0.1.9","scannedAt":"2026-09-20T16:20:02.023Z","components":{"code":{"score":-1,"max":25,"notes":["remote-only server, no package to scan"]},"reliability":{"score":20,"max":20,"notes":["remote reachable in 447ms"]},"poisoning":{"score":15,"max":15,"notes":["3 tool descriptions checked"]},"auth":{"score":10,"max":15,"notes":["open endpoint, read-only tools"]},"maintenance":{"score":15,"max":15,"notes":["last push 6 days ago"]},"identity":{"score":7,"max":10,"notes":["namespace and repository owner differ","GitHub account older than a year","website matches verified namespace"]}},"findings":[],"inputs":{"probes":[{"url":"https://fitllm.run/api/mcp","reachable":true,"authRequired":false,"latencyMs":447,"serverInfo":{"name":"fitllm","version":"1.1.0"}}],"packages":[],"repo":{"found":true,"owner":"click6067-ship-it","repo":"fitllm-engine","archived":false,"pushedAt":"2026-09-14T08:38:18Z","stars":8,"forks":2,"openIssues":3,"ownerType":"User","ownerAvatarUrl":"https://avatars.githubusercontent.com/u/226941778?v=4","ownerCreatedAt":"2025-08-17T02:29:15Z","license":"MIT"},"icon":{"url":"data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 32 32'%3E%3Crect width='32' height='32' rx='8' fill='%23030712'/%3E%3Ccircle cx='16' cy='16' r='6' fill='%23a78bfa'/%3E%3C/svg%3E","source":"site"},"presence":{"stars":8,"forks":2,"downloadsWeek":null,"license":"MIT","lastPushAt":"2026-09-14T08:38:18.000Z","score":32}}}},"grade_history":[],"reviews":[]}