llama-cpp skill
llama.cpp local GGUF inference + HF Hub model discovery.
Is the llama-cpp skill safe?
Clean: nothing in its files matched our rules. We read 7 files in the folder on 2026-09-28.
No findings.
Install the llama-cpp skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/kevinnft/ai-agent-skills.git /tmp/ai-agent-skills mkdir -p ~/.claude/skills cp -r /tmp/ai-agent-skills/skills/mlops/inference/llama-cpp ~/.claude/skills/llama-cpp
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
llama.cpp + GGUF
Use this skill for local GGUF inference, quant selection, or Hugging Face repo discovery for llama.cpp.
When to use
- Run local models on CPU, Apple Silicon, CUDA, ROCm, or Intel GPUs
- Find the right GGUF for a specific Hugging Face repo
- Build a llama-server or llama-cli command from the Hub
- Search the Hub for models that already support llama.cpp
- Enumerate available .gguf files and sizes for a repo
- Decide between Q4/Q5/Q6/IQ variants for the user's RAM or VRAM
Model Discovery workflow
Prefer URL workflows before asking for hf, Python, or custom scripts.
- Search for candidate repos on the Hub:
- Base: https://huggingface.co/models?apps=llama.cpp&sort=trending
- Add search= for a model family
- Add num_parameters=min:0,max:24B or similar when the user has size constraints
- Open the repo with the llama.cpp local-app view:
- https://huggingface.co/?local-app=llama.cpp
- Treat the local-app snippet as the source of truth when it is visible:
- copy the exact llama-server or llama-cli command
- report the recommended quant exactly as HF shows it
- Read the same ?local-app=llama.cpp URL as page text or HTML and extract the section under Hardware compatibility:
- prefer its exact quant labels and sizes over generic tables
- keep repo-specific labels such as UD-Q4KM or IQ4NLXL
- if that section is not visible in the fetched page source, say so and fall back to the tree API plus generic quant guidance
- Query the tree API to confirm what actually exists:
- https://huggingface.co/api/models//tree/main?recursive=true
- keep entries where type is file and path ends with .gguf
- use path and size as the source of truth for filenames and byte sizes
- separate quantized checkpoints from mmproj-*.gguf projector files and BF16/ shard files
- use https://huggingface.co//tree/main only as a human fallback
- If the local-app snippet is not text-visible, reconstruct the command from the repo plus the chosen quant:
- shorthand quant selection: llama-server -hf :
- exact-file fallback: llama-server --hf-repo --hf-file
- Only suggest conversion from Transformers weights if the repo does not already expose GGUF files.
Quick start
Install llama.cpp
# macOS / Linux (simplest)
brew install llama.cppwinget install llama.cppgit clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build
cmake --build build --config ReleaseRun directly from the Hugging Face Hub
llama-cli -hf bartowski/Llama-3.2-3B-Instruct-GGUF:Q8_0llama-server -hf bartowski/Llama-3.2-3B-Instruct-GGUF:Q8_0Run an exact GGUF file from the Hub
Use this when the tree API shows custom file naming or the exact HF snippet is missing.
llama-server \
--hf-repo microsoft/Phi-3-mini-4k-instruct-gguf \
--hf-file Phi-3-mini-4k-instruct-q4.gguf \
-c 4096OpenAI-compatible server check
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [
{"role": "user", "content": "Write a limerick about Python exceptions"}
]
}'Python bindings (llama-cpp-python)
pip install llama-cpp-python (CUDA: CMAKEARGS="-DGGMLCUDA=on" pip install llama-cpp-python --force-reinstall --no-cache-dir; Metal: CMAKEARGS="-DGGMLMETAL=on" ...).
Basic generation
from llama_cpp import Llama
llm = Llama(
model_path="./model-q4_k_m.gguf",
n_ctx=4096,
n_gpu_layers=35, # 0 for CPU, 99 to offload everything
n_threads=8,
)
out = llm("What is machine learning?", max_tokens=256, temperature=0.7)
print(out["choices"][0]["text"])Chat + streaming
llm = Llama(
model_path="./model-q4_k_m.gguf",
n_ctx=4096,
n_gpu_layers=35,
chat_format="llama-3", # or "chatml", "mistral", etc.
)
resp = llm.create_chat_completion(
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is Python?"},
],
max_tokens=256,
)
print(resp["choices"][0]["message"]["content"])
# Streaming
for chunk in llm("Explain quantum computing:", max_tokens=256, stream=True):
print(chunk["choices"][0]["text"], end="", flush=True)Embeddings
llm = Llama(model_path="./model-q4_k_m.gguf", embedding=True, n_gpu_layers=35)
vec = llm.embed("This is a test sentence.")
print(f"Embedding dimension: {len(vec)}")You can also load a GGUF straight from the Hub:
llm = Llama.from_pretrained(
repo_id="bartowski/Llama-3.2-3B-Instruct-GGUF",
filename="*Q4_K_M.gguf",
n_gpu_layers=35,
)Choosing a quant
Use the Hub page first, generic heuristics second.
- Prefer the exact quant that HF marks as compatible for the user's hardware profile.
- For general chat, start with Q4KM.
- For code or technical work, prefer Q5KM or Q6_K if memory allows.
- For very tight RAM budgets, consider Q3KM, IQ variants, or Q2 variants only if the user explicitly prioritizes fit over quality.
- For multimodal repos, mention mmproj-*.gguf separately. The projector is not the main model file.
- Do not normalize repo-native labels. If the page says UD-Q4KM, report UD-Q4KM.
Extracting available GGUFs from a repo
When the user asks what GGUFs exist, return:
- filename
- file size
- quant label
- whether it is a main model or an auxiliary projector
Ignore unless requested:
- README
- BF16 shard files
- imatrix blobs or calibration artifacts
Use the tree API for this step:
- https://huggingface.co/api/models//tree/main?recursive=true
For a repo like unsloth/Qwen3.6-35B-A3B-GGUF, the local-app page can show quant chips such as UD-Q4KM, UD-Q5KM, UD-Q6K, and Q80, while the tree API exposes exact file paths such as Qwen3.6-35B-A3B-UD-Q4KM.gguf and Qwen3.6-35B-A3B-Q8_0.gguf with byte sizes. Use the tree API to turn a quant label into an exact filename.
Search patterns
Use these URL shapes directly:
https://huggingface.co/models?apps=llama.cpp&sort=trending
https://huggingface.co/models?search=<term>&apps=llama.cpp&sort=trending
https://huggingface.co/models?search=<term>&apps=llama.cpp&num_parameters=min:0,max:24B&sort=trending
https://huggingface.co/<repo>?local-app=llama.cpp
https://huggingface.co/api/models/<repo>/tree/main?recursive=true
https://huggingface.co/<repo>/tree/mainOutput format
When answering discovery requests, prefer a compact structured result like:
Repo: <repo>
Recommended quant from HF: <label> (<size>)
llama-server: <command>
Other GGUFs:
- <filename> - <size>
- <filename> - <size>
Source URLs:
- <local-app URL>
- <tree API URL>References
More skills from kevinnft/ai-agent-skills
- Aaddyosmani-tddDrives development with tests. Use when implementing any logic, fixing any bug, or changing any behavior. Use when you need to prove that code works, when a bug report arrives, or when you're about to modify existing functionality.
- AairtableAirtable REST API via curl. Records CRUD, filters, upserts.
- Aapi-and-interface-designGuides stable API and interface design. Use when designing APIs, module boundaries, or any public interface. Use when creating REST or GraphQL endpoints, defining type contracts between modules, or establishing boundaries between frontend and backend.
- Aapi-monitoring-botsBuild monitoring bots that poll APIs and send notifications on state changes (new listings, price alerts, status updates)
- Aapple-notesManage Apple Notes via memo CLI: create, search, edit.
- Aapple-remindersApple Reminders via remindctl: add, list, complete.
- Aarchitecture-diagramDark-themed SVG architecture/cloud/infra diagrams as HTML.
- AarxivSearch arXiv papers by keyword, author, category, or ID.
- Aascii-artASCII art: pyfiglet, cowsay, boxes, image-to-ascii.
- Aascii-videoASCII video: convert video/audio to colored ASCII MP4/GIF.
- AaudiocraftAudioCraft: MusicGen text-to-music, AudioGen text-to-sound.
- CaxolotlAxolotl: YAML LLM fine-tuning (LoRA, DPO, GRPO).