Mmcp.market

llm-wiki skill

by kevinnft·kevinnft/ai-agent-skills·14 stars·MIT

Karpathy's LLM Wiki: build/query interlinked markdown KB.

A90/100content scan

Is the llm-wiki skill safe?

Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.

  • mediumSKILL.md:473

    Edits shell startup files, cron or launch agents, so something runs again after the skill is done.

    systemctl --user enable --now obsidian-wiki-sync

Install the llm-wiki skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/kevinnft/ai-agent-skills.git /tmp/ai-agent-skills
mkdir -p ~/.claude/skills
cp -r /tmp/ai-agent-skills/skills/research/llm-wiki ~/.claude/skills/llm-wiki
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Karpathy's LLM Wiki

Build and maintain a persistent, compounding knowledge base as interlinked markdown files. Based on Andrej Karpathy's LLM Wiki pattern.

Unlike traditional RAG (which rediscovers knowledge from scratch per query), the wiki compiles knowledge once and keeps it current. Cross-references are already there. Contradictions have already been flagged. Synthesis reflects everything ingested.

Division of labor: The human curates sources and directs analysis. The agent summarizes, cross-references, files, and maintains consistency.

When This Skill Activates

Use this skill when the user:

  • Asks to create, build, or start a wiki or knowledge base
  • Asks to ingest, add, or process a source into their wiki
  • Asks a question and an existing wiki is present at the configured path
  • Asks to lint, audit, or health-check their wiki
  • References their wiki, knowledge base, or "notes" in a research context

Wiki Location

Location: Set via WIKI_PATH environment variable (e.g. in ~/.hermes/.env).

If unset, defaults to ~/wiki.

WIKI="${WIKI_PATH:-$HOME/wiki}"

The wiki is just a directory of markdown files — open it in Obsidian, VS Code, or any editor. No database, no special tooling required.

Architecture: Three Layers

wiki/
├── SCHEMA.md           # Conventions, structure rules, domain config
├── index.md            # Sectioned content catalog with one-line summaries
├── log.md              # Chronological action log (append-only, rotated yearly)
├── raw/                # Layer 1: Immutable source material
│   ├── articles/       # Web articles, clippings
│   ├── papers/         # PDFs, arxiv papers
│   ├── transcripts/    # Meeting notes, interviews
│   └── assets/         # Images, diagrams referenced by sources
├── entities/           # Layer 2: Entity pages (people, orgs, products, models)
├── concepts/           # Layer 2: Concept/topic pages
├── comparisons/        # Layer 2: Side-by-side analyses
└── queries/            # Layer 2: Filed query results worth keeping

Layer 1 — Raw Sources: Immutable. The agent reads but never modifies these. Layer 2 — The Wiki: Agent-owned markdown files. Created, updated, and cross-referenced by the agent. Layer 3 — The Schema: SCHEMA.md defines structure, conventions, and tag taxonomy.

Resuming an Existing Wiki (CRITICAL — do this every session)

When the user has an existing wiki, always orient yourself before doing anything:

① Read SCHEMA.md — understand the domain, conventions, and tag taxonomy. ② Read index.md — learn what pages exist and their summaries. ③ Scan recent log.md — read the last 20-30 entries to understand recent activity.

WIKI="${WIKI_PATH:-$HOME/wiki}"
# Orientation reads at session start
read_file "$WIKI/SCHEMA.md"
read_file "$WIKI/index.md"
read_file "$WIKI/log.md" offset=<last 30 lines>

Only after orientation should you ingest, query, or lint. This prevents:

  • Creating duplicate pages for entities that already exist
  • Missing cross-references to existing content
  • Contradicting the schema's conventions
  • Repeating work already logged

For large wikis (100+ pages), also run a quick search_files for the topic at hand before creating anything new.

Initializing a New Wiki

When the user asks to create or start a wiki:

  1. Determine the wiki path (from $WIKI_PATH env var, or ask the user; default ~/wiki)
  2. Create the directory structure above
  3. Ask the user what domain the wiki covers — be specific
  4. Write SCHEMA.md customized to the domain (see template below)
  5. Write initial index.md with sectioned header
  6. Write initial log.md with creation entry
  7. Confirm the wiki is ready and suggest first sources to ingest

SCHEMA.md Template

Adapt to the user's domain. The schema constrains agent behavior and ensures consistency:

# Wiki Schema

## Domain
[What this wiki covers — e.g., "AI/ML research", "personal health", "startup intelligence"]

## Conventions
- File names: lowercase, hyphens, no spaces (e.g., `transformer-architecture.md`)
- Every wiki page starts with YAML frontmatter (see below)
- Use `[[wikilinks]]` to link between pages (minimum 2 outbound links per page)
- When updating a page, always bump the `updated` date
- Every new page must be added to `index.md` under the correct section
- Every action must be appended to `log.md`
- **Provenance markers:** On pages that synthesize 3+ sources, append `^[raw/articles/source-file.md]`
  at the end of paragraphs whose claims come from a specific source. This lets a reader trace each
  claim back without re-reading the whole raw file. Optional on single-source pages where the
  `sources:` frontmatter is enough.

## Frontmatter

title: Page Title created: YYYY-MM-DD updated: YYYY-MM-DD type: entity | concept | comparison | query | summary tags: [from taxonomy below] sources: [raw/articles/source-name.md] # Optional quality signals: confidence: high | medium | low # how well-supported the claims are contested: true # set when the page has unresolved contradictions contradictions: [other-page-slug] # pages this one conflicts with

`confidence` and `contested` are optional but recommended for opinion-heavy or fast-moving
topics. Lint surfaces `contested: true` and `confidence: low` pages for review so weak claims
don't silently harden into accepted wiki fact.

### raw/ Frontmatter

Raw sources ALSO get a small frontmatter block so re-ingests can detect drift:

source_url: https://example.com/article # original URL, if applicable ingested: YYYY-MM-DD sha256:

The `sha256:` lets a future re-ingest of the same URL skip processing when content is unchanged,
and flag drift when it has changed. Compute over the body only (everything after the closing
`---`), not the frontmatter itself.

## Tag Taxonomy
[Define 10-20 top-level tags for the domain. Add new tags here BEFORE using them.]

Example for AI/ML:
- Models: model, architecture, benchmark, training
- People/Orgs: person, company, lab, open-source
- Techniques: optimization, fine-tuning, inference, alignment, data
- Meta: comparison, timeline, controversy, prediction

Rule: every tag on a page must appear in this taxonomy. If a new tag is needed,
add it here first, then use it. This prevents tag sprawl.

## Page Thresholds
- **Create a page** when an entity/concept appears in 2+ sources OR is central to one source
- **Add to existing page** when a source mentions something already covered
- **DON'T create a page** for passing mentions, minor details, or things outside the domain
- **Split a page** when it exceeds ~200 lines — break into sub-topics with cross-links
- **Archive a page** when its content is fully superseded — move to `_archive/`, remove from index

## Entity Pages
One page 

index.md Template

The index is sectioned by type. Each entry is one line: wikilink + summary.

# Wiki Index

> Content catalog. Every wiki page listed under its type with a one-line summary.
> Read this first to find relevant pages for any query.
> Last updated: YYYY-MM-DD | Total pages: N

## Entities
<!-- Alphabetical within section -->

## Concepts

## Comparisons

## Queries

Scaling rule: When any section exceeds 50 entries, split it into sub-sections by first letter or sub-domain. When the index exceeds 200 entries total, create a _meta/topic-map.md that groups pages by theme for faster navigation.

log.md Template

# Wiki Log

> Chronological record of all wiki actions. Append-only.
> Format: `## [YYYY-MM-DD] action | subject`
> Actions: ingest, update, query, lint, create, archive, delete
> When this file exceeds 500 entries, rotate: rename to log-YYYY.md, start fresh.

## [YYYY-MM-DD] create | Wiki initialized
- Domain: [domain]
- Structure created with SCHEMA.md, index.md, log.md

Core Operations

1. Ingest

When the user provides a source (URL, file, paste), integrate it into the wiki:

① Capture the raw source:

On re-ingest of the same URL: recompute the sha256, compare to the stored value — skip if identical, flag drift and update if different. This is cheap enough to do on every re-ingest and catches silent source changes.

  • URL → use web_extract to get markdown, save to raw/articles/
  • PDF → use web_extract (handles PDFs), save to raw/papers/
  • Pasted text → save to appropriate raw/ subdirectory
  • Name the file descriptively: raw/articles/karpathy-llm-wiki-2026.md
  • Add raw frontmatter (source_url, ingested, sha256 of the body).

② Discuss takeaways with the user — what's interesting, what matters for the domain. (Skip this in automated/cron contexts — proceed directly.)

③ Check what already exists — search index.md and use search_files to find existing pages for mentioned entities/concepts. This is the difference between a growing wiki and a pile of duplicates.

④ Write or update wiki pages:

in SCHEMA.md (2+ source mentions, or central to one source)

  • New entities/concepts: Create pages only if they meet the Page Thresholds

When new info contradicts existing content, follow the Update Policy.

  • Existing pages: Add new information, update facts, bump updated date.

pages via [[wikilinks]]. Check that existing pages link back.

  • Cross-reference: Every new or updated page must link to at least 2 other

markers to paragraphs whose claims trace to a specific source.

  • Tags: Only use tags from the taxonomy in SCHEMA.md
  • Provenance: On pages synthesizing 3+ sources, append ^[raw/articles/source.md]

confidence: medium or low in frontmatter. Don't mark high unless the claim is well-supported across multiple sources.

  • Confidence: For opinion-heavy, fast-moving, or single-source claims, set

⑤ Update navigation:

  • Add new pages to index.md under the correct section, alphabetically
  • Update the "Total pages" count and "Last updated" date in index header
  • Append to log.md: ## [YYYY-MM-DD] ingest | Source Title
  • List every file created or updated in the log entry

⑥ Report what changed — list every file created or updated to the user.

More skills from kevinnft/ai-agent-skills

  • Aaddyosmani-tddDrives development with tests. Use when implementing any logic, fixing any bug, or changing any behavior. Use when you need to prove that code works, when a bug report arrives, or when you're about to modify existing functionality.
  • AairtableAirtable REST API via curl. Records CRUD, filters, upserts.
  • Aapi-and-interface-designGuides stable API and interface design. Use when designing APIs, module boundaries, or any public interface. Use when creating REST or GraphQL endpoints, defining type contracts between modules, or establishing boundaries between frontend and backend.
  • Aapi-monitoring-botsBuild monitoring bots that poll APIs and send notifications on state changes (new listings, price alerts, status updates)
  • Aapple-notesManage Apple Notes via memo CLI: create, search, edit.
  • Aapple-remindersApple Reminders via remindctl: add, list, complete.
  • Aarchitecture-diagramDark-themed SVG architecture/cloud/infra diagrams as HTML.
  • AarxivSearch arXiv papers by keyword, author, category, or ID.
  • Aascii-artASCII art: pyfiglet, cowsay, boxes, image-to-ascii.
  • Aascii-videoASCII video: convert video/audio to colored ASCII MP4/GIF.
  • AaudiocraftAudioCraft: MusicGen text-to-music, AudioGen text-to-sound.
  • CaxolotlAxolotl: YAML LLM fine-tuning (LoRA, DPO, GRPO).

All agent skills → · MCP servers