scaffold-exercises skill
Scaffold a graded problem set with sections, problems, worked solutions, and short "why this matters" explainers across analytical, empirical, and coding types. Use when user says "make a problem set on X", "scaffold exercises for this lecture", "create practice problems", "generate homework with a solution key", "build a graded assignment on topic Y". Emits a clean student set plus a separate solution key — NOT for grading submissions or auto-checking student answers.
Is the scaffold-exercises skill safe?
Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.
No findings.
Install the scaffold-exercises skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git /tmp/claude-code-my-workflow mkdir -p ~/.claude/skills cp -r /tmp/claude-code-my-workflow/.claude/skills/scaffold-exercises ~/.claude/skills/scaffold-exercises
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
/scaffold-exercises — Problem Set Scaffolder
Generate a graded problem set as two files: a clean student set (problems only) and a solution key (worked solutions + a one-line explainer per problem). Pattern imported from mattpocock/skills, adapted for economics teaching — the primary lens is graded coursework that mixes derivation, estimation, and code.
Input: $ARGUMENTS — a topic (e.g., "instrumental variables", "consumer theory", "quantile regression") and optional flags. See Flags.
When to use
- You have a lecture or reading and want a matching assignment with an answer key.
- You want a mix of problem types (derive, estimate, code) at a controlled difficulty, with solutions emitted separately so the student file stays clean.
Do not use this to grade submissions, auto-check answers, or build a timed exam — it scaffolds practice/graded material, not assessment infrastructure.
Problem types
If no dataset is supplied for an empirical problem, generate a small simulated one with a fixed seed (YYYYMMDD) so the answer key is deterministic and reproducible.
Workflow
Phase 0: Set topic, difficulty, counts, types (Pre-Flight)
Read any source material the user points at (lecture .tex/.qmd, a paper, a dataset header) and produce a Pre-Flight Report before generating problems:
## Pre-Flight Report — Problem Set
**Topic:** [topic]
**Source(s) read:** [lecture/paper/dataset — one-line takeaway each]
**Difficulty:** intro | core | advanced
**Counts by type:** analytical=N, empirical=N, coding=N (total = `--count`)
**Dataset:** [provided path | simulated with seed YYYYMMDD | none]
**Learning objectives:** [2-4 bullets the set should exercise]Resolve every flag here (interactive choices are gathered before generation, not mid-run). If the topic is too vague to write objectives, ask one clarifying question and stop. Otherwise proceed.
Phase 1: Generate problems
For each problem, write a number, a section heading, the prompt, and any data/notation it needs. Conventions:
- Motivation before mechanics — one sentence on why the problem is worth solving, matching create-lecture's pedagogy.
- Notation reuse — match symbols to the source lecture; never introduce a clashing symbol for an already-defined object.
- Difficulty calibration — intro checks one concept; core chains 2-3 steps; advanced requires a non-obvious insight or identification argument.
- Self-contained — each problem states its own assumptions; no "as in lecture 4" dangling references.
Phase 2: Generate worked solutions + explainers
For every problem, write:
- A worked solution — full derivation, expected estimate, or runnable code (depending on type). Coding solutions must actually run; if Bash + R/Stata are available, execute the snippet and paste real output.
- A "why this matters" explainer — 1-2 sentences linking the answer to the broader concept (the imported pattern's signature: every problem ships with a short rationale, not just a number).
Phase 3: Write student set + solution key
Emit two files (paths configurable; default under qualityreports/teaching/, beside /respond-to-eval's teaching plans. A new top-level directory such as exercises/ fails scripts/check-repo-hygiene.py once committed, unless it is added to that script's ROOTALLOW_DIRS):
- qualityreports/teaching/problems.md — the student set: sections, problems, any data, NO answers.
- qualityreports/teaching/solutions.md — the solution key: each problem restated, its worked solution, and its explainer.
The split is load-bearing: never leak a solution into the student file. With --no-solutions, write only the student set and stop.
Output / Report format
Student set:
# Problem Set: [Topic] (Difficulty: core)
## Section 1 — Analytical
**1.** [Motivation sentence.] [Prompt.]
## Section 2 — Empirical
**2.** Using `data/<file>` (vars: ...), [estimate + interpret prompt].
## Section 3 — Coding (R)
**3.** [Implement-X prompt.]Solution key mirrors the numbering, adding ### Solution and > Why this matters: blocks per problem. Close your chat reply with a one-line manifest: files written, problem count by type, and whether code solutions were executed or only drafted.
Exit behavior
- Print the two output paths (absolute), the per-type counts, and the seed if a dataset was simulated.
- If a coding solution could not be executed (no R/Stata, or it errored), flag it as DRAFTED — NOT RUN rather than implying it was verified.
- If any empirical problem references variables not present in the supplied dataset, stop and surface the mismatch instead of inventing columns.
Flags
- --difficulty — intro | core | advanced (default core); calibrates step depth as in Phase 1.
- --count — total number of problems (default 6); split across types per the Pre-Flight counts.
- --types — comma-separated subset of analytical,empirical,coding (default all three).
- --dataset — path to a real dataset for empirical problems; omit to simulate one with a seeded DGP.
- --no-solutions — write only the student set; skip the solution key (Phase 2/3 key file).
Cross-references
- .claude/skills/create-lecture/SKILL.md — build the lecture these exercises practice; shares notation-reuse + motivation-first conventions.
- .claude/skills/data-analysis/SKILL.md — for empirical problems whose reference solution needs a full R estimation pipeline.
- .claude/skills/simulation-study/SKILL.md — when a problem demonstrates an estimator's finite-sample behavior; reuse its seeded-DGP discipline.
- .claude/skills/lit-review/SKILL.md — source advanced problems from current papers on the topic.
- .claude/skills/interview-me/SKILL.md — turn a fuzzy "I want a set on…" into concrete learning objectives first.
- templates/skill-template.md — house style for authoring/extending this skill.
What this skill does NOT do
- Does not grade student submissions or auto-check answers against a key.
- Does not run a timed exam or enforce assessment policy (point weights, rubrics, proctoring).
- Does not invent data — empirical problems use a supplied dataset or an explicitly seeded simulation, never fabricated numbers.
- Does not leak solutions into the student file, and does not deploy/publish anything (no /deploy).
- Does not auto-invoke other skills — it references siblings; it does not call them.
More skills from pedrohcgs/claude-code-my-workflow
- Aadjudicate-reviewTurn an incoming set of findings — from an AI reviewer, a referee report, a code review, a linter, or a second model — into verified fixes, without letting a confident misread damage correct work. Every finding is a CANDIDATE until checked against the actual source. Use whenever you receive review comments, audit findings, or a critique you did not write yourself, especially when the reviewer is a model or when the volume is too large to check by feel.
- Aaudit-reproducibilityEnforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs. Report PASS/FAIL per claim against tolerance thresholds. Use before submission and before releasing a replication package.
- Ablast-radiusBefore and after changing anything shared — a function's return value, a signature, a schema, a label set, a config default, a constant, a file format — find every consumer and actually run them. Catches the change that looks purely additive but silently breaks a contract in a file you never opened. Use when editing shared code, adding a field/column/return element, renaming, changing units or defaults, or touching a pipeline that produces reported numbers.
- Acapture-environmentSnapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt / environment.yml / uv.lock, Stata version + ado package list), records seeds and RNG kind, optionally writes a pinning Dockerfile, and produces a paste-ready "Computational requirements" block. Use when user says "capture the environment", "snapshot my dependencies", "pin the versions", "make a renv.lock / requirements.txt", "make this byte-reproducible", or before releasing a replication package to openICPSR / the AEA Data Editor.
- AchallengeStress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse", "how sensitive is this", "what if I'd used a different measure", "stress-test my estimate", or before a result becomes a headline claim. NOT a reviewer of prose or code — it challenges the CLAIM.
- AcheckpointSave a structured state snapshot before stopping or handing off. Captures the active plan, recent decisions, file pointers (with line numbers), open questions, and the next 1–3 actions into a checkpoint file under `quality_reports/checkpoints/`. Optionally proposes `[LEARN]` entries to add to MEMORY.md. Use when user says "checkpoint", "save state", "snapshot before I stop", "where am I", "wrap up the session for handoff", or before a long break / model switch / collaborator handoff. Companion to (NOT replacement for) the narrative session-log workflow.
- Acoauthor-briefGenerate a co-author / collaborator handoff brief for a multi-author, multi-machine project — summarizing what changed since the last brief (git delta), the current state of each artifact (manuscript, analysis, slides), open questions, how to reproduce locally, and any restricted-data access steps. Use when user says "coauthor brief", "handoff brief", "bring my coauthor up to speed", "what changed since last week", "onboard a collaborator", "write a handoff for [name]", or before sending a co-author the repo. NOT a commit or a checkpoint — it is the cross-machine, cross-person summary `meta-governance.md` only partially covers.
- AcommitCommit the current work — runs the quality, consistency and passport gates, branches off main if needed, stages specific files, and writes a commit whose subject states what is now true. Pushes and opens a pull request only with --pr or when the user asks; never merges — a merge happens only when the user explicitly says to merge. Use ONLY on explicit commit intent — user says "commit", "let's commit this", "open a PR", or prefixes with `/commit`. Do NOT auto-invoke on vague end-of-task phrases ("we're done", "wrap up") — those require explicit confirmation first. Never force-pushes or skips hooks.
- Acompile-latexCompile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex). Use when user says "compile", "build the slides", "rebuild the PDF", "run latex", "render the tex", or asks why a `.tex` file isn't producing a PDF. Operates on `Slides/*.tex`.
- Acompress-sessionDistill the current conversation into a structured note (decisions made, open questions, file pointers with line numbers, next 1–3 actions) and save to `quality_reports/session_logs/` before auto-compression. Differs from `/checkpoint` (explicit stop-point snapshot) and from auto-compaction (which truncates rather than distills). Use when context is approaching auto-compact threshold, when a long pipeline has accumulated many decisions, or when the user says "compress", "distil this session", "before we hit auto-compact", "structured handoff before context resets".
- Acontext-statusShow current context status and session health. Use to check how much context has been used, whether auto-compact is approaching, and what state will be preserved.
- Acreate-lectureCreate a new Beamer lecture `.tex` from source papers and materials, with notation consistency checks and the project's preamble wired in. Use when user says "create a lecture on X", "new lecture from these papers", "start a deck on topic Y", "scaffold a new Beamer file", "build me a lecture from these PDFs". Scaffolds the full deck — NOT for compiling existing `.tex` (use `/compile-latex`).