respond-to-eval skill
Turn student course evaluations (free-text + numeric) into an actionable teaching-improvement plan — the teaching analogue of /respond-to-referees. Clusters comments into themes, separates signal from noise, classifies each theme Keep / Change / Investigate / Out-of-scope, and drafts concrete changes mapped to the syllabus and slide decks. Use when user says "respond to my evals", "what do these course evaluations tell me", "turn my teaching feedback into a plan", or after a semester's evals arrive.
Is the respond-to-eval skill safe?
Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.
No findings.
Install the respond-to-eval skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git /tmp/claude-code-my-workflow mkdir -p ~/.claude/skills cp -r /tmp/claude-code-my-workflow/.claude/skills/respond-to-eval ~/.claude/skills/respond-to-eval
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Respond to Evaluations
Convert a semester's course evaluations into a defensible teaching-improvement plan. Cluster free-text comments into themes, weight each theme by how many independent students raised it, classify what to do about it, and draft specific changes pointed at the syllabus and deck — so next semester's revision is a checklist, not a vibe.
Posture (echoes /respond-to-referees): one angry comment is not a trend, and a comment you disagree with is a signal to investigate, not a license to ignore or to auto-act. A single student's frustration may be the only one willing to say what twenty felt — frequency weights the theme, it does not gate it. Ground truth here is a process: the plan records why a theme was kept or changed, so the reasoning survives to the next round.
When to use
- End of term, when numeric scores + open-text comments land and you want a revision plan, not a mood.
- Assembling a teaching dossier / tenure file where you must show you acted on feedback.
- Mid-stream (early-semester feedback) to course-correct before the term ends.
Not for: writing the syllabus from scratch (use /syllabus), or reviewing one deck's pedagogy (use /pedagogy-review).
Inputs
- $0 — the evaluation file(s): a CSV/TSV export, .txt/.md of pasted comments, or a .pdf/.docx report.
- $1 (optional) — the prior improvement plan, so this round is a diff (did last term's changes land?).
If extraction fails or a tool is missing, ask for a plain-text export and stop. A scanned or partly scanned PDF extracts with exit 0 and blank pages, so also compare pdfinfo "$0" | grep Pages with the pages that returned text (awk 'BEGIN{RS="\f"} NF{n++} END{print n+0}' "$TMP"); read any blank pages directly with Read, or ask for a text version, and if you go on without them, say which pages were not read. When you are done, delete the extracted copy (rm -rf "$(dirname "$TMP")"): it is a plaintext copy of the document.
Phases
Phase 0: Load evals + prior plan (Pre-Flight)
Read the eval file(s) and the prior plan (if given). Produce a short Pre-Flight block before clustering:
## Pre-Flight Report
**Evals loaded:** N responses (M with free-text), instrument: [name/term]
**Numeric items:** [list each scale item + mean, and the institution/department mean if present]
**Prior plan:** [path, or "none — first round"] — changes promised last term: [bullet list]
**Course artifacts in scope:** [syllabus path] · [deck(s) under Slides/ or Quarto/]Numbers anchor the read but do not override text: a 4.2/5 with ten "I was lost by week 6" comments is a problem the mean is hiding.
Phase 1: Theme-cluster + weight by frequency
- Split free-text into atomic comments (one student may raise several themes; one comment may belong to several themes).
- Cluster into themes (e.g., pacing, problem-set difficulty, grading clarity, real-world relevance, office hours, prerequisite gaps). Name each theme in the instructor's words, not the student's.
- For each theme record: mention count (distinct students), representative verbatim quote (~25 words, anonymized — strip names/identifying detail), valence (positive / negative / mixed), and numeric corroboration (which scale item, if any, moves with it).
- Signal vs noise: a theme with < --min-mentions (default 2) distinct students is tagged low-frequency, not dropped — it carries to Phase 2 for a Keep/Investigate call. Frequency weights; it never silences.
Phase 2: Classify + propose changes
Assign each theme exactly one label (the teaching analogue of /respond-to-referees' coverage matrix):
For each Change, write a concrete revision mapped to a target: syllabus §X / Slides/LectureNN.tex slide K / a new worked example / an assessment reweighting — the same point-to-the-location discipline /respond-to-referees uses for "we added X on page Y". For Investigate, name the evidence you'll collect and the decision rule. For Out-of-scope, write the one-sentence rationale you'd stand behind in a dossier.
Disagreement is explicit and reasoned: "Students asked to drop proofs; retained because the course's stated objective is derivation fluency — added two scaffolded worked examples (LectureNN slide K) to ease the on-ramp instead" is a Keep-with-mitigation, not an Out-of-scope dismissal.
Phase 3: Save the improvement plan
Write the plan to qualityreports/teaching/YYYY-MM-DD[course]_improvement-plan.md. Structure:
- Header — course, term, instrument, response rate, numeric summary vs benchmark.
- Prior-plan retrospective (if $1 given) — for each change promised last term: Landed / Partial / Not done, with the evidence from this term's evals.
- Theme matrix — one row per theme: theme · mentions · valence · numeric corroboration · classification · target (syllabus §/deck slide) · representative quote.
- Change list — the concrete edits, ordered by mention count then severity, each pointing at a syllabus section or deck/slide.
- Investigate list — open questions + the evidence to collect next.
The plan is a deliverable, not a transient report, so it lives under quality_reports/teaching/ and feeds next term's $1.
Phase 3.5: Post-Flight Verification (quotes + targets)
The plan's hallucination-prone content is (a) verbatim quotes attributed to students and (b) "edit syllabus §X / LectureNN slide K" targets that must actually exist. Run the fresh-context verifier protocol in .claude/rules/post-flight-verification.md: spawn claim-verifier (fresh context, never a conversation fork) with the quotes + the eval source and the edit-targets + the syllabus/deck paths. Reconcile — a quote that isn't in the source, or a "slide K" that doesn't exist, is corrected or dropped before the plan is final. Opt-out: --no-verify (not recommended).
Output / Report
After writing the plan, surface this in your final chat message (not inside the plan file):
## Teaching-improvement summary — [course], [term]
Themes: K total — C Change · I Investigate · P Keep · O Out-of-scope
Top 3 changes (by mentions): 1) … 2) … 3) …
Open investigations: …
Prior plan: x of y promised changes landed.If all themes are classified and every Change names a target, say All themes classified; every Change mapped to a syllabus or deck target.
Exit behavior
- A theme with no classification halts the report — there are no orphans, exactly as /respond-to-referees admits no unclassified concern.
- Read-only on the syllabus and decks: this skill plans edits and writes the plan file; it does not edit teaching materials. Apply changes deliberately afterward (with /create-lecture or direct edits).
- Numbers never auto-override text and text never auto-overrides numbers; conflicts become Investigate, not a silent winner.
Flags
- --min-mentions — distinct-student threshold below which a theme is tagged low-frequency (default 2). Lowering it surfaces more singletons; it never drops them.
- --no-verify — skip Phase 3.5 Post-Flight Verification of quotes and edit-targets. Not recommended for a dossier-bound plan.
Cross-references
- .claude/skills/respond-to-referees/SKILL.md — the research analogue; this skill borrows its map-classify-respond shape and "signal to investigate, not auto-act" posture.
- .claude/skills/pedagogy-review/SKILL.md — once a Change targets a specific deck, run pedagogy-review on it before re-teaching.
- .claude/skills/create-lecture/SKILL.md — to execute deck-level changes the plan proposes.
- .claude/rules/post-flight-verification.md — the fresh-context verifier protocol Phase 3.5 reuses.
- templates/skill-template.md — house style for skills.
What this skill does NOT do
- It does not edit the syllabus or any deck — it produces a plan; you (or /create-lecture) apply it.
- It does not compute new numeric scores or re-weight the instrument; it reads the institution's numbers as given.
- It does not identify students or attempt to de-anonymize comments — quotes are stripped of identifying detail.
- It does not auto-act on disagreement or on a single comment; both route to Investigate or a documented rationale.
More skills from pedrohcgs/claude-code-my-workflow
- Aadjudicate-reviewTurn an incoming set of findings — from an AI reviewer, a referee report, a code review, a linter, or a second model — into verified fixes, without letting a confident misread damage correct work. Every finding is a CANDIDATE until checked against the actual source. Use whenever you receive review comments, audit findings, or a critique you did not write yourself, especially when the reviewer is a model or when the volume is too large to check by feel.
- Aaudit-reproducibilityEnforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs. Report PASS/FAIL per claim against tolerance thresholds. Use before submission and before releasing a replication package.
- Ablast-radiusBefore and after changing anything shared — a function's return value, a signature, a schema, a label set, a config default, a constant, a file format — find every consumer and actually run them. Catches the change that looks purely additive but silently breaks a contract in a file you never opened. Use when editing shared code, adding a field/column/return element, renaming, changing units or defaults, or touching a pipeline that produces reported numbers.
- Acapture-environmentSnapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt / environment.yml / uv.lock, Stata version + ado package list), records seeds and RNG kind, optionally writes a pinning Dockerfile, and produces a paste-ready "Computational requirements" block. Use when user says "capture the environment", "snapshot my dependencies", "pin the versions", "make a renv.lock / requirements.txt", "make this byte-reproducible", or before releasing a replication package to openICPSR / the AEA Data Editor.
- AchallengeStress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse", "how sensitive is this", "what if I'd used a different measure", "stress-test my estimate", or before a result becomes a headline claim. NOT a reviewer of prose or code — it challenges the CLAIM.
- AcheckpointSave a structured state snapshot before stopping or handing off. Captures the active plan, recent decisions, file pointers (with line numbers), open questions, and the next 1–3 actions into a checkpoint file under `quality_reports/checkpoints/`. Optionally proposes `[LEARN]` entries to add to MEMORY.md. Use when user says "checkpoint", "save state", "snapshot before I stop", "where am I", "wrap up the session for handoff", or before a long break / model switch / collaborator handoff. Companion to (NOT replacement for) the narrative session-log workflow.
- Acoauthor-briefGenerate a co-author / collaborator handoff brief for a multi-author, multi-machine project — summarizing what changed since the last brief (git delta), the current state of each artifact (manuscript, analysis, slides), open questions, how to reproduce locally, and any restricted-data access steps. Use when user says "coauthor brief", "handoff brief", "bring my coauthor up to speed", "what changed since last week", "onboard a collaborator", "write a handoff for [name]", or before sending a co-author the repo. NOT a commit or a checkpoint — it is the cross-machine, cross-person summary `meta-governance.md` only partially covers.
- AcommitCommit the current work — runs the quality, consistency and passport gates, branches off main if needed, stages specific files, and writes a commit whose subject states what is now true. Pushes and opens a pull request only with --pr or when the user asks; never merges — a merge happens only when the user explicitly says to merge. Use ONLY on explicit commit intent — user says "commit", "let's commit this", "open a PR", or prefixes with `/commit`. Do NOT auto-invoke on vague end-of-task phrases ("we're done", "wrap up") — those require explicit confirmation first. Never force-pushes or skips hooks.
- Acompile-latexCompile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex). Use when user says "compile", "build the slides", "rebuild the PDF", "run latex", "render the tex", or asks why a `.tex` file isn't producing a PDF. Operates on `Slides/*.tex`.
- Acompress-sessionDistill the current conversation into a structured note (decisions made, open questions, file pointers with line numbers, next 1–3 actions) and save to `quality_reports/session_logs/` before auto-compression. Differs from `/checkpoint` (explicit stop-point snapshot) and from auto-compaction (which truncates rather than distills). Use when context is approaching auto-compact threshold, when a long pipeline has accumulated many decisions, or when the user says "compress", "distil this session", "before we hit auto-compact", "structured handoff before context resets".
- Acontext-statusShow current context status and session health. Use to check how much context has been used, whether auto-compact is approaching, and what state will be preserved.
- Acreate-lectureCreate a new Beamer lecture `.tex` from source papers and materials, with notation consistency checks and the project's preamble wired in. Use when user says "create a lecture on X", "new lecture from these papers", "start a deck on topic Y", "scaffold a new Beamer file", "build me a lecture from these PDFs". Scaffolds the full deck — NOT for compiling existing `.tex` (use `/compile-latex`).