Mmcp.market

remove-ai-slops skill

by code-yeongyu·code-yeongyu/oh-my-openagent·70k stars

Removes AI-generated code smells from branch changes or an explicit file list behind regression tests. Use when the user asks to clean up, deslop, or remove AI-slop patterns from recent changes.

A100/100content scan

Is the remove-ai-slops skill safe?

Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.

No findings.

Install the remove-ai-slops skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/code-yeongyu/oh-my-openagent.git /tmp/oh-my-openagent
mkdir -p ~/.claude/skills
cp -r /tmp/oh-my-openagent/packages/shared-skills/skills/remove-ai-slops ~/.claude/skills/remove-ai-slops
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Remove AI Slops Skill

Inputs

  • Default scope: branch diff vs merge-base main (no arguments needed)
  • Optional scope: explicit file list passed by the caller (e.g., a Ralph workflow's changed-files set)

What this skill does

Cleans AI-generated slop from a bounded set of changed files while strictly preserving behavior. Locks behavior with regression tests first, then runs a categorized multi-pass cleanup, then verifies with quality gates and a critical review. Reverts and direct-edits when verification fails.

The core safety invariant: behavior is locked by green tests before a single line is removed. A checklist alone is not safety; a passing regression test is.

Categories (what counts as slop)

The agent looks for these nine categories. The first three are stylistic, the next three are structural, the next two are about hidden cost, and the last is about behavior coverage.

Stylistic

  1. Obvious comments — comments restating code, trivial docstrings, section dividers, commented-out code, vague TODOs/Notes.
  • KEEP: comments explaining WHY (business logic, edge cases, workarounds), ticket links, regex/algorithm explanations.
  • KEEP: BDD markers (# given, # when, # then, # when/then).
  1. Over-defensive code — null checks for guaranteed values, try/except around code that cannot raise, isinstance checks for statically typed params, default values for required params, backward-compat shims, redundant validation duplicated at multiple layers, broad exception catching (except Exception/except BaseException in Python, empty catch {} or catch (e) { console.error(e) } without narrowing in TypeScript/JavaScript).
  • KEEP: validation at system boundaries (user input, external APIs), I/O error handling, nullable DB fields. Top-level boundary catch-all (CLI main(), HTTP handler) with explicit logging + re-raise is acceptable.
  • REFACTOR: except Exception → catch the specific exception you expect. Empty catch {} → add instanceof narrowing or re-throw. catch (e) { log(e) } → narrow with instanceof, handle known cases, re-throw unknown.
  • PROOF REQUIRED: before deleting any validation or error handling at a trust boundary, Phase 2 must include an adversarial regression (malformed or hostile input) that fails if the guard is removed. No adversarial test → the guard stays. Redundant defense to remove is a duplicate of a check that already runs inside the boundary; a guard with no proof of redundancy is load-bearing.
  1. Excessive complexity — deep nesting (>3 levels), nested ternaries, complex boolean expressions (combine 4+ predicates), long parameter lists (>5 args without a struct/dataclass/object), god functions (>50 lines doing many things), overly clever one-liners that sacrifice readability, if/elif/else chains for type/enum/literal discrimination (must be match/case + assert_never), object used as a type annotation (must be Protocol, TypeVar, or explicit union).
  • KEEP: established complexity patterns in this codebase, performance-critical hot paths that intentionally use a complex idiom. if/else for boolean conditions and range checks (not variant discrimination).
  • REFACTOR: nested if-chains → guard clauses / early returns. Complex ternaries → explicit if/else. isinstance/enum if/elif chains → match/case with assert_never on the wildcard. object annotations → Protocol (structural), TypeVar (generic), or union (known variants).

Structural

  1. Needless abstraction — pass-through wrappers, single-use helpers, speculative indirection ("we might need this later"), interfaces with one implementer where the interface adds no testability win, factory functions that just call a constructor.
  • KEEP: abstractions that provide a real seam (testability, multiple implementers, framework-required boundaries).
  1. Boundary violations — wrong-layer imports (UI importing DB driver), leaky responsibilities (handler doing business logic that belongs in a service), hidden coupling (module A reads module B's private state), side effects in pure-named functions.
  • KEEP: pragmatic short-circuits already established as a pattern in this codebase. Flag for human judgment if unsure.
  1. Dead code — unused imports, unused private functions/methods, unreachable branches, stale feature flags, debug leftovers (console.log, print(...), dbg!), removed-but-still-referenced code.
  • KEEP: code referenced via reflection, dynamic dispatch, or string lookup. Code intentionally kept as a feature flag rollback path (verify with the user).

Hidden cost

  1. Duplication — copy-pasted branches with trivial differences, redundant helpers that do the same thing in two places, repeated literal/magic-number sequences.
  • KEEP: incidental duplication (two pieces of code that look similar but serve different intents that could diverge). Prefer leaving them separate over forcing a premature shared abstraction.
  1. Performance equivalences (behavior-preserving optimizations) — changes that are provably equivalent in semantics but cheaper in time/space:
  • O(n²) → O(n) when correctness preserved (e.g., set lookup vs list scan)
  • Repeated computation inside a loop → hoist outside
  • Unnecessary intermediate collections (eager list(...) when only iterated once → generator)
  • String concatenation in loop → join
  • Redundant DB/API calls in a loop → batch
  • Redundant deep copies / clones
  • .length / len() recomputed inside loop → cache

Hard rule: only apply when behavior equivalence is obvious. Do NOT change algorithms with subtle correctness implications. Do NOT micro-optimize hot paths without a benchmark. If in doubt, SKIP.

Behavior coverage

  1. Missing tests — behavior present in changed files that is not locked by any regression test. The fix is not to remove code but to ADD the narrowest test that pins the behavior. EXCEPTION: a PROSE file (prompt, SKILL.md, rule, markdown) has no behavioral seam — do NOT add a text/word-count/phrase pin for it; that guards a diff, not behavior. Cover only a machine-consumed value (parsed field, sentinel a runtime greps, a doc JSON sample through its real validator) or leave it to review.

Structural

  1. Oversized modules — any source file exceeding 250 pure LOC (non-blank, non-comment lines). This is an architectural defect, not a style preference. Measure: awk '!/^[[:space:]]$/ && !/^[[:space:]](#|\/\/)/' | wc -l.

When found, do NOT just flag it. Execute a full modular refactoring:

  1. Run check-no-excuse-rules.py recursively on scope to list all violations.
  2. For each oversized file, identify distinct responsibilities (single-responsibility principle).
  3. Plan the split: name each new file after the concept it owns (never utils.py, helpers.py, common.py, part_1.py).
  4. Present the split plan to the user before executing.
  5. Extract into clean modules with explicit init.py re-exports (re-exports ONLY, no logic in init.py).
  6. Verify: run check-no-excuse-rules.py again — every file must be ≤250 pure LOC. Run tests, typecheck, lint.

Forbidden escapes:

  • Counting blanks/comments toward budget.
  • Splitting by token count (foo1.py, foo2.py) — split by what each file DOES.
  • Catch-all dump files (utils.py, helpers.py, service.py).
  • "It's generated" — only valid if the file lives in a build output directory.
  • "230 LOC, close enough" — a 230-LOC file about to grow is already over. Split now.

KEEP: genuinely self-contained single-responsibility scripts (e.g., a standalone CLI checker). Opt out with # noqa: SIZE_OK in first 5 lines and a comment explaining why.

Quality Gates

A pass is complete only when all applicable gates are green. Skip gates that are genuinely N/A for the project (e.g., no security scanner configured), and report N/A explicitly — do not silently skip.

Process

Phase 0: Plan with TodoWrite

Create todos for all phases below. Mark in_progress one at a time.

Phase 1: Determine scope

If file paths were passed as arguments, that is the scope. Otherwise:

git diff $(git merge-base main HEAD)..HEAD --name-only

Filter out: deleted files, binary files, generated/vendored files (node_modules/, dist/, target/, lockfiles). List the final scope.

Phase 2: Lock behavior with regression tests (NEW — non-negotiable)

For each in-scope source file:

  1. Identify the public/observable behavior the file exposes (exported functions, HTTP handlers, CLI commands, classes used elsewhere).
  2. Check whether existing tests cover that behavior. Use git grep / project test conventions to find related test files.
  3. If behavior is uncovered or weakly covered, write the narrowest regression test that pins current behavior BEFORE editing the file. Tests should pin observable outputs, not implementation details. A PROSE file (prompt/SKILL.md/rule/markdown) is exempt — its wording is not behavior; skip the test and rely on review, or assert only a machine-consumed value.
  4. Run the test suite (or at minimum the relevant tests). They must be green before any cleanup begins.

If you cannot establish a green baseline (e.g., test runner is broken), STOP and report. Do not proceed with cleanup on unverified ground.

Phase 3: Cleanup plan — existence first, then smells

The largest, safest deletion is code that should not have existed. Before categorizing smells, run the deletion ladder on each changed unit:

  • Delete entirely — the behavior is not needed (YAGNI, speculative, dead on arrival).
  • Reuse — an existing helper or pattern in this repo already does it; replace the reimplementation with a call to it.
  • Platform / stdlib / native / dependency — the language stdlib, the runtime, or an already-installed dependency already does it (a hand-rolled date picker → , a custom query parser → URLSearchParams, a bespoke debounce → the util already imported).
  • Simplify in place — it must exist; make it smaller.

Only code that lands on Simplify in place proceeds to the smell categories. This turns the pass from "find smells to trim" into "first decide whether the code should exist, then trim what survives." One function replaced by a platform call is a bigger, safer win than any in-place cleanup — and it needs no per-line smell analysis.

For a diff that fixes a bug, grep the callers of every shared function it touches. Prefer one root-cause fix at the shared seam over repeated guards at each caller — a per-caller patch that leaves a sibling caller broken is a partial fix, not a cleanup.

Then produce an explicit plan before spawning the removal agents:

File: src/foo.py
  Ladder: 2 units simplify-in-place; 1 unit delete (native <input> replaces custom picker)
  Categories: dead code, excessive complexity, performance
  Order: dead code → complexity → performance
  Risk: medium (touches caching layer)

File: src/bar.py
  Ladder: all simplify-in-place
  Categories: obvious comments, over-defensive
  Order: comments → defensive
  Risk: low

Intentional shortcuts: if the plan deliberately keeps a bounded simplification (a naive scan fine under N rows, a global lock, an O(n²) path), mark it in-code with a debt: comment naming the ceiling and the upgrade trigger (in omo, prefix with // @allow so the comment-checker treats it as intentional), and list it under "Remaining Risks / Deferred" in the report. That section is the debt ledger — a simplification with a known ceiling and no marker is indistinguishable from a bug.

Order rule (safest → riskiest): comments → dead code → defensive → duplication → complexity → abstraction/boundary → performance → tests → oversized-modules. This minimizes blast radius of any one change.

Phase 4: Parallel slop removal via deep-low agents in batches of 5

More skills from code-yeongyu/oh-my-openagent

  • Aast-grepSearches and rewrites code by AST shape across 25 languages. Use when the target is a syntax pattern (every call/class/import shaped like X, a codemod, a YAML rule) rather than literal text; for plain strings, comments, or filenames, use rg.
  • AbrowserDrives a real browser through the omowright library from the js eval kernel: sites the user is already signed into, forms and clicks, JS-rendered pages, screenshots, web QA, extension popups, a human handoff for login, CAPTCHA or OTP, and a browser you own for scraping, bot-scored targets, network capture and QA traces. Use for any interactive browser task; not for a plain search or an unblocked static fetch.
  • Acodex-qaQA the omo Codex Light edition (lazycodex / packages/omo-codex) itself, in strict isolation so ONLY our plugin is exercised, never the user's real ~/.codex. The first-party method drives the real `codex app-server` against an isolated CODEX_HOME plus a LOCAL mock model (no real API call), and proves a plugin hook fired by asserting hook/started + hook/completed notifications. Also: isolated install verification, per-component hook probes, a tmux TUI smoke, and runtime log observation (RUST_LOG / logs SQLite / /debug-config). Ships tested helper scripts each with a --self-test. Use whenever someone changes anything under packages/omo-codex or wants to QA, smoke-test, verify, or debug the Codex plugin, its hooks/components, the installer/config.toml, the app-server flow, or the Codex TUI. Triggers: codex qa, qa codex, codex-qa, test codex plugin, verify codex hook, codex app-server, lazycodex qa, isolated CODEX_HOME, prove codex hook fired, codex tui test.
  • Acoding-agent-sessionsFinds, reads, and reconstructs coding-agent sessions across Codex, Claude, OpenCode, OMO/Senpi, and other local agent logs. Use when asked to find or search past sessions, transcripts, or subagent runs, or to recover what an earlier session did.
  • Acomment-checkerUse when Codex needs to understand or respond to automatic comment-checker feedback emitted after an edit-like PostToolUse hook.
  • Adag-libraryStores a DAG definition once and re-runs it by name, instead of pasting the definition into every run. Use when the user wants to save a DAG, run a saved one, or schedule the same multi-agent graph repeatedly.
  • Ddata-scientistProcesses and analyzes data with resident-kernel engines (DuckDB, Polars) and one-shot tools. Use for CSV/parquet/JSON analysis, group-by/join/aggregation, time series, distributions, cleaning, or plotting a dataset.
  • AdebuggingRuns a hypothesis-driven debugging loop across any language or binary, escalating to orthogonal oracle angles and locking the fix with a failing test. Use for crashes, silent failures, hangs, wrong responses, memory leaks, async misbehavior, or reverse engineering.
  • Adev-browserBrowser automation with persistent page state. Use when users ask to navigate websites, fill forms, take screenshots, extract web data, test web apps, or automate browser workflows. Trigger phrases include "go to [url]", "click on", "fill out the form", "take a screenshot", "scrape", "automate", "test the website", "log into", or any browser interaction request.
  • CfrontendBuilds, styles, and polishes web UI and UX. Use for any frontend, page, component, styling, layout, animation, or visual-quality task, or when asked to make an interface look or feel a certain way.
  • AfrontendBuilds, styles, and polishes web UI and UX. Use for any frontend, page, component, styling, layout, animation, or visual-quality task, or when asked to make an interface look or feel a certain way.
  • Aget-unpublished-changesCompare HEAD with the latest published npm versions and list all unpublished changes by release layer. Triggers: unpublished changes, changelog, what changed, whats new.

All agent skills → · MCP servers