ulw-execute skill
Executes a written ulw-plan work plan with Boulder state, evidence ledger, worktree discipline, and parallel subagents. Use when the user says ulw-execute or asks to run a .omo/plans plan.
Is the ulw-execute skill safe?
Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.
No findings.
Install the ulw-execute skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/code-yeongyu/oh-my-openagent.git /tmp/oh-my-openagent mkdir -p ~/.claude/skills cp -r /tmp/oh-my-openagent/packages/shared-skills/skills/ulw-execute ~/.claude/skills/ulw-execute
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
ABSOLUTE RULE: YOU ARE AN ORCHESTRATOR — NEVER THE IMPLEMENTER
YOU DO NOT WRITE CODE. YOU DO NOT EDIT PRODUCT FILES. YOU DO NOT RUN QA YOURSELF. EVERY unit of implementation, test, QA, and review work MUST be delegated to a spawned subagent. NO EXCEPTIONS. Your hands touch only plan selection, .omo/ state (Boulder, ledger, plan checkboxes), decomposition, dispatch, verdicts, and evidence records. About to edit a product file or run an implementation command yourself? STOP. SPAWN A WORKER INSTEAD. Orchestrate at MAXIMUM PARALLELISM: every independent unit runs concurrently; only named dependencies serialize.
Codex Harness Tool Compatibility
Translate any OpenCode-only tool name in an inherited example to its Codex equivalent:
When translating load_skills=[...], name the skills inside the spawned agent's message. If a code block below conflicts with this section, this section wins.
Codex exposes ONE of two subagent tool surfaces per session; check your own tool list and route accordingly. If multiagentv1. tools exist, use the table above as written. If instead a flat spawnagent with a required taskname exists (multiagentv2), rewrite every multiagentv1. example: multiagentv1.spawnagent({...,"forkcontext":false}) becomes spawnagent({"taskname":"","message":...,"agenttype":...,"forkturns":"none"}) ("all" only when full parent history is truly required); sendinput becomes sendmessage; do not call closeagent/resumeagent (finished agents end on their own; followuptask re-tasks one, interruptagent stops one); waitagent takes only timeoutms and returns on any child mailbox activity. On the v2 surface agent_type may be absent from the spawn schema — when absent, omit it and describe the role inside message. If a code block below conflicts with this section, this section wins.
Codex tier mapping for the delegation router
When tier worker agents are installed, map the delegation router's parenthesized difficulty to agenttype: (low) -> lazycodex-worker-low; (medium) -> lazycodex-worker-medium; (high) -> lazycodex-worker-high. Explorer/librarian research lanes keep their own roles. On spawn surfaces without agenttype, state the tier inside message. Difficulty (model power) is orthogonal to the LIGHT/HEAVY rigor tier in step 4 — judge each on its own facts.
Codex Subagent Reliability
Every multiagentv1.spawnagent message is a self-contained executable assignment: TASK: , then DELIVERABLE, SCOPE, and VERIFY, with role instructions inside message. Use forkcontext: false unless full history is truly required; paste only the context the child needs.
Plan and reviewer agents may run for a long time: spawn them in the background and keep doing independent root work. Between multiagentv1.waitagent calls, back off — double the timeout up to ~5 minutes — instead of spinning short cycles. A timeout only means no new mailbox update arrived; treat a running child as alive. Require WORKING: - before long passes and BLOCKED: only when progress stops. Keep the parent visibly alive with active subagent count, names, and latest WORKING: phase. Fallback only when the child is completed without the deliverable, ack-only after followup, explicitly BLOCKED:, or no longer running — then record inconclusive (never a pass), close if safe, and respawn a smaller forkcontext: false task with the missing deliverable.
ulw-execute
Execute a work plan until every top-level checkbox is complete. This skill pairs with the harness's ulw-execute continuation hook, which re-injects the next turn while .omo/boulder.json says this codex: still has unchecked plan work.
Usage
$ulw-execute [plan-name] [--worktree <absolute-path>] [--make-pr] [--ship]- plan-name (optional): a full or partial file stem under .omo/plans/.
- --worktree (optional): reuse an existing task-owned worktree for the first phase instead of creating one; every phase runs in a task-owned worktree regardless.
- --make-pr (optional): deliver each phase's worktree as a pull request — push the branch, open a reviewer-readable PR, hand off with the URL, and merge only if the user asks.
- --ship (optional): full delivery lifecycle; implies --make-pr. After the PR opens, stay on the job until it is MERGED: watch CI and review gates, fix failures and address feedback from the worktree (fresh QA evidence for behavior changes), merge per the repository's merge policy, then remove the worktree and sync .omo/ state back.
Goal and todo discipline (MANDATORY)
Do ALL of this immediately after the plan is selected, BEFORE the first implementation dispatch. Skipping any step is a defect.
- Set the goal, in detail. When a goal tool is available (create_goal), call it with a DETAILED objective: the plan name and path, the concrete end state, the phase and task counts, the delivery mode (direct, --make-pr, or --ship), and how completion will be verified. One work session = one registered goal (the goal tool holds one active goal); each phase then carries its own concrete goal — the ledger entry Phase 2 records before the wave's first dispatch, defined from the previous phase's landed and verified evidence. No goal tool -> record the same objective as the first ledger entry.
- Register every phase and task as todos. Mirror the plan into the todo/plan tool of your harness: one phase per plan wave, one todo per column-zero checkbox (including the final verification wave). Register ALL of them up front - never keep tasks in memory only.
- Keep them current at every moment. Mark a todo in_progress when its work dispatches and done immediately after its verification passes. Never batch-complete at the end, never execute work that is not a registered todo; discovered work — a pre-existing bug, failing test, stale doc, or wrong guidance — is appended to the plan as a checkbox, mirrored as a todo before it runs, and fixed to the ideal state, never deferred as a follow-up. A worker that meets a defect outside its assigned files reports it instead of fixing it; only the orchestrator appends the checkbox and dispatches it to a correctly scoped unit. The todo list, Boulder state, and plan checkboxes must always tell the same story.
Phase 1: Select the plan
- Read .omo/boulder.json if it exists.
- List work plan files under .omo/plans/.
- If plan-name was provided, select the matching plan.
- If exactly one active or paused Boulder work exists for this session, resume it.
- If no active work exists and exactly one plan exists, select it.
- If no active work exists and there is no selectable plan, enter No-plan bootstrap.
- If multiple plans remain possible, ask one focused selection question.
No-plan bootstrap
When the user explicitly said start work / $ulw-execute and no selectable plan exists, treat that phrase as approval: bootstrap ulw-plan to create the approved plan before execution and implementation, instead of stalling or asking for generic approval again. A brief or notes file without waves, checkboxes, and acceptance criteria is NOT decision-complete — enter this bootstrap too.
- Invoke the ulw-plan skill from the current request and require its dynamic adversarial workflow: collect, verify, design, adversarial plan-review, synthesize.
- The generated work plan must be saved under .omo/plans/.md before implementation or Boulder state writes that point at plan work.
- Use maximum safe parallelism in the generated plan: independent files/tasks fan out; same-file writes, shared state, and named dependencies serialize.
- Preserve safety boundaries. Ask one focused question only when the objective is missing, destructive, or has a safety/product ambiguity that repository exploration cannot resolve.
- After the plan exists, continue directly to Phase 2.
Phase 2: Create or update Boulder state
Write .omo/boulder.json before implementation starts. Prefix session ids with codex: so the continuation hook can identify its own session.
{
"schema_version": 2,
"active_work_id": "<work-id>",
"works": {
"<work-id>": {
"work_id": "<work-id>",
"active_plan": ".omo/plans/<plan-name>.md",
"plan_name": "<plan-name>",
"session_ids": ["codex:<session_id>"],
"status": "active",
"worktree_path": null
}
}
}Every phase (plan wave) runs in its own task-owned worktree with its own goal: before the wave's first dispatch, record the wave's goal — its checkboxes and their acceptance criteria — as a ledger entry, then git worktree add -wt/- (or verify a --worktree path with git worktree list --porcelain), store the absolute path as worktree_path, run every edit, command, test, and evidence capture inside it; the wave lands on the integration base once its checkboxes are verified (direct merge, or the PR under --make-pr/--ship), and the next wave branches from that landed base.
Parallel delivery lanes (teams and worktrees)
Solo orchestration with parallel background workers is the default topology. Decide once, when the wave's lanes are known, and record the verdict in the ledger:
- Independent lanes -> parallel workers. Separate files, no shared contract: one parallel spawn burst; no team.
- Dependency-ordered lanes -> one workflow run per wave. Sub-tasks with real ordering between them (C needs A and B finished first) and a harness with a native workflow tool: dispatch the wave as ONE run (one producer node per lane plus a verification node); recover inside it with retry/amend/send; let node completions wake you instead of arming per-lane watchers; the next wave is a NEW run (or amend when only the definition changed) — never one graph for the whole plan. Read the mass-ulw skill's SKILL.md and references/planning.md IN FULL before defining any graph.
- Overlapping lanes -> a team. The lanes touch the same module or contract AND running them concurrently actually finishes sooner: stand up a team (where the harness has one) so one lane's discoveries relay through you mid-flight.
- PR-mode independent lanes -> a worktree per lane. Under --make-pr/--ship, when a wave holds independent checkboxes, give each lane its own branch and task-owned worktree, delivered as its own PR.
Landing rules, regardless of topology:
- Merge per verified unit. A lane lands the moment its own gates pass — it never waits for the slowest sibling. Integrate landed work back into the base the remaining lanes branch from.
- Only the orchestrator merges. Workers and team members never merge and never push the base branch.
- Conflicts are the orchestrator's job. Decide the landing order and tell the later lane what changed; a worker never resolves a sibling's conflict blind.
Phase 3: Execute the next checkbox
- Read the full selected plan.
- Find the first unchecked column-0 checkbox in ## TODOs or ## Final Verification Wave.
- Ignore nested checkboxes under acceptance criteria, evidence, and definition-of-done sections.
- Classify the checkbox tier and record it in its ledger entry. Default is LIGHT — a narrow change inside existing layers. Take HEAVY only on a fact you can point to: a new module / abstraction / domain model; auth, security, or session; an external integration; a DB schema or migration; concurrency or transaction boundaries; a cross-domain refactor; or the plan or user signals care. When unsure, take HEAVY; upgrade and redo skipped gates the moment a HEAVY fact surfaces; never downgrade.
- Decompose that checkbox into atomic sub-tasks sized for ONE worker in ONE run — a sub-task that would need mid-flight steering is two sub-tasks. Collect every other unchecked checkbox in the same plan wave whose dependencies are met — their lanes execute concurrently. A wave that could split further but holds fewer than 3 independent sub-tasks is under-split.
- DELEGATE EVERYTHING. YOU NEVER IMPLEMENT. Route every sub-task through the delegation router below, then dispatch ALL independent sub-tasks across those checkboxes in one parallel worker-spawn burst (a single batched spawn call where the harness supports it); route named dependencies per the lane-topology decision above. Verification and checkbox marking stay per-checkbox.
- Give every dispatched sub-task its completion condition and watch for it per the section below. A dispatch whose completion nobody watches is an unfinished dispatch.
Monitor every dispatched subagent to its completion condition
A spawned worker is not fire-and-forget. For EACH subagent in the burst, name the observable state that ends its lane — the file written, the PR opened, the checkbox's gates green — and put a watcher on THAT state, never on a clock.
- Arm one watcher per lane, at spawn time. The worker's own completion arrives on its own as an injected notification; arm an explicit monitor on top of it only when the lane's completion condition lives OUTSIDE the child's final message — CI turning green, a log line, a build artifact appearing, a branch landing. Watch the state itself (monitor with a command that exits or emits on that condition), and keep the burst's watchers distinct so one lane firing never reads as another's.
- NEVER poll and NEVER sleep. No sleep, no timed retry loop, no re-reading the same status hoping it changed. Between waves, do independent root work or end the turn; an idle session is always woken. A single task_output({ mode: "tail" }) peek is allowed only when a midpoint decision genuinely depends on it.
- Tear the watcher down the instant it resolves. The moment a monitor fires, or you discover it was armed on the WRONG condition (it watches a path the lane never touches, a pattern that can never match, a lane you already cancelled), stop it with kill_bash and say so in the ledger. A stale watcher re-fires on unrelated output and corrupts the next wave's verdict.
- Then advance. Fired watcher plus verified evidence means that lane's gates run and its checkbox closes; a mis-set watcher means re-arm it on the right condition or drop it, and continue. Never let a dead watcher hold the run open, and never treat watcher silence as a pass.
Delegation router — recommended task executor category
When the plan annotates a todo with Recommended task executor category:, follow that annotation; deviate only for a reason recorded in the ledger entry. Otherwise route by shape, in the omo category vocabulary (category-capable harnesses pass it directly on the worker-spawn tool, e.g. task(category="quick", ...); others map the parenthesized difficulty):
Sizing is a two-branch decision made per checkbox, before dispatch:
- Splittable work splits. When the checkbox decomposes into independent pieces, dispatch them as a swarm of quick/unspecified-low workers in ONE parallel burst — many small cheap workers in parallel beat one large delegation.
- Cohesive hard work stays whole. When splitting would sever shared reasoning (one algorithm, one migration, one subtle bug), send the WHOLE problem to deep-low, deep-high or ultrabrain as ONE delegation. Never force-split work whose parts share one insight.
Each sub-task message must include:
- Goal and exact files or directories in scope.
- The tests already covering the touched behavior, READ before any edit as the behavior of record (intent, coverage, pass); a bug's reproduction captured before the fix. A new test ONLY where the repository keeps tests for this behavior AND a regression would otherwise pass unnoticed by the sub-task's Manual-QA scenario — never one that mirrors its implementation (mock-call assertions, pinned constants) or restates the change.
- Implementation constraints from the plan and project rules.
- Automated verification commands to run.
- One Manual-QA channel, named with the exact tool and exact invocation (the literal curl, send-keys, page.click / session.click, payload, selectors, and the binary observable that decides PASS/FAIL), not "verify it works". A LIGHT checkbox needs one real-surface proof of its deliverable, and auxiliary surfaces (CLI stdout, DB state diff, parsed config dump) are first-class when the surface is CLI- or data-shaped:
- HTTP call: curl -i against the live endpoint.
- Terminal / TUI: drive a real pty; tmux send-keys is fine for a boot/behavior smoke, but color/layout/CJK evidence goes through the xterm.js web terminal below, NEVER tmux capture-pane.
- Browser use: omowright from js eval (staged in the browser skill) — the owned engine (connectPipe on a task-owned profile, connectCloakProfile for bot-scored targets) for unauthenticated pages, the attached engine (connectBrowserSkill() in the user's signed-in browser) when the page needs their login; never a clone of or a launch against the live profile.
- Computer use: OS-level GUI automation against the running desktop app when the surface is not a page.
- TUI visual evidence: when a TUI claim needs visual QA or PR proof, run bun script/qa/web-terminal-visual-qa.mjs --command "" --input "{Enter}" --evidence-dir (real pty rendered through xterm.js in Chrome) and attach terminal.png plus metadata.json.
- The adversarial classes that apply to this sub-task (from the 9 ultraqa classes) and how each is probed.
- Required artifact path and cleanup receipt.
- Tool-use expectations: batch independent tool calls in parallel; when the harness exposes a code-execution surface (eval), use it for multi-call steps instead of one-by-one calls.
The 9 ultraqa classes are trigger-mapped: new input parsing → malformed input; untrusted external text → prompt injection; resumable or long-running flows → cancel/resume; generated or cached artifacts → stale state; uncommitted user files in scope → dirty worktree; long external commands → hung or long commands; new or timing-sensitive tests → flaky tests; log-based success claims → misleading success output; mid-operation interrupts → repeated interruptions. A class applies when its trigger fact holds. Probe each appl
More skills from code-yeongyu/oh-my-openagent
- Aast-grepSearches and rewrites code by AST shape across 25 languages. Use when the target is a syntax pattern (every call/class/import shaped like X, a codemod, a YAML rule) rather than literal text; for plain strings, comments, or filenames, use rg.
- AbrowserDrives a real browser through the omowright library from the js eval kernel: sites the user is already signed into, forms and clicks, JS-rendered pages, screenshots, web QA, extension popups, a human handoff for login, CAPTCHA or OTP, and a browser you own for scraping, bot-scored targets, network capture and QA traces. Use for any interactive browser task; not for a plain search or an unblocked static fetch.
- Acodex-qaQA the omo Codex Light edition (lazycodex / packages/omo-codex) itself, in strict isolation so ONLY our plugin is exercised, never the user's real ~/.codex. The first-party method drives the real `codex app-server` against an isolated CODEX_HOME plus a LOCAL mock model (no real API call), and proves a plugin hook fired by asserting hook/started + hook/completed notifications. Also: isolated install verification, per-component hook probes, a tmux TUI smoke, and runtime log observation (RUST_LOG / logs SQLite / /debug-config). Ships tested helper scripts each with a --self-test. Use whenever someone changes anything under packages/omo-codex or wants to QA, smoke-test, verify, or debug the Codex plugin, its hooks/components, the installer/config.toml, the app-server flow, or the Codex TUI. Triggers: codex qa, qa codex, codex-qa, test codex plugin, verify codex hook, codex app-server, lazycodex qa, isolated CODEX_HOME, prove codex hook fired, codex tui test.
- Acoding-agent-sessionsFinds, reads, and reconstructs coding-agent sessions across Codex, Claude, OpenCode, OMO/Senpi, and other local agent logs. Use when asked to find or search past sessions, transcripts, or subagent runs, or to recover what an earlier session did.
- Acomment-checkerUse when Codex needs to understand or respond to automatic comment-checker feedback emitted after an edit-like PostToolUse hook.
- Adag-libraryStores a DAG definition once and re-runs it by name, instead of pasting the definition into every run. Use when the user wants to save a DAG, run a saved one, or schedule the same multi-agent graph repeatedly.
- Ddata-scientistProcesses and analyzes data with resident-kernel engines (DuckDB, Polars) and one-shot tools. Use for CSV/parquet/JSON analysis, group-by/join/aggregation, time series, distributions, cleaning, or plotting a dataset.
- AdebuggingRuns a hypothesis-driven debugging loop across any language or binary, escalating to orthogonal oracle angles and locking the fix with a failing test. Use for crashes, silent failures, hangs, wrong responses, memory leaks, async misbehavior, or reverse engineering.
- Adev-browserBrowser automation with persistent page state. Use when users ask to navigate websites, fill forms, take screenshots, extract web data, test web apps, or automate browser workflows. Trigger phrases include "go to [url]", "click on", "fill out the form", "take a screenshot", "scrape", "automate", "test the website", "log into", or any browser interaction request.
- CfrontendBuilds, styles, and polishes web UI and UX. Use for any frontend, page, component, styling, layout, animation, or visual-quality task, or when asked to make an interface look or feel a certain way.
- AfrontendBuilds, styles, and polishes web UI and UX. Use for any frontend, page, component, styling, layout, animation, or visual-quality task, or when asked to make an interface look or feel a certain way.
- Aget-unpublished-changesCompare HEAD with the latest published npm versions and list all unpublished changes by release layer. Triggers: unpublished changes, changelog, what changed, whats new.