ulw-loop skill
A goal-like loop that decomposes work into systematic, evidence-bound ultrawork steps. Use when the user wants a goal loop or durable, checkpointed execution.
Is the ulw-loop skill safe?
Clean: nothing in its files matched our rules. We read 3 files in the folder on 2026-09-28.
No findings.
Install the ulw-loop skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/code-yeongyu/oh-my-openagent.git /tmp/oh-my-openagent mkdir -p ~/.claude/skills cp -r /tmp/oh-my-openagent/packages/omo-senpi/skills/ulw-loop ~/.claude/skills/ulw-loop
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
ulw-loop
Use this skill when the user asks for ulw-loop, ulw, durable goal execution, evidence-led work, manual QA, or checkpointed long-running delivery.
This skill is compact by design: the run contract below is the whole bootstrap. references/full-workflow.md and references/define-goal.md carry the full doctrine; open a section only when the phase you are in needs it.
Run contract
- Create goals from a JS eval cell: agentToolkit.createGoals({ brief }) after the SDK import below. The SDK binds this session from the host env, so never pass a session id or a plan path. If the envelope reports ULWLOOPPLANEXISTSCOMPLETE (this session's aggregate is already complete), unrelated new work needs a fresh session; createGoals({ brief, force: true }) is only for deliberately overwriting completed evidence.
- Register the aggregate objective from the returned handoff with create_goal, shaped by references/define-goal.md. Goal creation is NEVER skipped.
- Mirror every atomic step into the live todo checklist: one granular step per action, exactly one in_progress, transitions marked the instant they happen.
- Treat each goal as a phase: create its own worktree off the integration base; dispatch its dependency-ordered lanes as ONE workflow run (read the mass-ulw skill first; ordering-free lanes stay a task batch); verify every criterion with real-surface evidence; land the worktree on the integration base at agentToolkit.checkpoint({ goalId, status: "complete", evidence }) per the repository's flow (direct merge or merged PR); define the next goal's run from what this one proved. Tests alone never prove done. When a mass-ulw pointer accompanies this skill, this contract still owns goals, criteria, evidence, and checkpoints.
- Stop when the goal's WHEN-TO-STOP line holds with evidence in hand.
When the injected ultrawork directive accompanies this skill, its goal/notepad/todo bootstrap is subsumed by this contract: the loop SDK owns goal state and the loop ledger is the notepad. Do not create a second one.
The SDK: one import, then method calls
Every ulw-loop operation runs inside a JS eval cell through the SDK the extension publishes at OMOAGENTTOOLKITSDKROOT. There is no omoagenttoolkit tool and no CLI to spawn on Senpi.
const { agentToolkit } = await import(`${env("OMO_AGENT_TOOLKIT_SDK_ROOT")}/sdk.js`)
print(await agentToolkit.status())Rules that keep it working:
- JS cells only. From a py, rb, or jl cell, run a separate eval with language js.
- Import once per kernel lifetime. agentToolkit stays bound in later cells; re-import only after a kernel restart or a ReferenceError: agentToolkit is not defined.
- Every call resolves to an envelope, never a throw: { ok: true, operation, result, nextActions, warnings? } or { ok: false, operation, error: { code, message }, warnings? }. nextActions are things to do next; warnings are facts to know (a fallback the binder took, a driver objective that differs). Read nextActions before deciding the next step, and branch on error.code, not on the message text.
- Never pass a session id or plan path. The SDK binds PISESSIONID and PISESSIONCWD from the host env on each call; state lives under .omo/ulw-loop//. When PISESSIONCWD is missing (seen after an extension reload restarted the kernel), the binder falls back to the cwd recorded in the PISESSIONFILE header, then to process.cwd(), and every envelope carries a warning naming the fix: env("PISESSIONCWD", ""). status().result.binding shows { cwd, cwdSource, sessionId, goalStorePaths }, so check it after any kernel restart and re-pin the env when cwdSource is not PISESSIONCWD.
- The driver snapshot is filled automatically from this session's goal store. Pass codexGoalJson only to override it.
Methods (argument fields are exact):
Non-Negotiables
- Write loop state only through the SDK; it lives under .omo/ulw-loop// and is never hand-edited. Mutations are serialized across processes by the session's .state.lock, so parallel recordEvidence calls from workers are safe.
- Register goals up front, shaped by references/define-goal.md (agentToolkit.createGoals({ brief }), then creategoal from the returned handoff), and mirror every atomic step into the live todo checklist: one ultra-granular step per action, exactly one inprogress, transitions marked the instant they happen.
- After any compaction or context loss, re-read brief + goals + ledger FIRST plus agentToolkit.status() (re-import the SDK if the kernel restarted; help() lists every method with its argument fields), confirm result.binding.cwdSource is PISESSIONCWD and re-pin it with env("PISESSIONCWD", result.binding.cwd) when it is not, then resume; never re-plan from scratch.
- If createGoals answers ULWLOOPPLANEXISTSCOMPLETE, this session's aggregate is already done: start unrelated new work in a fresh session instead of steering or forcing the completed state. Use force: true only to intentionally overwrite completed evidence.
- Every success criterion needs observable evidence from a real surface: a channel (terminal/TUI via the xterm.js web terminal, HTTP, browser, computer-use) or, for CLI- or data-shaped criteria, an auxiliary surface (CLI stdout, DB diff, parsed config dump).
- Evidence is bound to the tree it was captured at (git rev-parse --short "HEAD^{tree}"); it goes stale only when tracked content changes — a rebase or amend that keeps the tree identical keeps it valid. When the tree differs, re-run at the current HEAD and re-record, never relabel or regenerate. Record only after cleanup receipts exist.
- Delegate code edits, test writes, fixes, and QA execution to right-sized omo-senpi subagents through the native task tool or through workflow nodes when the phase's lanes carry ordering.
- Use git-master for git-tracked edits: inspect recent and touched-path commit history, then commit each verified work unit atomically in the repository's observed language, scope, and message style with only that unit's files staged. Never carry verified units into a later omnibus commit.
Team mode: decide it, do not default to it
Solo execution with parallel background task workers is the default: fan independent units out in one batched spawn, each routed to the category (or configured subagenttype) that fits it, with scopes cut so no two workers write the same files. A team (teamcreate) adds per-member briefing, shared-state, and relay overhead, so it must be paid for by the work's shape. Decide ONCE, when the plan's work units are known, and record the verdict plus its reason in the notepad.
Stand up a team when BOTH hold:
- The units' scopes overlap in a way you cannot cleanly cut. They touch the same module, contract, or migration, so one unit's discovery changes what another should do. Fire-and-forget workers cannot exchange that mid-flight; teammates can, because the lead relays it.
- Running them at the same time actually finishes sooner. The units are each substantial and none is merely waiting on another's output. Two units where the second only consumes the first's result are a sequence, not a team.
When the units are genuinely independent — separate files, no shared contract — spawn parallel background task workers instead and avoid the team coordination overhead entirely. When the work is one cohesive unit, do it yourself. Overlap alone is not enough: near-identical units that would collide on the same lines are faster done in sequence by one worker.
Under team mode, isolate and land per unit:
- One git worktree per member, never a shared checkout — concurrent members editing one working tree corrupt each other's diffs and evidence. Give each member its own branch off the base and its own worktree path.
- Merge per work unit, as each unit is verified. A member's unit lands when its own evidence is captured and its gates are green; it does not wait for the slowest sibling. Integrate each merged unit back into the base the others branch from, so overlapping members rebase onto real merged work rather than guessing at it.
- Conflicts are the lead's job. When two members' units touch the same lines, the lead decides the order they land and tells the later member what changed; members never resolve a sibling's conflict blind.
Native Senpi Task Contract
Senpi already exposes its real subagent spawn surface through the omo-senpi task component. Use it directly. Do not route delegation through external app-server threads or another harness.
Every worker prompt starts with TASK: and names DELIVERABLE, SCOPE, VERIFY, and STOP WHEN. Put requested skill names and all required context inside prompt; children do not inherit interview context automatically.
Driver goal lifecycle
The session goal you create with creategoal is a DRIVER the loop instructs, never a gate. A checkpoint accepts an optional driver snapshot, records it verbatim, and never rejects on its status or objective; the advice comes back in nextActions. The snapshot is filled from this session's goal store automatically, so pass codexGoal only to override it. If the driver was completed early, the advice tells you to creategoal again with the plan's objective verbatim - a completed goal is replaced, not reused. If the driver is paused or usage/budget limited, the advice tells you to resume it. A different objective is a warning, not a refusal, and it is reported once: the first checkpoint under that driver records the objective in the plan (acknowledgedDriverObjectives) and later calls stay quiet; status().result.driver shows the relation at any time. create_goal is advised only when the goal store really holds no goal.
More skills from code-yeongyu/oh-my-openagent
- Aast-grepSearches and rewrites code by AST shape across 25 languages. Use when the target is a syntax pattern (every call/class/import shaped like X, a codemod, a YAML rule) rather than literal text; for plain strings, comments, or filenames, use rg.
- AbrowserDrives a real browser through the omowright library from the js eval kernel: sites the user is already signed into, forms and clicks, JS-rendered pages, screenshots, web QA, extension popups, a human handoff for login, CAPTCHA or OTP, and a browser you own for scraping, bot-scored targets, network capture and QA traces. Use for any interactive browser task; not for a plain search or an unblocked static fetch.
- Acodex-qaQA the omo Codex Light edition (lazycodex / packages/omo-codex) itself, in strict isolation so ONLY our plugin is exercised, never the user's real ~/.codex. The first-party method drives the real `codex app-server` against an isolated CODEX_HOME plus a LOCAL mock model (no real API call), and proves a plugin hook fired by asserting hook/started + hook/completed notifications. Also: isolated install verification, per-component hook probes, a tmux TUI smoke, and runtime log observation (RUST_LOG / logs SQLite / /debug-config). Ships tested helper scripts each with a --self-test. Use whenever someone changes anything under packages/omo-codex or wants to QA, smoke-test, verify, or debug the Codex plugin, its hooks/components, the installer/config.toml, the app-server flow, or the Codex TUI. Triggers: codex qa, qa codex, codex-qa, test codex plugin, verify codex hook, codex app-server, lazycodex qa, isolated CODEX_HOME, prove codex hook fired, codex tui test.
- Acoding-agent-sessionsFinds, reads, and reconstructs coding-agent sessions across Codex, Claude, OpenCode, OMO/Senpi, and other local agent logs. Use when asked to find or search past sessions, transcripts, or subagent runs, or to recover what an earlier session did.
- Acomment-checkerUse when Codex needs to understand or respond to automatic comment-checker feedback emitted after an edit-like PostToolUse hook.
- Adag-libraryStores a DAG definition once and re-runs it by name, instead of pasting the definition into every run. Use when the user wants to save a DAG, run a saved one, or schedule the same multi-agent graph repeatedly.
- Ddata-scientistProcesses and analyzes data with resident-kernel engines (DuckDB, Polars) and one-shot tools. Use for CSV/parquet/JSON analysis, group-by/join/aggregation, time series, distributions, cleaning, or plotting a dataset.
- AdebuggingRuns a hypothesis-driven debugging loop across any language or binary, escalating to orthogonal oracle angles and locking the fix with a failing test. Use for crashes, silent failures, hangs, wrong responses, memory leaks, async misbehavior, or reverse engineering.
- Adev-browserBrowser automation with persistent page state. Use when users ask to navigate websites, fill forms, take screenshots, extract web data, test web apps, or automate browser workflows. Trigger phrases include "go to [url]", "click on", "fill out the form", "take a screenshot", "scrape", "automate", "test the website", "log into", or any browser interaction request.
- CfrontendBuilds, styles, and polishes web UI and UX. Use for any frontend, page, component, styling, layout, animation, or visual-quality task, or when asked to make an interface look or feel a certain way.
- AfrontendBuilds, styles, and polishes web UI and UX. Use for any frontend, page, component, styling, layout, animation, or visual-quality task, or when asked to make an interface look or feel a certain way.
- Aget-unpublished-changesCompare HEAD with the latest published npm versions and list all unpublished changes by release layer. Triggers: unpublished changes, changelog, what changed, whats new.