Mmcp.market

browser skill

by code-yeongyu·code-yeongyu/oh-my-openagent·70k stars

Drives a real browser through the omowright library from the js eval kernel: sites the user is already signed into, forms and clicks, JS-rendered pages, screenshots, web QA, extension popups, a human handoff for login, CAPTCHA or OTP, and a browser you own for scraping, bot-scored targets, network capture and QA traces. Use for any interactive browser task; not for a plain search or an unblocked static fetch.

A100/100content scan

Is the browser skill safe?

Clean: nothing in its files matched our rules. We read 13 files in the folder on 2026-09-28.

No findings.

Install the browser skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/code-yeongyu/oh-my-openagent.git /tmp/oh-my-openagent
mkdir -p ~/.claude/skills
cp -r /tmp/oh-my-openagent/packages/shared-skills/skills/browser ~/.claude/skills/browser
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Browser

One library, two engines. omowright ships inside this skill; choose the engine before you act:

Attached is the default, because it is the only engine carrying the user's logins and the only one where a human is a single call away. Never substitute one engine for the other silently: if the attached engine is not set up, run the onboarding script and tell the user its one remaining step.

Step 0 — load omowright and prove the stack

const { loadOmowright } = await import("<skill-root>/scripts/omowright.mjs")
const { omowright } = await loadOmowright()          // { connectBrowserSkill, bskSnapshot, connectPipe, ... }
node "<skill-root>/scripts/browser-doctor.mjs" --json

Install into the browser the user actually uses, never into whatever happens to be on disk. Before installing, check your memory for the user's browser; otherwise read the doctor's browser (picked from the OS default browser, running apps and recent use — candidates shows the evidence). If memory and the doctor disagree, or the doctor says choose-browser, ask the user. Pass the answer as --browser= and record it in memory. A Chrome that is merely installed is not their browser.

Never launch a headless browser because the attached one is missing. It has none of the user's sessions, so every login turns into a ladder you should not be climbing. Say which state you hit and ask.

The loop (attached)

const session = await omowright.connectBrowserSkill({ name: "<task>", focused: false })
try {
  await session.navigate("https://example.com/", { waitUntil: "load" })
  const { tree, refs, css } = await omowright.bskSnapshot(session, { interactive: true })  // OmOWright tree + refs, no trace in the page
  await session.click({ selector: css.e3 })                                                 // css[ref] is null inside shadow roots:
  const vom = await session.observe({ maxTokens: 4000 })                                    //   then read the daemon's own tree ...
  await session.click("@e7")                                                                //   ... and click its @eN ref
  await session.fill(css.e5, "hello")
  await session.press("Enter")
  await session.waitForNavigation({ waitUntil: "load" })
  const shot = await session.screenshot()                                                   // { buffer, width, height, captureId }
} finally {
  await session.stop()                                                                      // success AND failure; returns borrowed tabs
}

call; use a ref in the same cycle you read it.

  1. Read before every action. bskSnapshot refs and observe @eN refs are reissued on each

Borrowing prompts the user; never invent tab ids and never repeat a denied borrow.

  1. Navigation and large DOM changes stale every ref. Read again rather than reusing.
  2. Two identical failures mean change approach, not retry. A third identical attempt is a defect.
  3. Borrow a user tab explicitly (tabList({ scope: "user" }), tabBorrow(id), tabReturn(id)).
  1. Always stop() the session, on success and on failure.

Every method, its options, and the failure codes are in references/commands.md.

When a human is the only way through

Login, CAPTCHA, OTP, a payment confirmation, a consent dialog:

const outcome = await session.requestHelp({ prompt: "<what you need done>", targets: ["@e4"], timeoutMs: 300_000 })

Then read the page again. Respect a cancelled or timed_out outcome; do not work around it by changing the extension's automation settings.

Rules

cookie or recovery code. The value of the attached engine is that the browser is already signed in.

  • Never read credentials through the page. No evaluate that extracts a password, token,

out everywhere. No flow here needs it.

  • Never clear cookies, cache or site data. It is the user's real profile; clearing it logs them

capture on every tab it drives, which is a known automation signal; CloakBrowser through connectCloakProfile() is the stealth path.

  • focused: false by default. The browser belongs to someone who is probably using it.
  • One short, named session per task, always stopped.
  • Bot-scored or WAF targets go to the owned engine. The attached engine's daemon enables console

Where the rest lives

More skills from code-yeongyu/oh-my-openagent

  • Aast-grepSearches and rewrites code by AST shape across 25 languages. Use when the target is a syntax pattern (every call/class/import shaped like X, a codemod, a YAML rule) rather than literal text; for plain strings, comments, or filenames, use rg.
  • Acodex-qaQA the omo Codex Light edition (lazycodex / packages/omo-codex) itself, in strict isolation so ONLY our plugin is exercised, never the user's real ~/.codex. The first-party method drives the real `codex app-server` against an isolated CODEX_HOME plus a LOCAL mock model (no real API call), and proves a plugin hook fired by asserting hook/started + hook/completed notifications. Also: isolated install verification, per-component hook probes, a tmux TUI smoke, and runtime log observation (RUST_LOG / logs SQLite / /debug-config). Ships tested helper scripts each with a --self-test. Use whenever someone changes anything under packages/omo-codex or wants to QA, smoke-test, verify, or debug the Codex plugin, its hooks/components, the installer/config.toml, the app-server flow, or the Codex TUI. Triggers: codex qa, qa codex, codex-qa, test codex plugin, verify codex hook, codex app-server, lazycodex qa, isolated CODEX_HOME, prove codex hook fired, codex tui test.
  • Acoding-agent-sessionsFinds, reads, and reconstructs coding-agent sessions across Codex, Claude, OpenCode, OMO/Senpi, and other local agent logs. Use when asked to find or search past sessions, transcripts, or subagent runs, or to recover what an earlier session did.
  • Acomment-checkerUse when Codex needs to understand or respond to automatic comment-checker feedback emitted after an edit-like PostToolUse hook.
  • Adag-libraryStores a DAG definition once and re-runs it by name, instead of pasting the definition into every run. Use when the user wants to save a DAG, run a saved one, or schedule the same multi-agent graph repeatedly.
  • Ddata-scientistProcesses and analyzes data with resident-kernel engines (DuckDB, Polars) and one-shot tools. Use for CSV/parquet/JSON analysis, group-by/join/aggregation, time series, distributions, cleaning, or plotting a dataset.
  • AdebuggingRuns a hypothesis-driven debugging loop across any language or binary, escalating to orthogonal oracle angles and locking the fix with a failing test. Use for crashes, silent failures, hangs, wrong responses, memory leaks, async misbehavior, or reverse engineering.
  • Adev-browserBrowser automation with persistent page state. Use when users ask to navigate websites, fill forms, take screenshots, extract web data, test web apps, or automate browser workflows. Trigger phrases include "go to [url]", "click on", "fill out the form", "take a screenshot", "scrape", "automate", "test the website", "log into", or any browser interaction request.
  • CfrontendBuilds, styles, and polishes web UI and UX. Use for any frontend, page, component, styling, layout, animation, or visual-quality task, or when asked to make an interface look or feel a certain way.
  • AfrontendBuilds, styles, and polishes web UI and UX. Use for any frontend, page, component, styling, layout, animation, or visual-quality task, or when asked to make an interface look or feel a certain way.
  • Aget-unpublished-changesCompare HEAD with the latest published npm versions and list all unpublished changes by release layer. Triggers: unpublished changes, changelog, what changed, whats new.
  • Agit-masterHandles git work: atomic commits, rebase, squash, blame, bisect, reflog, and history questions. Use whenever a task needs a commit or a git-history investigation; skip for ordinary code edits.

All agent skills → · MCP servers