browser-automation skill
统一自动化入口。覆盖浏览器自动化(Playwright)和 Windows 桌面应用自动化(OpenReverse)。 浏览器场景:打开网页、点击、填表、爬取、截图、自动化登录、渗透页面交互。 桌面场景:操作 IDA/x64dbg 等 GUI 工具、Windows UI Automation、视觉驱动交互、桌面应用网络抓包。 触发关键词:浏览器自动化、桌面自动化、打开网页、填表、爬取、截图、自动化登录、Playwright、agent-browser、headless、OpenReverse、UIA、CUA、桌面操作、Windows 自动化。
Is the browser-automation skill safe?
Clean: nothing in its files matched our rules. We read 3 files in the folder on 2026-09-28.
No findings.
Install the browser-automation skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/zhaoxuya520/reverse-skill.git /tmp/reverse-skill mkdir -p ~/.claude/skills cp -r /tmp/reverse-skill/skills/browser-automation ~/.claude/skills/browser-automation
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
自动化操作 (Desktop & Browser Automation)
ACTION REQUIRED(读完后立刻执行)
- NOW:确认当前任务是否命中本 skill 的适用范围
- NOW:读取 ../tool-index.md,校验工具可用性和实际路径
- NEXT:缺工具时调用 bootstrap,不要猜路径
- ACT:进入"工作流"第一步并执行,不要停在确认状态
适用范围
当任务属于以下场景时使用本 skill:
浏览器场景(Playwright / agent-browser)
- 打开网页并操作页面元素(点击、填表、提交)
- 爬取页面内容或截图
- 自动化登录流程
- 渗透测试中与 Web 页面交互(提交 payload、触发 XSS)
- 验证码页面的自动化处理
- 批量表单提交
桌面应用场景(OpenReverse)
- 操作 Windows 桌面应用(IDA Pro、x64dbg、Wireshark 等)
- 需要视觉驱动交互(CUA 模式)
- 需要结构化 UI 操作(UIA 模式)
- 桌面应用的网络流量观察(内置 mitmproxy)
- 自动化逆向工具的 GUI 操作
- 黑盒测试桌面软件
与其他工具的分工
简单判断:
- 目标是网页 → Playwright
- 目标是 Windows 桌面应用 → OpenReverse
- 两者都需要 → 组合使用
Part 1: 浏览器自动化(Playwright / agent-browser)
核心工作流
# 1. 打开页面
agent-browser open <url>
# 2. 获取可交互元素(返回 @e1, @e2... 引用)
agent-browser snapshot -i
# 3. 用引用操作元素
agent-browser click @e1
agent-browser fill @e2 "text"
# 4. 完成后关闭
agent-browser close命令参考
# 导航
agent-browser open <url>
agent-browser close
# 页面快照
agent-browser snapshot # 完整无障碍树
agent-browser snapshot -i # 仅可交互元素(推荐)
# 交互操作
agent-browser click @e1
agent-browser fill @e2 "text"
agent-browser type @e2 "text"
agent-browser press Enter
agent-browser scroll down 500
# 获取信息
agent-browser get text @e1
agent-browser get title
agent-browser get url
# 等待
agent-browser wait @e1
agent-browser wait 2000
agent-browser wait --load networkidle注意事项
- 必须执行 agent-browser close,否则进程泄漏
- 操作前先 snapshot,不要猜元素引用
- 提交表单后用 wait --load networkidle 等页面稳定
Part 2: 桌面应用自动化(OpenReverse)
概述
OpenReverse 是面向 AI Agent 的桌面交互与证据采集框架,支持:
- UIA 模式:Windows UI Automation,结构化桌面控件操作
- CUA 模式:视觉驱动交互(Computer Use Agent),适合复杂 GUI
- 网络观察:内置 mitmproxy 代理 + 本地抓取
交互模式选择
网络观察模式
安装与配置
# 1. Clone 项目
git clone https://github.com/zhexulong/openreverse.git
cd openreverse
# 2. 安装依赖
npm install
# 3. 接入 Agent 宿主(Claude Code / Codex / Zed)
npm run init:agents -- --target=all /path/to/project
# 4. 安装 CUA runtime(如果需要视觉驱动模式)
npm run install:cua-runtime
npm run doctor:cua-runtime
# 5. 安装网络观察依赖(如果需要抓包)
npm run install:mitmproxy
npm run doctor:network常见组合
逆向场景示例
场景:自动化操作 IDA Pro 进行批量分析
1. 用 OpenReverse CUA 模式打开 IDA Pro
2. 自动加载目标二进制
3. 等待分析完成
4. 通过 UI 操作导出函数列表
5. 同时用 network lane 观察 IDA 的网络行为(如 Lumina 请求)场景:自动化操作 x64dbg 调试
1. 用 OpenReverse UIA 模式启动 x64dbg
2. 加载目标程序
3. 设置断点
4. 运行并观察寄存器/内存变化
5. 截图保存证据按需自举(On-Demand Bootstrap)
自动化能力边界
自举触发
- 浏览器操作缺 Playwright → 自动 bootstrap
- 桌面操作需要 OpenReverse → 引导用户手动安装(给出完整步骤)
OpenReverse 手动安装引导
如果 AI 检测到需要桌面应用自动化但 OpenReverse 未安装:
⚠️ **需要 OpenReverse 进行桌面应用自动化**
**安装步骤**:
1. `git clone https://github.com/zhexulong/openreverse.git`
2. `cd openreverse && npm install`
3. `npm run init:agents -- --target=all <你的项目路径>`
4. 如需视觉模式:`npm run install:cua-runtime`
5. 如需网络观察:`npm run install:mitmproxy`
**验证**:`npm run doctor:cua-runtime` 和 `npm run doctor:network`路由上下文
上游入口: skills/SKILL.md(总控)、routing.md 适用场景: 任何需要自动化操作浏览器或桌面应用的任务 下游出口:
- 抓到的请求需要分析 → anything-analyzer 或 js-reverse
- 需要 JS 调试/Hook → jshookmcp
- 需要还原签名算法 → js-reverse
- 桌面应用是逆向工具 → ida-reverse/
同级关联模块: js-reverse(浏览器操作后可能需要分析 JS)、ida-reverse(OpenReverse 可以自动化操作 IDA GUI)
任务完成自检(声称完成前 MUST 通过)
- [ ] 我是否执行了工作流中的每一步(而不是只阅读)?
- [ ] 我是否基于 tool-index 使用了真实工具路径?
- [ ] 我是否产出了可复现证据(命令/脚本/截图/报告)?
- [ ] 我是否完成并回写了 RULES 要求的 Checklist 项?
More skills from zhaoxuya520/reverse-skill
- Fapi-securityUse for authorized security assessment of REST, GraphQL, WebSocket, or SOAP APIs, including discovery, authentication, authorization, rate-limit, and CI/CD testing.
- Capk-reverse在 CLI 环境下做 Android APK 逆向时使用。适用于 APK 解包、Java 反编译、smali 修改、重打包、Frida 动态 Hook,以及按需切换到 so/native 分析。优先使用本机已安装的 jadx、apktool、frida、adb、ida-reverse、radare2。
- Cattack-chainUse for authorized multi-stage attack-path planning and orchestration when a task spans reconnaissance, initial access, privilege escalation, lateral movement, or impact assessment. Route single-stage tasks directly to their specialist skill.
- Abinary-diff跨版本符号迁移与二进制差分。当你有旧版本的符号/逆向结果,需要快速迁移到新版本时使用。 适用场景:内核缺 PDB 用旧版符号推导、程序更新后批量迁移函数名、应用更新后快速定位新偏移。 核心方法:用 LLM 做结构化差异比对,程序化输入输出,成本极低(200 函数 ~1 元)。 触发关键词:符号迁移、bindiff、跨版本、PDB 缺失、函数偏移迁移、symbol migration、binary diff、版本对比。
- Abinary-ninja-reverseUse for authorized binary analysis in Binary Ninja, including HLIL/MLIL/LLIL inspection, strings/imports/exports, cross-references, types, patch review, Python API automation, and optional Binary Ninja MCP or localhost HTTP integration.
- Abrowser-extension-reverseUse for authorized reverse engineering of browser extensions (Chrome/Firefox) including manifest analysis, background workers, and extension-based credential or traffic logic recovery.
- Acase-reviewReviews a reverse-skill case package for scope readiness, Evidence to Finding to Path traceability, work item coverage, timeline references, and optional artifact hash integrity before report handoff.
- Acloud-k8sUse for authorized cloud, container, and Kubernetes security assessment including metadata SSRF, IAM misconfig, container escape paths, and cluster RBAC review.
- Acode-auditUse for authorized source-code security review and SAST workflows including Semgrep, CodeQL patterns, dangerous API hunting, and fix verification.
- Acompetition-ad-certificate-abuseInternal downstream skill for ctf-sandbox-orchestrator. CTF-sandbox workflow for AD CS, certificate templates, enrollment rights, EKUs, SAN controls, PKINIT, certificate mapping, and cert-based privilege paths. Use when the user asks about ESC-style abuse, certificate templates, enrollment agents, EKUs, SAN or subject controls, smartcard or PKINIT logon, CA policy, or how an issued cert turns into accepted privilege. Use only after `$ctf-sandbox-orchestrator` has already established sandbox assumptions and routed here.
- Acompetition-agent-cloudInternal downstream skill for ctf-sandbox-orchestrator. CTF-sandbox workflow for AI-agent, prompt-injection, MCP or toolchain, cloud, container, CI/CD, and supply-chain challenges. Use when the user asks to analyze prompt-to-tool flows, retrieval poisoning, mounted secrets, deployment drift, runtime-vs-manifest mismatches, registry provenance, or CI-produced artifacts under sandbox assumptions. Use only after `$ctf-sandbox-orchestrator` has already established sandbox assumptions and routed here.
- Acompetition-android-hookingInternal downstream skill for ctf-sandbox-orchestrator. CTF-sandbox workflow for Android APK hooking, Frida tracing, request-signing recovery, SSL pinning bypass, JNI boundary inspection, and app trust-boundary analysis. Use when the user asks to hook an APK, inspect signer logic, trace Java or native boundaries, bypass pinning or root checks, inspect shared prefs or app databases, or replay accepted mobile requests. Use only after `$ctf-sandbox-orchestrator` has already established sandbox assumptions and routed here.