Mmcp.market

debug skill

by nanocoai·nanocoai/nanoclaw·31k stars·MIT

Debug container agent issues. Use when things aren't working, container fails, authentication problems, or to understand how the container system works. Covers logs, session DBs, mounts, and common issues.

A100/100content scan

Is the debug skill safe?

Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.

No findings.

Install the debug skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/nanocoai/nanoclaw.git /tmp/nanoclaw
mkdir -p ~/.claude/skills
cp -r /tmp/nanoclaw/.claude/skills/debug ~/.claude/skills/debug
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

NanoClaw Container Debugging

This guide covers debugging the containerized agent execution system.

Architecture Overview

The host is a single Node process that orchestrates per-session agent containers. The two session DBs are the sole IO surface between host and container — there is no IPC, no file watcher, and no stdin piping.

Host (Node)                                Container (Bun, Linux VM)
──────────────────────────────────────────────────────────────────────
src/container-runner.ts                    container/agent-runner/src/
    │                                          │
    │ spawns one container per session          │ polls inbound.db for work,
    │ with the session folder mounted          │ calls the agent provider,
    │ at /workspace                            │ writes replies to outbound.db
    │                                          │
    ├── data/v2-sessions/<group>/<session>/ ──> /workspace
    │     ├── inbound.db   (host writes, container reads RO)
    │     ├── outbound.db  (container writes, host reads)
    │     └── .heartbeat   (container touches → /workspace/.heartbeat)
    ├── groups/<folder> ─────────────────────> /workspace/agent  (cwd)
    ├── <group>/.claude-shared ──────────────> /home/node/.claude
    └── agent-runner src + skills ───────────> /app/src, /app/skills

Message flow: host writes a row to inbound.db (messagesin) and wakes the container; the container's poll loop picks it up, runs the agent, and writes the reply to outbound.db (messagesout); the host's delivery poll reads messages_out and sends it through the channel adapter. See docs/db.md and docs/db-session.md for the full two-DB model.

Container identity: the container runs as user node with HOME=/home/node. Per-group Claude state (settings, session history) lives in /.claude-shared on the host, mounted to /home/node/.claude.

Log Locations

Containers run with --rm, so the container's own filesystem is gone after it exits. The host streams container stderr into logs/nanoclaw.log at debug level, tagged with container=; raise the log level (below) to see it. If the agent silently failed inside an exited container, there is no persistent in-container log — reconstruct from the session DBs and the host log.

Enabling Debug Logging

Set LOG_LEVEL=debug for verbose output, including streamed container stderr:

# For development
LOG_LEVEL=debug pnpm run dev

# For launchd service (macOS), add to plist EnvironmentVariables:
<key>LOG_LEVEL</key>
<string>debug</string>
# For systemd service (Linux), add to unit [Service] section:
# Environment=LOG_LEVEL=debug

Debug level shows full mount configurations, the container spawn command, and streamed container stderr lines.

Inspecting Session DBs

The two session DBs are where the message flow lives. Use the in-tree query wrapper (it goes through the better-sqlite3 dep that setup already installs, avoiding a dependency on the sqlite3 CLI):

# List sessions and their agent group / messaging group from the central DB
pnpm exec tsx scripts/q.ts data/v2.db "SELECT id, agent_group_id, messaging_group_id, status, container_status, last_active FROM sessions"

# Or via the admin CLI
ncl sessions list

# Did the message reach the container? (inbound.db, host writes / container reads)
pnpm exec tsx scripts/q.ts data/v2-sessions/<group>/<session>/inbound.db \
  "SELECT seq, kind, status, timestamp FROM messages_in ORDER BY seq DESC LIMIT 10"

# Did the agent produce a reply? (outbound.db, container writes / host reads)
pnpm exec tsx scripts/q.ts data/v2-sessions/<group>/<session>/outbound.db \
  "SELECT seq, kind, timestamp FROM messages_out ORDER BY seq DESC LIMIT 10"

# Container-side processing status for each inbound message
pnpm exec tsx scripts/q.ts data/v2-sessions/<group>/<session>/outbound.db \
  "SELECT message_id, status, status_changed FROM processing_ack ORDER BY status_changed DESC LIMIT 10"

Reading the flow:

  • messagesin has the message but no matching messagesout → the container never produced a reply (check processing_ack, then logs/nanoclaw.log for spawn/exit and container stderr).
  • messages_out has a reply but the user never received it → a delivery problem (see issue 1 below).
  • messages_in is empty → routing never reached this session (check the router log lines and the central wiring with ncl wirings list).

Common Issues

1. "No adapter for channel type" / Messages silently lost (null platformmessageid)

Symptom: The bot stops replying. logs/nanoclaw.error.log shows repeated:

WARN No adapter for channel type channelType="telegram"
WARN No adapter for channel type channelType="signal"

The main log shows "Message delivered" entries with platformMsgId=undefined — meaning the delivery poll ran, found no adapter, and marked the message delivered without sending it.

Root cause: two NanoClaw service instances running simultaneously.

When a second service instance is active with a stale binary, it has no channel adapters registered. Its delivery poll races the working instance and wins — marking outbound messages delivered without ever sending them.

Diagnosis:

# Check for duplicate running instances
ps aux | grep 'nanoclaw/dist/index.js' | grep -v grep

# Check which services are active (Linux)
systemctl --user list-units 'nanoclaw*' --all

# Confirm channel adapters registered by the current process
grep "Channel adapter started" logs/nanoclaw.log | tail -10

Fix:

  1. Identify which service has the correct binary and EnvironmentFile (the one whose log shows the expected channels — e.g. signal, telegram, cli — all started).
  2. Stop and disable the stale duplicate service:
systemctl --user stop nanoclaw.service   # or whichever is the old one
   systemctl --user disable nanoclaw.service
  1. If the remaining service unit is missing EnvironmentFile, add it:
# Edit the service unit — add this line under [Service]:
   # EnvironmentFile=/home/[user]/nanoclaw/.env
   systemctl --user daemon-reload
   systemctl --user restart nanoclaw-v2-<id>.service
  1. Verify only one instance runs: ps aux | grep nanoclaw/dist/index.js | grep -v grep

Messages marked delivered with a null platformmessageid are not automatically retried. Ask the user to resend.

2. Container exits immediately / agent produces no reply

A spawned container that exits without writing to outbound.db shows up in logs/nanoclaw.log as a Container exited line with a non-zero code, often preceded by streamed container= stderr (at debug level).

Authentication errors: secrets are injected per request by the OneCLI gateway — none are passed in env vars or chat context. A 401 from an API whose credential is in the vault usually means the agent is in selective secret mode and that secret was never assigned:

onecli agents list                                        # check secretMode
onecli agents set-secret-mode --id <agent-id> --mode all  # inject all matching secrets

If the gateway itself is unreachable, the container runner refuses to spawn (OneCLI gateway not applied — refusing to spawn container without credentials in the host log). Confirm the gateway is up at http://127.0.0.1:10254.

MCP server failures: a misconfigured MCP server can abort the agent run. Look for MCP initialization errors in the streamed container stderr (LOG_LEVEL=debug).

3. Mount Issues

Session and group folders are bind-mounted into the container. To see the resolved mounts for a spawn, run with LOG_LEVEL=debug and read the spawn command in logs/nanoclaw.log, or grep the mount targets directly:

grep -n "containerPath" src/container-runner.ts

Expected mount targets inside the container:

/workspace            ← session folder (inbound.db, outbound.db, .heartbeat, inbox/, outbox/)
/workspace/agent      ← agent group folder (cwd; CLAUDE.md, skills, working files)
/home/node/.claude    ← per-group .claude-shared (Claude state, settings, history)
/app/src              ← agent-runner source (read-only)
/app/skills           ← container skills (read-only)

To inspect what a fresh container sees:

docker run --rm --entrypoint /bin/bash nanoclaw-agent:latest -c 'whoami; ls -la /workspace/ /app/'

All of /workspace/ and /app/ should be owned by node. Use :ro on a -v mount for read-only.

4. Heartbeat / stale-session detection

Liveness is a file touch on /workspace/.heartbeat (host path: data/v2-sessions///.heartbeat), not a DB write. The host sweep reads its mtime plus the processing_ack claim age to decide whether a container is alive or stale. A session stuck "processing" with a stale .heartbeat mtime means the container died mid-run:

stat -f '%Sm' data/v2-sessions/<group>/<session>/.heartbeat   # macOS
stat -c '%y'  data/v2-sessions/<group>/<session>/.heartbeat   # Linux

Container CLI (ncl) inside a session

The agent reaches the central DB from inside the container via ncl, which uses the session DB transport (container/agent-runner/src/cli/ncl.ts). On the host, ncl connects over a Unix socket (src/cli/socket-server.ts). If ncl calls fail from inside a container, check the agent group's cli_scope in its container config:

ncl groups config get --id <group-id>   # look at cli_scope: disabled | group | global

disabled rejects every cli_request; group scopes the agent to its own group's groups/sessions/destinations/members; global is unrestricted.

Restarting a session's container

# Restart all containers for an agent group
ncl groups restart --id <group-id>

# Restart and rebuild the image first (after package/Dockerfile changes)
ncl groups restart --id <group-id> --rebuild

# Restart and wake immediately with a message
ncl groups restart --id <group-id> --message "on_wake test"

Without --message, the container comes back on the next user message. From inside a container, --id is auto-filled and only the calling session restarts.

Manual Container Probes

The container's entry point is exec bun run /app/src/index.ts; it talks only to the mounted session DBs, so there is no JSON to pipe in. To probe the image directly:

More skills from nanocoai/nanoclaw

  • Aadd-anydocAdd local office-document-to-Markdown conversion to NanoClaw agent containers with the pinned Firecrawl AnyDoc CLI. Use when agents need to read attached Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, or text-based PDF files without uploading them to a hosted parser.
  • Aadd-atomic-chat-toolAdd Atomic Chat MCP server so the container agent can call local models served by the Atomic Chat desktop app via its OpenAI-compatible API.
  • Fadd-clidashAdd clidash — a zero-dependency, read-only web dashboard that derives its tabs and tables at runtime from any CLI that lists resources as JSON. Ships pre-wired for NanoClaw's ncl CLI (agent groups, sessions, channels, users, roles), plus message-activity charts, a log tail, and a read-only file viewer for group skills/CLAUDE.md/profiles.
  • Aadd-codexUse Codex (OpenAI's codex app-server) as a full agent provider — planning, tool orchestration, MCP tools, server-side history, session resume — alongside or instead of Claude. ChatGPT subscription or OpenAI API key, vault-only via the selected gateway. Per-group via `ncl groups config update --provider codex`. Distinct from using OpenAI as an MCP tool (where Claude remains the planner).
  • Aadd-dashboardAdd a monitoring dashboard to NanoClaw. Installs @nanoco/nanoclaw-dashboard and a pusher that sends periodic JSON snapshots.
  • Aadd-deltachatAdd DeltaChat channel integration via @deltachat/stdio-rpc-server. Native adapter — no Chat SDK bridge. Email-based messaging with end-to-end encryption.
  • Cadd-dialAdd Dial channel integration — a real phone number for SMS and AI voice calls via the Dial platform (getdial.ai). Native adapter — no Chat SDK bridge.
  • Aadd-dial-numberAdd another phone number to an existing Dial channel — a second (or third) public line for the agent, so one NanoClaw install answers SMS and AI voice calls on multiple numbers. Use when Dial is already installed and the operator wants an additional number (e.g. a personal line plus a support line). Requires the Dial channel to already be installed (see /add-dial).
  • Aadd-dial-toolGive chosen NanoClaw agents a real phone number as a container tool — the `dial` CLI baked into the agent image plus OneCLI credential injection for api.getdial.ai, scoped per agent, so the agents you pick can send SMS, place AI voice calls, and receive verification codes from inside the sandbox. Independent of the Dial channel; idempotent; re-run to change which agents may use it. Use when the user wants agents to text, call, or run `dial …` from a chat, without wiring Dial as a messaging channel.
  • Aadd-discordAdd Discord bot channel integration via Chat SDK.
  • Aadd-emacsAdd Emacs as a channel. Opens an interactive chat buffer and org-mode integration so you can talk to NanoClaw from within Emacs (Doom, Spacemacs, or vanilla). Local HTTP bridge — no bot token or external service needed.
  • Aadd-gchatAdd Google Chat channel integration via Chat SDK.

All agent skills → · MCP servers