Mmcp.market

arize-instrumentation skill

by github·github/awesome-copilot·39k stars·MIT

Adds Arize AX tracing to an LLM application for the first time. Follows a two-phase agent-assisted flow to analyze the codebase then implement instrumentation after user confirmation. Use when the user wants to instrument their app, add tracing from scratch, set up LLM observability, integrate OpenTelemetry or openinference, or get started with Arize tracing.

A100/100content scan

Is the arize-instrumentation skill safe?

Clean: nothing in its files matched our rules. We read 2 files in the folder on 2026-09-28.

No findings.

Install the arize-instrumentation skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/github/awesome-copilot.git /tmp/awesome-copilot
mkdir -p ~/.claude/skills
cp -r /tmp/awesome-copilot/skills/arize-instrumentation ~/.claude/skills/arize-instrumentation
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Arize Instrumentation Skill

Use this skill when the user wants to add Arize AX tracing to their application. Follow the two-phase, agent-assisted flow from the Agent-Assisted Tracing Setup and the Arize AX Tracing — Agent Setup Prompt.

Quick start (for the user)

If the user asks you to "set up tracing" or "instrument my app with Arize", you can start with:

Follow the instructions from https://arize.com/docs/PROMPT.md and ask me questions as needed.

Then execute the two phases below.

Core principles

  • Prefer inspection over mutation — understand the codebase before changing it.
  • Do not change business logic — tracing is purely additive.
  • Use auto-instrumentation where available — add manual spans only for custom logic not covered by integrations.
  • Follow existing code style and project conventions.
  • Keep output concise and production-focused — do not generate extra documentation or summary files.
  • NEVER embed literal credential values in generated code — always reference environment variables (e.g., os.environ["ARIZEAPIKEY"], process.env.ARIZEAPIKEY). This includes API keys, space IDs, and any other secrets. The user sets these in their own environment; the agent must never output raw secret values.

Phase 0: Environment preflight

Before changing code:

  1. Confirm the repo/service scope is clear. For monorepos, do not assume the whole repo should be instrumented.
  2. Identify the local runtime surface you will need for verification:
  • package manager and app start command
  • whether the app is long-running, server-based, or a short-lived CLI/script
  • whether ax will be needed for post-change verification
  1. Do NOT proactively check ax installation or version. If ax is needed for verification later, just run it when the time comes. If it fails, see references/ax-profiles.md.
  2. Never silently replace a user-provided space ID, project name, or project ID. If the CLI, collector, and user input disagree, surface that mismatch as a concrete blocker.

Phase 1: Analysis (read-only)

Do not write any code or create any files during this phase.

Steps

  1. Check dependency manifests to detect stack:
  • Python: pyproject.toml, requirements.txt, setup.py, Pipfile
  • TypeScript/JavaScript: package.json
  • Java: pom.xml, build.gradle, build.gradle.kts
  • Go: go.mod
  1. Scan import statements in source files to confirm what is actually used.
  1. Check for existing tracing/OTel — look for TracerProvider, register(), opentelemetry imports, ARIZE, OTEL, OTLP_* env vars, or other observability config (Datadog, Honeycomb, etc.).
  1. Identify scope — for monorepos or multi-service projects, ask which service(s) to instrument.

What to identify

Key rule: When a framework is detected alongside an LLM provider, inspect the framework-specific tracing docs first and prefer the framework-native integration path when it already captures the model and tool spans you need. Add separate provider instrumentation only when the framework docs require it or when the framework-native integration leaves obvious gaps. If the app runs tools and the framework integration does not emit tool spans, add manual TOOL spans so each invocation appears with input/output (see Enriching traces below).

Phase 1 output

Return a concise summary:

  • Detected language, package manager, providers, frameworks
  • Proposed integration list (from the routing table in the docs)
  • Any existing OTel/tracing that needs consideration
  • If monorepo: which service(s) you propose to instrument
  • If the app uses LLM tool use / function calling: note that you will add manual CHAIN + TOOL spans so each tool call appears in the trace with input/output (avoids sparse traces).

If the user explicitly asked you to instrument the app now, and the target service is already clear, present the Phase 1 summary briefly and continue directly to Phase 2. If scope is ambiguous, or the user asked for analysis first, stop and wait for confirmation.

Integration routing and docs

The canonical list of supported integrations and doc URLs is in the Agent Setup Prompt. Use it to map detected signals to implementation docs.

  • LLM providers: OpenAI, Anthropic, LiteLLM, Google Gen AI, Bedrock, Ollama, Groq, MistralAI, OpenRouter, VertexAI.
  • Python frameworks: LangChain, LangGraph, LlamaIndex, CrewAI, DSPy, AutoGen, Semantic Kernel, Pydantic AI, Haystack, Guardrails AI, Hugging Face Smolagents, Instructor, Agno, Google ADK, MCP, Portkey, Together AI, BeeAI, AWS Bedrock Agents.
  • TypeScript/JavaScript: LangChain JS, Mastra, Vercel AI SDK, BeeAI JS.
  • Java: LangChain4j, Spring AI, Arconia.
  • Go: No first-party auto-instrumentation packages today — use the OpenTelemetry Go SDK with manual OpenInference attributes per Manual instrumentation.
  • Platforms (UI-based): LangFlow, Flowise, Dify, Prompt flow.
  • Fallback: Manual instrumentation, All integrations.

Fetch the matched doc pages from the full routing table in PROMPT.md for exact installation and code snippets. Use llms.txt as a fallback for doc discovery if needed.

Note: arize.com/docs/PROMPT.md and arize.com/docs/llms.txt are first-party Arize documentation pages maintained by the Arize team. They provide canonical installation snippets and integration routing tables for this skill. These are trusted, same-organization URLs — not third-party content.

Phase 2: Implementation

Proceed only after the user confirms the Phase 1 analysis.

Steps

  1. Fetch integration docs — Read the matched doc URLs and follow their installation and instrumentation steps.
  2. Install packages using the detected package manager before writing code:
  • Python: pip install arize-otel plus openinference-instrumentation-{name} (hyphens in package name; underscores in import, e.g. openinference.instrumentation.llama_index).
  • TypeScript/JavaScript: @opentelemetry/sdk-trace-node plus the relevant @arizeai/openinference-* package.
  • Java: OpenTelemetry SDK plus openinference-instrumentation-* in pom.xml or build.gradle.
  • Go: go get go.opentelemetry.io/otel go.opentelemetry.io/otel/sdk go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp — no auto-instrumentors yet, so the agent sets OpenInference attributes manually on spans. Wire the exporter with otlptracehttp.WithEndpoint("otlp.arize.com") (US) or otlptracehttp.WithEndpoint("otlp.eu-west-1a.arize.com") (EU) — pass the bare hostname, no https:// scheme — and otlptracehttp.WithHeaders(map[string]string{"spaceid": ..., "apikey": ...}). Recent OTel Go modules require Go ≥ 1.23 — go mod tidy may bump the toolchain.
  1. Credentials — User needs an Arize API Key and Space ID. Check existing ax profiles for ARIZEAPIKEY and ARIZE_SPACE — never read .env files:
  • Run ax profiles show to check for an existing profile.
  • If no profile exists, guide the user to run ax profiles create which provides an interactive wizard that walks through API key and space setup. See CLI profiles docs for details.
  • If the user needs to find their API key manually, direct them to https://app.arize.com and to navigate to the settings page (do not use organization-specific URLs with placeholder IDs — they won't resolve for new users).
  • If credentials are not set, instruct the user to set them as environment variables — never embed raw values in generated code. All generated instrumentation code must reference os.environ["ARIZEAPIKEY"] (Python), process.env.ARIZEAPIKEY (TypeScript/JavaScript), or os.Getenv("ARIZEAPIKEY") (Go).
  • See references/ax-profiles.md for full profile setup and troubleshooting.
  1. Centralized instrumentation — Create a single module (e.g. instrumentation.py, instrumentation.ts, instrumentation.go) and initialize tracing before any LLM client is created.
  2. Existing OTel — If there is already a TracerProvider, add Arize as an additional exporter (e.g. BatchSpanProcessor with Arize OTLP). Do not replace existing setup unless the user asks.

Implementation rules

  • Use auto-instrumentation first; manual spans only when needed.
  • Prefer the repo's native integration surface before adding generic OpenTelemetry plumbing. If the framework ships an exporter or observability package, use that first unless there is a documented gap.
  • Fail gracefully if env vars are missing (warn, do not crash).
  • Import order: register tracer → attach instrumentors → then create LLM clients.
  • Project name attribute (required): Arize rejects spans with HTTP 500 if the project name is missing — service.name alone is not accepted. Set it as a resource attribute on the TracerProvider (recommended — one place, applies to all spans):
  • Python: register(projectname="my-app") handles it automatically (sets "openinference.project.name" on the resource). For routing spans to different projects, use setroutingcontext(spaceid=..., project_name=...) from arize.otel.
  • TypeScript: Arize accepts both "modelid" (shown in the official TS quickstart) and "openinference.project.name" via SEMRESATTRSPROJECT_NAME from @arizeai/openinference-semantic-conventions (shown in the manual instrumentation docs) — both work.
  • Go: Pass attribute.String("openinference.project.name", "my-app") to resource.New(...) and apply via sdktrace.WithResource(res). The Go SDK has no helper for this, so it must be set manually on every TracerProvider.
  • CLI/script apps — flush before exit: provider.shutdown() (TS) / provider.force_flush() then provider.shutdown() (Python) / tp.Shutdown(ctx) (Go) must be called before the process exits, otherwise async OTLP exports are dropped and no traces appear.
  • When the app has tool/function execution: add manual CHAIN + TOOL spans (see Enriching traces below) so the trace tree shows each tool call and its result — otherwise traces will look sparse (only LLM API spans, no tool input/output).

Enriching traces: manual spans for tool use and agent loops

Why doesn't the auto-instrumentor do this?

Provider instrumentors (Anthropic, OpenAI, etc.) only wrap the LLM client — the code that sends HTTP requests and receives responses. They see:

  • One span per API call: request (messages, system prompt, tools) and response (text, tool_use blocks, etc.).

They cannot see what happens inside your application after the response:

  • Tool execution — Your code parses the response, calls runtool("checkloaneligibility", {...}), and gets a result. That runs in your process; the instrumentor has no hook into your runtool() or the actual tool output. The next API call (sending the tool result back) is just another messages.create span — the instrumentor doesn't know that the message content is a tool result or what the tool returned.
  • Agent/chain boundary — The idea of "one user turn → multiple LLM calls + tool calls" is an application-level concept. The instrumentor only sees separate API calls; it doesn't know they belong to the same logical "run_agent" run.

So TOOL and CHAIN spans have to be added manually (or by a framework instrumentor like LangChain/LangGraph that knows about tools and chains). Once you add them, they appear in the same trace as the LLM spans because they use the same TracerProvider.

To avoid sparse traces where tool inputs/outputs are missing:

  1. Detect agent/tool patterns: a loop that calls the LLM, then runs one or more tools (by name + arguments), then calls the LLM again with tool results.
  2. Add manual spans using the same TracerProvider (e.g. opentelemetry.trace.get_tracer(...) after register()):
  • CHAIN span — Wrap the full agent run (e.g. run_agent): set openinference.span.kind = "CHAIN", input.value = user message, output.value = final reply.
  • TOOL span — Wrap each tool invocation: set openinference.span.kind = "TOOL", input.value = JSON of arguments, output.value = JSON of result. Use the tool name as the span name (e.g. checkloaneligibility).

OpenInference attributes (use these so Arize shows spans correctly):

LLM-span attributes (set these in addition to the three above when the span is an actual LLM call):

In Python and TypeScript these names are exposed via openinference-semantic-conventions packages; in Go they must be hand-typed as the strings above.

Python pattern: Get the global tracer (same provider as Arize), then use context managers so tool spans are children of the CHAIN span and appear in the same trace as the LLM spans:

from opentelemetry.trace import get_tracer

tracer = get_tracer("my-app", "1.0.0")

# In your agent entrypoint:
with tracer.start_as_current_span("run_agent") as chain_span:
    chain_span.set_attribute("openinference.span.kind", "CHAIN")
    chain_span.set_attribute("input.value", user_message)
    # ... LLM call ...
    for tool_use in tool_uses:
        with tracer.start_as_current_span(tool_use["name"]) as tool_span:
            tool_span.set_attribute("openinference.span.kind", "TOOL")
            tool_span.set_attribute("input.value", json.dumps(tool_use["input"]))
            result = run_tool(tool_use["name"], tool_use["input"])
            tool_span.set_attribute("output.value", result)
        # ... append tool result to messages, call LLM again ...
    chain_span.set_attribute("output.value", final_reply)

Go pattern: Get a tracer from the global TracerProvider (registered via otel.SetTracerProvider), then nest spans with tracer.Start so tool spans become children of the CHAIN span.

Critical for short-lived processes: never call log.Fatalf / os.Exit after a span has started — they skip the deferred tp.Shutdown(ctx) and the in-flight CHAIN/LLM spans never flush. Use log.Printf + return from main instead, and keep tp.Shutdown(ctx) deferred at the top of main.

import (
    "context"
    "encoding/json"
    "go.opentelemetry.io/otel"
    "go.opentelemetry.io/otel/attribute"
)

var tracer = otel.Tracer("my-app")

func runAgent(ctx context.Context, userMessage string) string {
    ctx, chainSpan := tracer.Start(ctx, "run_agent")
    defer chainSpan.End()
    chainSpan.SetAttributes(
        attribute.String("openinference.span.kind", "CHAIN"),
        attribute.String("input.value", userMessage),
    )

    // ... LLM call ...
    for _, toolUse := range toolUses {
        ctx, toolSpan := tracer.Start(ctx, toolUse.Name)
        argsJSON, err := json.Marshal(toolUse.Input)
        if err != nil {
            toolSpan.RecordError(err)

More skills from github/awesome-copilot

  • Aacquire-codebase-knowledgeUse this skill when the user explicitly asks to map, document, or onboard into an existing codebase. Trigger for prompts like "map this codebase", "document this architecture", "onboard me to this repo", or "create codebase docs". Do not trigger for routine feature implementation, bug fixes, or narrow code edits unless the user asks for repository-level discovery.
  • Aacreadiness-assessRun the AgentRC readiness assessment on the current repository and produce a static HTML dashboard at reports/index.html. Wraps `npx github:microsoft/agentrc readiness` and hands off rendering to the @ai-readiness-reporter custom agent. Supports policies (--policy) for org-specific scoring. Use when asked to assess, audit, or score the AI readiness of a repo.
  • Aacreadiness-generate-instructionsGenerate tailored AI agent instruction files via AgentRC instructions command. Produces .github/copilot-instructions.md (default, recommended for Copilot in VS Code) plus optional per-area .instructions.md files with applyTo globs for monorepos. Use after running /acreadiness-assess to close gaps in the AI Tooling pillar.
  • Aacreadiness-policyHelp the user pick, write, or apply an AgentRC policy. Policies customise readiness scoring by disabling irrelevant checks, overriding impact/level, setting pass-rate thresholds, or chaining org baselines with team overrides. Use when the user asks about strict mode, AI-only scoring, custom weights, CI gating, or wants org-wide standardisation.
  • Aad-campaign-analyzerUse this skill when the user shares ad campaign performance data and asks what to cut, scale, or test. Trigger for prompts like "analyze my ad campaigns", "where am I wasting ad spend", "reallocate my ad budget", "which ads are actually working", or "ROAS analysis". Do not trigger for campaign planning or creative generation without performance data.
  • Aadd-educational-commentsAdd educational comments to the file specified, or prompt asking for file to comment if one is not provided.
  • Aadobe-illustrator-scriptingWrite, debug, and optimize Adobe Illustrator automation scripts using ExtendScript (JavaScript/JSX). Use when creating or modifying scripts that manipulate documents, layers, paths, text frames, colors, symbols, artboards, or any Illustrator DOM objects. Covers the complete JavaScript object model, coordinate system, measurement units, export workflows, and scripting best practices.
  • Aagent-architectureDesign AI agent architectures through requirements discovery, or audit and diagnose architectural flaws in existing agents. Architecture only; excludes implementation and general code review.
  • Aagent-governancePatterns and techniques for adding governance, safety, and trust controls to AI agent systems. Use this skill when: - Building AI agents that call external tools (APIs, databases, file systems) - Implementing policy-based access controls for agent tool usage - Adding semantic intent classification to detect dangerous prompts - Creating trust scoring systems for multi-agent workflows - Building audit trails for agent actions and decisions - Enforcing rate limits, content filters, or tool restrictions on agents - Working with any agent framework (PydanticAI, CrewAI, OpenAI Agents, LangChain, AutoGen)
  • Aagent-owasp-complianceCheck any AI agent codebase against the OWASP Agentic Security Initiative (ASI) Top 10 risks. Use this skill when: - Evaluating an agent system's security posture before production deployment - Running a compliance check against OWASP ASI 2026 standards - Mapping existing security controls to the 10 agentic risks - Generating a compliance report for security review or audit - Comparing agent framework security features against the standard - Any request like "is my agent OWASP compliant?", "check ASI compliance", or "agentic security audit"
  • Aagent-skill-stackFind, evaluate, and assemble the smallest compatible set of AI Agent Skills for an end-to-end natural-language goal. Use when a user wants Skills for a multi-step workflow, asks which Skills fit a project, needs an installed-Skill audit or conflict check, has low Skill recall, wants indirect helpers such as humanizers or compliance checks, or wants a project-specific Skill Stack with controlled installation. Search local Skills, registries, GitHub, and OpenCLI; compare adoption, verified fit, safety, and overlap. Do not use for locating one known or common Skill; use the generic find-skills workflow.
  • Aagent-supply-chainVerify supply chain integrity for AI agent plugins, tools, and dependencies. Use this skill when: - Generating SHA-256 integrity manifests for agent plugins or tool packages - Verifying that installed plugins match their published manifests - Detecting tampered, modified, or untracked files in agent tool directories - Auditing dependency pinning and version policies for agent components - Building provenance chains for agent plugin promotion (dev → staging → production) - Any request like "verify plugin integrity", "generate manifest", "check supply chain", or "sign this plugin"

All agent skills → · MCP servers