agent-evaluation-reporting skill
Use when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.
Is the agent-evaluation-reporting skill safe?
Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.
No findings.
Install the agent-evaluation-reporting skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git /tmp/agentic-awesome-skills mkdir -p ~/.claude/skills cp -r /tmp/agentic-awesome-skills/plugins/agentic-awesome-skills-claude/skills/agent-evaluation-reporting ~/.claude/skills/agent-evaluation-reporting
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Agent Evaluation Reporting
Overview
Turn raw agent evaluation runs into a decision-ready report without hiding failures or overstating capability. Keep outcome populations, denominators, latency populations, and experiment conditions explicit so readers can reproduce every headline number.
When to Use This Skill
- Use when reporting benchmark, regression, pilot, or production evaluation runs for an AI agent.
- Use when autonomous and human-assisted completions appear in the same result set.
- Use when failures, timeouts, infrastructure-invalid runs, retries, or partial results affect the denominator.
- Use when comparing two agents, prompts, harnesses, or releases and deciding whether the comparison is valid.
How It Works
Step 1: Freeze the comparison contract
Record the task set and sampling, model and provider, prompt or policy version, tool and harness versions, evaluator rubric, timeout and retry policy, token or cost budget, environment, and human-intervention policy. Assign the configuration a stable label or digest.
If a material condition differs between runs, mark the comparison as non-equivalent. Report a directional observation only; do not claim that the changed agent caused the difference.
Step 2: Build a mutually exclusive outcome ledger
Classify every scheduled attempt exactly once:
Preserve attempt ID, task ID or seed, retry index, parent attempt ID, configuration label, outcome, intervention count, duration, cost, evaluator evidence, and invalid reason when available. Never silently drop invalid or retried runs.
Also build a unique-task rollup. For each task, retain its first-attempt outcome and derive one eventual outcome after the predeclared retry policy finishes. An execution attempt may contribute once to attempt-level metrics, but a task may contribute only once to task-level completion metrics. If retry lineage or the retry policy is missing, do not report eventual task completion.
Step 3: Lock each metric to a denominator
Let Nall be all execution attempts, including retries, and Neval = Nall - Ninvalid be evaluable attempts. Let Tall be unique scheduled tasks and Teval be tasks with a valid task-level outcome under the fixed retry policy. Report counts beside every rate.
autonomous attempt success = N_autonomous / N_eval
assisted attempt success = N_assisted / N_eval
attempt non-completion = (N_failure + N_timeout) / N_eval
invalid-attempt rate = N_invalid / N_all
first-attempt completion = T_first_attempt_completed / T_all
eventual task completion = T_eventual_completed / T_eval
operational task delivery = T_eventual_completed / T_allLabel attempt-level and unique-task metrics explicitly; never call an attempt-level rate workflow completion. Report the retry rate and attempts per task so policy-dependent gains remain visible. Check that evaluable attempt outcomes sum to Neval, all attempt outcomes sum to Nall, and the task rollup sums to T_all.
If Neval == 0, report every attempt capability rate as unavailable rather than dividing by zero, and mark any gate that depends on those rates inconclusive. Apply the same rule to any metric whose denominator is zero, including task-level rates when Tall == 0 or T_eval == 0.
Step 4: Keep latency and cost populations honest
Report autonomous-completion latency, assisted end-to-end latency, and failure time-to-terminal separately. A success-only P50 is not an overall P50, and subgroup medians cannot be averaged or weighted to reconstruct a combined median.
Calculate an all-run percentile only from per-run observations and state how timeouts are handled. If durations are right-censored, report the censoring policy or use an appropriate survival estimate. Apply the same population labels to token and cost metrics.
Step 5: Quantify uncertainty and comparability
For stochastic evaluations, show sample size and an interval or repeated-run distribution beside headline rates. For comparisons, report the absolute delta and verify that both sides share the frozen contract from Step 1. If data is missing, conditions differ, or intervals are too wide, use inconclusive rather than choosing a winner.
Step 6: Map evidence to predeclared decision gates
Define readiness gates before reading the result, such as minimum autonomous success, maximum timeout rate, zero critical safety violations, and latency or cost bounds. Return pass, fail, or inconclusive for each gate.
Do not infer production readiness from a success rate alone. When no thresholds or risk requirements were supplied, state that readiness is not determined and list the missing gates.
Example
For 120 unique tasks with one attempt each, including 12 infrastructure-invalid runs, 48 autonomous successes, 24 assisted successes, 20 failures, and 16 timeouts:
Evaluable attempts: 108 / 120
Autonomous success: 48 / 108 = 44.4%
Assisted success: 24 / 108 = 22.2%
Attempt non-completion: 36 / 108 = 33.3%
First-attempt completion: 72 / 120 = 60.0%
Eventual task completion: 72 / 108 = 66.7% (no retries)
Operational task delivery: 72 / 120 = 60.0%
Infrastructure-invalid: 12 / 120 = 10.0%
Overall latency P50: unavailable from subgroup aggregates
Readiness: inconclusive until gates are declaredBest Practices
- Report counts, formulas, denominator labels, and exclusions together.
- Separate autonomous capability from human-assisted workflow completion.
- Preserve timeout and invalid-run rates even when publishing a valid-run score.
- Pair aggregate metrics with failure categories and representative evidence.
- Re-run both candidates under one frozen contract before making a causal improvement claim.
Limitations
- This skill structures and interprets supplied evaluation evidence; it does not validate the evaluator or recreate missing run records.
- Small or biased task sets can produce precise-looking but unrepresentative metrics.
- Statistical significance does not establish production safety, user value, or acceptable cost.
- Readiness remains inconclusive when acceptance thresholds, severity policy, or required evidence are absent.
Security & Safety Notes
- Redact credentials, private prompts, personal data, and sensitive tool output from reports while retaining stable evidence references.
- Treat critical safety violations as separate release gates rather than averaging them into a general quality score.
Common Pitfalls
Solution: Publish separate autonomous, assisted, and workflow-completion rates.
- Problem: Assisted completions are presented as autonomous success.
Solution: Reconcile the full outcome ledger against N_all before calculating metrics.
- Problem: Timeouts or invalid runs disappear from the denominator.
Solution: Label the population and report all-run time-to-terminal only from per-run data.
- Problem: A faster success-only P50 is presented as a faster system.
Solution: Apply predeclared gates or return inconclusive.
- Problem: A release verdict is improvised after seeing results.
Related Skills
- @agent-evaluation - Design behavioral tests, benchmarks, and reliability evaluations.
- @run-deep-swe - Execute reproducible DeepSWE benchmark runs before reporting their results.
More skills from sickn33/agentic-awesome-skills
- A00-andruia-consultantArquitecto de Soluciones Principal y Consultor Tecnológico de Andru.ia. Diagnostica y traza la hoja de ruta óptima para proyectos de IA en español.
- F007Security audit, hardening, threat modeling (STRIDE/PASTA), Red/Blue Team, OWASP checks, code review, incident response, and infrastructure security for any project.
- A10-andruia-skill-smithIngeniero de Sistemas de Andru.ia. Diseña, redacta y despliega nuevas habilidades (skills) dentro del repositorio siguiendo el Estándar de Diamante.
- A20-andruia-niche-intelligenceEstratega de Inteligencia de Dominio de Andru.ia. Analiza el nicho específico de un proyecto para inyectar conocimientos, regulaciones y estándares únicos del sector. Actívalo tras definir el nicho.
- A2slides-ppt-generatorAI-powered presentation generation via the 2slides API — create slides from text, match a reference image style, summarize documents into decks, add AI voice narration, and export pages/audio. Use for any \"make slides\", \"create a deck\", or \"slides from this document\" request.
- A3d-web-experienceExpert in building 3D experiences for the web - Three.js, React
- Aab-test-setupUse when designing an A/B or split test: define the hypothesis, control and variants, estimate sample size, verify tracking, and predeclare metrics and stopping rules.
- Aab-testingWhen the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.
- Aacceptance-orchestratorUse when a coding task should be driven end-to-end from issue intake through implementation, review, deployment, and acceptance verification with minimal human re-intervention.
- Aaccess-reviewConduct periodic access reviews and certifications. Implement access
- Aaccessibility-compliance-accessibility-auditYou are an accessibility expert specializing in WCAG compliance, inclusive design, and assistive technology compatibility. Conduct audits, identify barriers, and provide remediation guidance.
- Aaccesslint-auditFind and fix WCAG 2.2 accessibility issues. Two modes — report (sweep a codebase or page, produce a prioritized written report, no edits) and fix (audit→edit→verify loop on a target). Prefers direct-CDP live-DOM auditing; falls back to a browser-MCP composition or HTML-string audits.