evaluation-methodology skill
PluginEval quality methodology, covering dimensions, rubrics, and scoring formulas. Use this skill when understanding how plugin quality is measured, when interpreting a low score on a specific dimension, when deciding how to improve a skill's triggering accuracy or orchestration fitness, when setting score thresholds for your marketplace, or when explaining quality badges to external partners like Neon.
Is the evaluation-methodology skill safe?
Clean: nothing in its files matched our rules. We read 3 files in the folder on 2026-09-28.
No findings.
Install the evaluation-methodology skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/wshobson/agents.git /tmp/agents mkdir -p ~/.claude/skills cp -r /tmp/agents/plugins/plugin-eval/skills/evaluation-methodology ~/.claude/skills/evaluation-methodology
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Evaluation methodology
PluginEval scores a skill or a plugin from 0 to 100 by combining up to three layers. The static layer is a lint. It's fast and deterministic, and it's useful for checking structure. The LLM judge and Monte Carlo layers are experimental, not validated against human labels, so treat their numbers as rough signals. For the trace-based eval program, see evals/README.md at the repository root.
The judge rubric anchors are in references/rubrics.md. Fixes for each anti-pattern and tips for each dimension are in references/improving-scores.md.
Evaluation depths
The label names the depth that ran, not a check against human judgment. A plugin directory gets the static layer only, whatever depth you ask for.
Static layer (lint)
The static layer reads SKILL.md and makes no model calls. It computes seven sub-scores, and the first six feed composite dimensions:
- frontmatterquality (feeds triggeringaccuracy)
- orchestrationwiring (feeds orchestrationfitness)
- progressivedisclosure, structuralcompleteness, tokenefficiency, ecosystemcoherence
- harness_portability
harness_portability maps to no dimension, so it doesn't change a skill's composite score. It does carry 6% of the static layer's own score. A plugin's score is built from that layer score, so portability findings can lower a plugin's score a little. Its findings are not counted as anti-patterns.
The static layer flags six anti-patterns: OVERCONSTRAINED, EMPTYDESCRIPTION, MISSINGTRIGGER, BLOATEDSKILL, ORPHANREFERENCE, and DEADCROSS_REF. Each flag cuts the score by 5%, down to a floor of 50%:
penalty = max(0.5, 1.0 - 0.05 * anti_pattern_count)The report prints a severity for each flag, but the penalty counts flags and ignores severity.
LLM judge layer (experimental)
The judge layer makes four model calls and returns four holistic scores from 0 to 1:
trigger and 5 that should not. It predicts the outcome for each prompt and reports its own F1. Nothing checks those predictions against real triggering.
- triggering_accuracy: Haiku reads the description and writes 10 test prompts, 5 that should
- orchestrationfitness and scopecalibration: Sonnet rates the skill on a five-point rubric.
- output_quality: Sonnet imagines three tasks and rates the output it expects.
The three Sonnet calls see only the first 3,000 characters of SKILL.md. Only one judge runs, because nothing reads the judges setting.
Monte Carlo layer (experimental)
Haiku writes 15 prompts that should trigger the skill, and the layer repeats them to reach 50 runs (100 at thorough). Each run sends the SKILL.md text and one prompt to the model, and the layer records four measures:
answered, not whether the skill should have fired.
- Activation rate is the share of runs with any non-empty reply, so it shows whether the model
- Output quality is reply length divided by 500, capped at 1.0.
- Failure rate is the share of runs that errored.
- Token efficiency is 1 - median_tokens / 8000.
Every prompt is one that should trigger, so the layer never checks that the skill stays out of unrelated requests. The layer's JSON includes Wilson, bootstrap, and Clopper-Pearson intervals for its own measures. The composite cilower and ciupper fields are always null.
Composite score
First, for each dimension, the engine blends the layer scores that exist, and it renormalizes the blend weights over those layers. Second, it sums the weighted dimension scores, and it renormalizes the dimension weights over the dimensions that have a score. Third, it multiplies the sum by 100 and by the anti-pattern penalty.
A cell reads "none" when its layer produces no score for the dimension, even where LAYERBLENDS lists a weight. No layer produces codetemplate_quality, so it's always unmeasured. For a plugin directory, the composite is the static layer's mean score across the plugin's skills and agents, times 100, times the penalty for the plugin's anti-pattern count.
Badges and grades
Badges come from the composite score alone. Platinum needs at least 90, Gold at least 80, Silver at least 70, and Bronze at least 60. Badge.from_scores accepts an Elo rating, but no command computes one. Plugin-level badges, including the ones in the weekly CI report, come from the static layer alone. Skill-level badges at standard depth or deeper also include the experimental layers.
Each measured dimension gets a letter grade on the 0 to 100 scale, from A+ at 97 down to D- at 60, and F below 60.
Usage
plugin-eval score ./path/to/skill --depth quick # static lint only
plugin-eval score ./path/to/skill # static and judge
plugin-eval certify ./path/to/skill # deep depth
plugin-eval compare ./skill-a ./skill-b # quick depth by default
plugin-eval score ./path/to/skill --depth quick --output json --threshold 70At standard depth or deeper, score, certify, and compare print a note on stderr that the judge and Monte Carlo layers are experimental. For a plugin directory, the CLI prints a warning that only the static layer runs instead. With --threshold, the command exits with code 1 when the composite is below the value. plugin-eval init writes a corpus index, but no other command reads it.
Examples
An abridged example of the JSON output follows. Scripts can read composite.score from it:
{
"layers": [{"layer": "static", "score": 0.75, "sub_scores": {}, "anti_patterns": []}],
"composite": {"score": 76.9, "ci_lower": null, "ci_upper": null, "badge": "silver",
"confidence_label": "Estimated", "dimensions": []},
"elo": null
}Troubleshooting
when" to the description, followed by several comma-separated contexts.
- When a score drops after you add content, check layers[0].anti_patterns in the JSON.
- If triggering_accuracy is low at quick depth, add a trigger phrase such as "Use this skill
time. Use the static layer for comparisons you want to repeat.
- Judge scores change between runs, because the model writes new test prompts and tasks each
uv sync --extra llm.
- If stderr says the judge could not measure some dimensions, install the LLM extra with
Related
The eval-judge agent scores the four judge dimensions inside Claude Code, and the eval-orchestrator agent runs the CLI and merges the results.
More skills from wshobson/agents
- Aairflow-dag-patternsBuild production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
- Aangular-migrationMigrate from AngularJS to Angular using hybrid mode, incremental component rewriting, and dependency injection updates. Use when upgrading AngularJS applications, planning framework migrations, or modernizing legacy Angular code.
- Aanti-reversing-techniquesUnderstand anti-reversing, obfuscation, and protection techniques encountered during software analysis. Use this skill when analyzing malware evasion techniques, when implementing anti-debugging protections for CTF challenges, when reverse engineering packed binaries, or when building security research tools that need to detect virtualized environments.
- Aapi-design-principlesMaster REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers. Use when designing new APIs, reviewing API specifications, or establishing API design standards.
- Aarchitecture-decision-recordsWrite and maintain Architecture Decision Records (ADRs) following best practices for technical decision documentation. Use when documenting significant technical decisions, reviewing past architectural choices, or establishing decision processes.
- Aarchitecture-patternsImplement proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design. Use this skill when designing clean architecture for a new microservice, when refactoring a monolith to use bounded contexts, when implementing hexagonal or onion architecture patterns, or when debugging dependency cycles between application layers.
- Aasync-python-patternsMaster Python asyncio, concurrent programming, and async/await patterns for high-performance applications. Use when building async APIs, concurrent systems, or I/O-bound applications requiring non-blocking operations.
- Aauth-implementation-patternsMaster authentication and authorization patterns including JWT, OAuth2, session management, and RBAC to build secure, scalable access control systems. Use when implementing auth systems, securing APIs, or debugging security issues.
- Aavoid-ai-writingAudit and rewrite prose so it stops reading as machine-generated. Use this skill when asked to remove AI-isms, clean up AI writing, edit a draft for AI tells, audit a README, changelog, release note, PR description, or blog post for machine-sounding prose, or make text sound less like AI. Supports a detect-only mode, a rewrite mode, and an edit-in-place mode, with optional voice and context profiles.
- Abacktesting-frameworksBuild robust backtesting systems for trading strategies with proper handling of look-ahead bias, survivorship bias, and transaction costs. Use when developing trading algorithms, validating strategies, or building backtesting infrastructure.
- Abazel-build-optimizationOptimize Bazel builds for large-scale monorepos. Use when configuring Bazel, implementing remote execution, or optimizing build performance for enterprise codebases.
- Abefore-you-buildPre-build product and feature risk review for founders, product managers, and AI-assisted builders. Use this skill when the user is about to build a landing page, MVP, SaaS product, internal tool, agent workflow, or major feature and needs to check demand, positioning, monetization, retention, trust, distribution, and adoption risk before implementation starts.