security-and-hardening skill
Hardens code against vulnerabilities. Use when auditing an input handler for vulnerabilities, when handling user input, authentication, data storage, or external integrations, or when checking a login flow is safe against the OWASP Top Ten. Use when building any feature that accepts untrusted data, manages user sessions, or interacts with third-party services. Use when auditing dependencies for known vulnerabilities, triaging package-manager audit findings, or assessing supply-chain risk in a new package. Use when personal data or privacy compliance (GDPR, CCPA) is involved.
Is the security-and-hardening skill safe?
Clean: nothing in its files matched our rules. We read 2 files in the folder on 2026-09-28.
No findings.
Install the security-and-hardening skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/addyosmani/agent-skills.git /tmp/agent-skills mkdir -p ~/.claude/skills cp -r /tmp/agent-skills/skills/security-and-hardening ~/.claude/skills/security-and-hardening
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Security and Hardening
Overview
Security-first development practices for web applications. Treat every external input as hostile, every secret as sacred, and every authorization check as mandatory. Security isn't a phase — it's a constraint on every line of code that touches user data, authentication, or external systems.
When to Use
- Building anything that accepts user input
- Implementing authentication or authorization
- Storing or transmitting sensitive data
- Integrating with external APIs or services
- Adding file uploads, webhooks, or callbacks
- Handling payment or PII data
Process: Threat Model First
Controls bolted on without a threat model are guesses. Before hardening, spend five minutes thinking like an attacker:
- Map the trust boundaries. Where does untrusted data cross into your system? HTTP requests, form fields, file uploads, webhooks, third-party APIs, message queues, and LLM output — plus the local values that look internal because the OS handed them to you: another process's command line or environment, filenames on a shared volume, a path in a job payload. Trust follows who wrote a value, not which channel delivered it. Every boundary is attack surface.
- Name the assets. What's worth stealing or breaking? Credentials, PII, payment data, admin actions, money movement.
- Run STRIDE over each boundary — a quick lens, not a ceremony:
- Write abuse cases next to use cases. For each feature, ask "how would I misuse this?" — then make that your first test.
If you can't name the trust boundaries for a feature, you're not ready to secure it. This is OWASP A04: Insecure Design — most breaches begin in design, not code.
The Three-Tier Boundary System
Always Do (No Exceptions)
- Validate all external input at the system boundary (API routes, form handlers)
- Parameterize all database queries — never concatenate user input into SQL
- Encode output to prevent XSS (use framework auto-escaping, don't bypass it)
- Use HTTPS for all external communication
- Hash passwords with bcrypt/scrypt/argon2 (never store plaintext)
- Set security headers (CSP, HSTS, X-Frame-Options, X-Content-Type-Options)
- Use httpOnly, secure, sameSite cookies for sessions
- Run the detected package manager's native audit against the committed lockfile before every release
Ask First (Requires Human Approval)
- Adding new authentication flows or changing auth logic
- Storing new categories of sensitive data (PII, payment info)
- Adding new external service integrations
- Changing CORS configuration
- Adding file upload handlers
- Modifying rate limiting or throttling
- Granting elevated permissions or roles
Never Do
- Never commit secrets to version control (API keys, passwords, tokens)
- Never log sensitive data (passwords, tokens, full credit card numbers)
- Never trust client-side validation as a security boundary
- Never disable security headers for convenience
- Never use eval() or innerHTML with user-provided data
- Never store sessions in client-accessible storage (localStorage for auth tokens)
- Never expose stack traces or internal error details to users
Hardening Controls
The rules below are the workflow; a concrete implementation of each lives in references/hardening-patterns.md. Open the section you need when you reach that code, not before.
Injection, XSS, and access control
- Parameterize every query. Never build SQL, NoSQL, or shell commands from input strings.
- Encode output through the framework's auto-escaping. If raw HTML is unavoidable, sanitize with an allowlist sanitizer first.
- Check authorization on every request, not just authentication: the authenticated user must own, or be permitted on, the specific resource (A01, IDOR).
Patterns: Injection, XSS, Access control.
Authentication and sessions
- Hash passwords with bcrypt (≥12 rounds), scrypt, or argon2. The session secret comes from the environment, never from code.
- Session cookies are httpOnly, secure, and sameSite: 'lax' or 'strict' (the CSRF defense; 'none' sends the cookie on cross-site requests), with a bounded maxAge.
Pattern: Authentication.
Headers, CORS, and responses
- Security headers on every response (helmet or the framework equivalent); CSP starts from default-src 'self' and is tightened, not loosened.
- CORS restricted to an explicit origin list from configuration. Never * with credentials.
- Strip sensitive fields (passwordHash, reset tokens) before any response. Error bodies are generic; internals go to server logs only.
Patterns: Misconfiguration, Sensitive data exposure.
Input validation and uploads
- Validate at the boundary with a schema: allowlisted shape, lengths, enums, formats. Reject with 422 and structured details; downstream code uses only the parsed, typed value.
- Uploads: allowlist MIME types, cap size, verify content (magic bytes) when it matters. The extension proves nothing.
Patterns: Schema validation, File upload.
Server-side fetches (SSRF)
Any URL the user influences — webhooks, import-from-URL, image proxies, link previews — can be aimed at internal services. Allowlist scheme and host, resolve all DNS records and reject any private or reserved address (loopback, link-local 169.254.169.254, private, unique-local, for IPv4 and IPv6), and forbid redirects. That check still has a DNS-rebinding TOCTOU gap: for high-risk surfaces, pin the resolved IP or put a filtering agent in front.
Pattern: SSRF.
Destructive operations on derived paths
A delete, move, or overwrite is only as safe as the value naming its target, and trust follows who wrote that value, not which channel delivered it: another process's command line is as attacker-controlled as a form field. A shape check proves well-formedness, not authorization. Before the call, require all three: the resolved target (symlinks resolved) sits under an allowlisted root; it is at least one level below that root; and it carries ownership evidence read before the operation. On refusal, log the rejected target and stop; never fall back to a broader default path.
Why the check is weaker than it reads (marker self-attestation, check/use races): Destructive paths. Worked code: ../../references/security-checklist.md.
Rate limiting
Limit the API generally and auth endpoints strictly (about 10 attempts per 15 minutes). Once more than one process serves traffic, in-memory counters silently become max × instances, or never fire on serverless: back the limiter with a shared store.
Pattern: Rate limiting.
Secrets
Secrets come from the environment. .env.example is committed with placeholders; real .env files and key material are gitignored; grep the staged diff before committing. A secret that reaches a remote is compromised the moment it lands: rotate it first, then purge history.**
Pattern: Secrets management.
Dependencies and supply chain
- Find the installation boundary and manager. Use the workspace root that owns the lockfile, or an independent nested project only when it is outside that workspace. Corroborate packageManager (when present), the lockfile, and CI; stop on disagreement or competing lockfiles. Pin the manager version.
- Block dependency scripts before first execution. Bootstrap with scripts disabled or a documented fail-closed policy, inspect the pending script source, approve only the minimum, commit the policy, then verify with a clean frozen/immutable install. Never blanket-approve.
- Run the native audit against the committed lockfile before every release. Triage critical/high by reachability (runtime, build, test, deploy paths) and fix availability. Never apply forced remediation (npm audit fix --force or equivalent) automatically, since forced fixes may cross declared dependency ranges; preview, read changelogs, test each upgrade. Document every deferral with a reason and a review date.
- Audits only match known advisories. They do not catch a newly malicious or typosquatted package (cross-env vs crossenv). Review new dependencies, lockfile diffs, and script-policy changes together: ownership, maintenance, release age, provenance, transitive graph. Verify registry signatures where supported (npm audit signatures, pnpm audit signatures) and treat their absence as a signal to investigate, not automatic proof of compromise (A06, LLM03).
Triage decision tree: Dependency audit triage. Manager matrix and install-script gate: ../../references/security-checklist.md.
Personal data and privacy
Hardening asks "can an attacker read it?" Privacy asks "should we hold it at all, and for how long?" The cheapest data to protect, breach, and comply over is the data you never collected; treat personal data as a liability to minimize.
- Classify fields as you add them (non-personal, PII, sensitive) and handle each class accordingly. You cannot protect, or honor a deletion request for, data you cannot find.
- Collect only against a stated purpose. "Might be useful later" is latent breach scope, not a purpose. Keep PII out of telemetry (the observability-and-instrumentation skill makes the same point from the ops side).
- Set retention up front, then actually delete. Every personal-data store needs a TTL and a working deletion path, including backups, caches, search indexes, and analytics copies.
- Support the data-subject rights your jurisdiction requires (GDPR, CCPA, and kin): export, correct, delete. Design the schema so a user's data is findable and erasable, not smeared irreversibly across systems.
- Consent gates collection and third-party sharing, and is auditable. Sending PII to an analytics, ad, or LLM vendor is sharing; the vendor needs a data-processing agreement. Make region a configurable policy, not a hardcoded assumption.
Classification table: Data classification. A privacy incident starts the breach-notification clock; run the postmortem with the debugging-and-error-recovery skill.
AI / LLM features
Calling an LLM — chatbots, summarizers, agents, RAG — adds a new attack surface; map it to the OWASP Top 10 for LLM Applications (2025):
- Model output is untrusted input (LLM05). Never into eval, SQL, a shell, innerHTML, or a file path; parse defensively, validate against a schema, then encode.
- Prompts can be hijacked (LLM01). Untrusted text in the context — a user message, a fetched page, a PDF — can carry instructions. The system prompt is not a security boundary; enforce permissions in code.
- Keep secrets, other tenants' data, and the full system prompt out of the context window (LLM02, LLM07); scope tool permissions, validate every tool argument, and confirm destructive actions (LLM06); cap tokens, request rate, and recursion depth (LLM10); partition RAG embeddings per tenant and validate documents before indexing (LLM08).
Pattern: LLM output handling.
Review Checklist
Before sign-off, walk ../../references/security-checklist.md: it covers authentication, authorization, input, data protection and privacy, headers and CORS, dependencies and supply chain, AI/LLM, and error handling, plus the OWASP quick-reference tables.
Common Rationalizations
Red Flags
- User input passed directly to database queries, shell commands, or HTML rendering
- A delete, move, or overwrite whose target comes from a payload, a config value, or another process's command line, guarded only by a shape check on the path
- Secrets in source code or commit history
- API endpoints without authentication or authorization checks
- Missing CORS configuration or wildcard (*) origins
- No rate limiting on authentication endpoints, or an in-memory limiter in front of more than one instance
- Stack traces or internal errors exposed to users
- Dependencies with known critical vulnerabilities, competing lockfiles at one installation boundary, non-reproducible installs, or blanket-approved scripts
- Server fetches user-supplied URLs without an allowlist (SSRF)
- LLM/model output passed into a query, the DOM, a shell, or eval
- Secrets, PII, or the full system prompt placed inside an LLM context window
- Personal data collected with no stated purpose, retention limit, or deletion path
Verification
More skills from addyosmani/agent-skills
- Aapi-and-interface-designGuides stable API and interface design. Use when designing APIs, module boundaries, or any public interface. Use when creating REST or GraphQL endpoints, defining type contracts between modules, or establishing boundaries between frontend and backend.
- Cbrowser-testing-with-devtoolsTests in real browsers via Chrome DevTools MCP. Use when building or debugging anything that runs in a browser. Use when you need to inspect the DOM, capture console errors, analyze network requests, profile performance, or verify visual output with real runtime data. Requires the chrome-devtools MCP server to be configured.
- Aci-cd-and-automationAutomates CI/CD pipeline setup. Use when setting up or modifying build and deployment pipelines. Use when you need to automate quality gates, configure test runners in CI, or establish deployment strategies.
- Acode-review-and-qualityConducts multi-axis code review. Use before merging any change. Use when reviewing code written by yourself, another agent, or a human. Use when you need to assess code quality across multiple dimensions before it enters the main branch. Use when asked to review a diff or a pull request, even when the diff is pasted inline.
- Acode-simplificationSimplifies code for clarity. Use when refactoring code for clarity without changing behavior. Use when code works but is harder to read, maintain, or extend than it should be. Use when reviewing code that has accumulated unnecessary complexity.
- Aconstraint-driven-developmentEstablishes a project's quality bar as a written contract and stops agents quietly lowering it. Interviews the user on which dimensions matter, supplies sane default thresholds when they have no number in mind, records everything in CONSTRAINTS.md, and watches the diff for a weakened bar — new @ts-ignore or eslint-disable suppressions, skipped or deleted tests, assertions stripped out, unimplemented stubs, thresholds edited down. Use when no quality bar is written down, when the user says "set up constraints" or "define our standards", when the user wants dimensions they care about — accessibility, web performance, coverage — set up as enforced constraints, when an agent keeps silencing checks or skipping tests to get to green, when you need a coverage or performance threshold and don't know what number to pick, or when an agent writes more code than anyone will read.
- Acontext-engineeringOptimizes agent context setup. Use when starting a new session, when agent output quality degrades, when switching between tasks, or when you need to configure rules files and context for a project.
- Adebugging-and-error-recoveryGuides systematic root-cause debugging. Use when tests fail, builds break, something that worked yesterday broke, behavior doesn't match expectations, or you encounter any unexpected error. Use when you need to figure out what broke and why — a systematic approach to finding and fixing the root cause rather than guessing.
- Adeprecation-and-migrationManages deprecation and migration. Use when removing old systems, APIs, or features. Use when migrating users from one implementation to another. Use when migrating a database schema in production, such as renaming or dropping a column without downtime (expand/contract). Use when deciding whether to maintain or sunset existing code.
- Adocumentation-and-adrsRecords decisions and documentation. Use when you need to document an architecture decision (ADR) or the reasoning behind a design choice, when changing public APIs, shipping features, or when you need to record context that future engineers and agents will need to understand the codebase.
- Adoubt-driven-developmentSubjects every non-trivial decision to a fresh-context adversarial review before it stands. Use when you want every assumption cross-examined before proceeding, when stress-testing a plan for hidden failure modes, when correctness matters more than speed, when working in unfamiliar code, when stakes are high (production auth, security-sensitive logic, a high-stakes migration, irreversible operations), or any time a confident output would be cheaper to verify now than to debug later.
- Afrontend-ui-engineeringBuilds production-quality, accessible, responsive user-facing UIs. Use when building or modifying interfaces and pages, creating components, implementing layouts, meeting WCAG accessibility requirements, managing state, or when the output needs to look and feel production-quality rather than AI-generated.