kubectl-investigator skill
Investigate a live or recent incident in a Kubernetes cluster. Anchor the window, bisect the change surface (rollouts, ConfigMaps/Secrets, RBAC, HPA/cluster changes, CronJobs), classify against four reference failure paths (OOM, DNS, cascading-failure, deploy-correlator), confirm the hypothesis with three independent signals, quantify blast radius, and propose mitigation before root cause. Use whenever an agent is asked "what is breaking in the cluster right now", "why did this pod/Deployment just page", "did the rollout cause Z", or to triage an active Kubernetes incident. Vendor-neutral by default (works with kubectl, kube-state-metrics, and whatever telemetry you have); an opt-in Anyshift integration is documented separately.
Is the kubectl-investigator skill safe?
Clean: nothing in its files matched our rules. We read 55 files in the folder on 2026-09-28.
No findings.
Install the kubectl-investigator skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/anyshift-io/sre-skills.git /tmp/sre-skills mkdir -p ~/.claude/skills cp -r /tmp/sre-skills/skills/kubectl-investigator ~/.claude/skills/kubectl-investigator
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
kubectl-investigator
Methodology skill for investigating a live or recent incident on Kubernetes. Produces a timeline, a ranked set of hypotheses, a blast-radius estimate, and a recommended mitigation. Hands off cleanly to postmortem-author once the incident is mitigated.
Scope: workloads running on Kubernetes (Deployments, StatefulSets, DaemonSets, Jobs/CronJobs) and the cluster primitives around them (Services, Ingress, CoreDNS, ConfigMaps/Secrets, RBAC, HPA, nodes). External dependencies (third-party APIs, partner TLS endpoints, managed databases) are in scope only as seen from a Kubernetes workload — the methodology investigates the cluster-side symptom and the in-cluster change surface.
When to invoke
- A PrometheusRule / Alertmanager alert just fired on a workload and the agent needs to triage before paging a human.
- A user asks "what is breaking in the cluster right now" or "why did Deployment X just page".
- A kubectl rollout / Helm release / Argo CD sync went out in the last hour and a metric moved; need to know whether they are linked.
- Pods are crash-looping, OOMKilled, or Pending, or customer impact is reported with no alert yet; need to find the failing surface.
The methodology, in order
The order matters. Skipping a step produces confident wrong answers.
1. Anchor the window
Lock two timestamps before doing anything else:
- T0: the trigger timestamp. Apply this order:
- If an alert is provided as the trigger, T0 = alert fire time. Use this verbatim. Do not substitute an earlier "first error in logs / first OOMKilled event" timestamp just because one exists; the alert fire time is the agreed-upon coordination point for the incident.
- If a customer report is the trigger, T0 = report timestamp.
- If neither exists (operator-initiated investigation, "pods slow all morning"), T0 = earliest unambiguous signal in the available telemetry (first OOMKilled event, first SERVFAIL, first error-rate inflection), and mark T0 as ambiguous (see below).
- Tnow: current time, or the timestamp the investigation was triggered.
Every later signal is filtered to [T0 - 15min, Tnow]. The 15-minute lead-in catches changes that landed just before the symptom surfaced (a rollout's pods take time to roll, an HPA scale-down takes time to bite).
If T0 is ambiguous (operator-triggered with no alert, or "slow all morning"-class reports), the methodology's recommended mitigation in step 6 must begin with "re-run the investigation with a widened window" before any irreversible action. The change identified within the original narrow window is likely incomplete; the actual causal change may sit outside it. Do not silently round and do not skip the re-run step.
2. Bisect the change surface
Pull every change event that overlaps the window. On Kubernetes the change surface is:
- Workload rollouts: kubectl rollout / kubectl apply / kubectl set image, new image tags, new ReplicaSets, Helm releases, Argo CD / Flux syncs.
- Cluster / capacity changes: node-pool scaling, node cordon/drain, resource requests/limits edits, HPA / VPA changes, PodDisruptionBudget edits, PV/PVC/StorageClass changes.
- RBAC / ServiceAccount changes: Role / ClusterRole / RoleBinding / ClusterRoleBinding edits, ServiceAccount or its token/permissions changed (these break Secret reads, API access, admission).
- Config / feature-flag changes: ConfigMap / Secret edits, CoreDNS Corefile ConfigMap edits, Ingress/NetworkPolicy changes, feature-flag flips.
- Admission / operator changes: Validating/MutatingWebhookConfiguration edits, CRD or controller upgrades.
- CronJobs / Jobs that ran in the window (batch, data migrations, cluster maintenance jobs).
If the window has zero change events, treat it as a strong signal in itself: the failure is likely external (upstream provider, certificate expiry, DNS, capacity drift from organic growth) rather than a self-inflicted regression.
3. Classify against the four reference paths
Match the failure shape to one of these four canonical paths first. They cover the majority of Kubernetes incidents; only branch out once they are ruled out.
If the failure does not match any of the four, classify as "outside reference paths" and document why. Outside-reference-paths means the methodology has no reference path for the failure shape, so its confidence in the root cause is low and escalation to a human is mandatory (step 6). It does not mean the agent does nothing: a pre-approved safe mitigation (traffic-shift to a healthy peer, feature-flag-off) is still recommended as the top action when available, with the root-cause investigation escalated in parallel. See step 6 for the exact ordering.
4. Confirm with three independent signals
Never declare a hypothesis on one signal. Require at least three of the following, drawn from independent sources:
- Pod / cluster events (kubectl get events, kubelet: OOMKilled, BackOff, Unhealthy, FailedScheduling).
- Logs (application container logs, system component logs).
- Metrics (request rate, error rate, latency, saturation, working-set, CoreDNS error rate — from Prometheus / kube-state-metrics).
- Traces (distributed traces showing the failing hop / Service).
- Change events (rollouts, ConfigMap/Secret, RBAC, HPA/cluster changes).
- External signals (customer reports, status pages of dependencies the workload calls).
Two signals from the same source (e.g. two log lines) count as one. The independence requirement is the guard against confirmation bias.
Split aggregate signals before trusting them. Error rate, latency, and saturation are usually reported as a single number across every region, cluster, AZ, shard, or canary/stable split. Before classifying, break each aggregate down along these dimensions. A per-dimension asymmetry — one region failing while its peer is healthy, one shard hot while the rest are flat — is a first-class diagnostic signal, and aggregate metrics actively hide it (a 25% failure in one of two equally-sized regions shows up as a moderate ~12% aggregate that matches no clean reference path).
When the failing and healthy slices run the same image tag / same code, a code-regression path (OOM, deploy-correlator) is ruled out by construction: identical code cannot fail in one slice and not the other. The cause is environmental — config/GitOps drift, a stale Service reference, per-region capacity, an external dependency reachable from only one slice. A confirmed asymmetry short-circuits the four-path search: stop trying to fit OOM/DNS/cascade/deploy-correlator on the aggregate, classify "outside reference paths" with a regional-asymmetry (or shard/AZ-asymmetry) reason, and move to step 5. Continuing to hunt for a reference-path match on aggregate signals after an asymmetry is detected wastes the investigation and is the single most common way this step runs long.
5. Quantify blast radius
Before recommending action, estimate:
- Users affected (count or percentage of traffic).
- Surfaces affected (which Services/endpoints, which namespaces, which clusters/regions, which customer segments).
- Business impact (revenue, SLO burn, contractual obligations if known).
A wrong mitigation that touches more surface than the incident itself is worse than the incident.
6. Propose mitigation before root cause
Mitigation comes first. Root cause comes after the bleeding stops.
Hard constraints, in order. Check these before ranking the standard actions below:
- If the classification from step 3 is "outside reference paths", escalation to a human is mandatory — but it is not automatically the top action. Two mitigations are pre-approved as safe because they are reversible and contained, and when one of them is available it becomes the top recommended action:
- Traffic-shift away from the failing slice to a healthy peer (other region/cluster/shard/replica). This is the canonical first move for a regional/shard asymmetry: it stops the bleeding immediately and is trivially reversible. When the asymmetry detector in step 4 has identified a healthy peer, recommend the traffic-shift as action #1, then escalate the root-cause investigation (config/GitOps drift, the failing dependency) to a human as the parallel follow-up.
- Feature-flag off the failing code path, if a flag exists.
Every other option — contacting an external provider, irreversible config/RBAC/state changes, anything touching the failing slice directly — is surfaced as an alternative for the human to approve, not executed by the agent. The principle (from FAILUREMODES M1): outside-reference-paths means low confidence in root cause*, so escalate the root cause; it does not forbid the safe, reversible mitigation that an on-call would reach for first.
- If T0 was flagged as ambiguous in step 1, the top recommended action is "re-run the investigation with a widened window". Only after that re-run identifies a fuller change surface should any irreversible mitigation (rollout undo, RBAC change, cluster/infra rollback) be recommended.
- If the implicated change has bundlesize > 1 (multiple changes shipped in one rollout), rollout undo remains the top recommendation but requires explicit human approval before execution.** Surface the asymmetry explicitly: "the rollback reverts N changes when the incident affects only K of them".
- If the classification is "cascading-failure", the top action is to break the amplification loop at its source, not to undo a rollout. A pure cascade typically has no rollout in the window (the trigger is a degraded dependency, not a deploy), so there is nothing to revert. Recommend, in order: open the circuit breaker on / shed load from the degraded dependency itself (the root of the cascade), then cap or disable the retry budget at the callers driving the retry storm. Shedding the callers' retries alone treats the symptom (the amplification) while leaving the degraded dependency saturated; opening the circuit at the dependency stops the loop at its origin and lets the dependency recover.
Standard mitigation order (applies when the constraints above do not fire):
- kubectl rollout undo the workload identified in step 2, if one rollout is clearly implicated and reversible (or revert the implicated ConfigMap / RBAC change).
- Feature-flag off the failing code path, if a flag exists.
- Scale the saturated resource (kubectl scale / raise the HPA ceiling / raise resources.limits), if the path is capacity-bound and not regression-bound.
- Traffic-shift away from the failing region / cluster / Service version / shard.
- Manual intervention (kubectl delete pod to force a fresh restart, kill a stuck Job) as a last resort, with explicit acknowledgement that it does not address cause — pods will recreate from the same broken spec.
If no safe mitigation exists even after applying the above, surface that explicitly and escalate.
7. Hand off
Produce a structured handoff for postmortem-author. All four elements below are mandatory and must appear as labelled sections, even when an element is empty (write "Open questions: none identified", not nothing — a silently missing section reads as "investigation incomplete" to the next responder):
- Timeline (T0, key events, mitigation timestamp, Tresolved).
- Ranked hypotheses with the evidence supporting each.
- Mitigation taken / recommended and observed effect.
- Open questions (gaps in signals, unverified assumptions, root-cause threads the mitigation did not close). This section is the most-often dropped and the most valuable to the postmortem: list every unresolved thread explicitly. If the investigation truly left no gaps, say so explicitly rather than omitting the heading.
Output format
The agent's final message in any invocation must include:
- Anchored window: T0 = ..., Tnow = ....
- Change surface: bulleted list of overlapping changes (rollouts, ConfigMap/Secret, RBAC, HPA/cluster, CronJobs), or "no changes in window".
- Classified path: one of the four, or "outside reference paths" with justification.
- Confirming signals: three or more, each cited with source.
- Blast radius: users + surfaces + business impact.
- Recommended mitigation: ordered, with explicit "do not address cause" notes where applicable.
- Handoff payload: structured for postmortem-author, containing all four labelled sections from step 7 — timeline, ranked hypotheses, mitigation, and open questions. Do not collapse or omit any of them; an absent "open questions" section is treated as an incomplete handoff.
Worked examples
Eleven end-to-end examples are committed under examples/, each with fixtures and a runnable replay test.
Reference paths (one canonical example per path):
- examples/01-oom-cascade.md: OOM in a payments Deployment triggering a retry storm from the API gateway.
- examples/02-dns-resolution-failure.md: CoreDNS Corefile ConfigMap misconfiguration causing intermittent SERVFAIL for an internal Service.
- examples/03-cascading-failure-retry-storm.md: pure cascade from an upstream DB query-plan slowdown; no rollout in window.
- examples/04-deploy-correlator-serialization.md: pure deploy-correlator regression (serialization change breaks downstream parsers).
Escalation cases (exercise the FAILURE_MODES.md rules):
- examples/05-outside-reference-paths-third-party-rate-limit.md: a third-party API rate-limits a workload; the methodology escalates rather than force-fitting one of the four paths (M1).
- examples/06-ambiguous-t0-slow-burn.md: slow-burn memory leak where T0 is genuinely ambiguous; escalation (M2) recommends re-running with a widened window.
- examples/07-blast-radius-asymmetric-revert.md: a rollout bundling six unrelated changes; rollout undo is the top mitigation but escalates (M3) because the rollback blast radius exceeds the incident.
- examples/08-deploy-correlator-confirmation-bias.md: a rollout and an RBAC change collide in time; the methodology rejects the deploy-correlator classification (M4 guard) because the rollout diff does not touch the failing surface.
Edge / boundary cases:
- examples/09-zero-changes-external-cert-expiry.md: zero changes in window, failure is an external partner TLS certificate expiry seen from a cluster workload.
- examples/10-multi-region-asymmetry.md: same image deployed to two clusters/regions, one fails; the methodology surfaces the per-region asymmetry as a first-class signal.
- examples/11-capacity-bound-organic-growth.md: organic traffic growth saturates capacity; the methodology recommends scaling (HPA) instead of rollout undo.
The examples mirror the seven methodology steps so contributors can see the methodology in motion, not just described.
Replay tests
Every example has a replay test in tests/ that runs the methodology against committed fixtures, with no external credentials (no live cluster needed). Run from the skill directory:
for t in tests/replay_*.py; do python "$t" || exit 1; doneThe 11 tests cover the four reference paths, the FAILURE_MODES.md escalation rules (M1, M2, M3, M4), and the edge cases (zero changes, multi-region asymmetry, capacity saturation). Tests exit non-zero if the methodology produces the wrong classification, mitigation, or escalation against known-good fixtures. See tests/README.md for the fixture schema and how to add a new replay test.
Failure modes
More skills from anyshift-io/sre-skills
- Aiam-deceptive-escalation-auditorAudit the union of every IAM policy attached to one principal for privilege-escalation paths that no single statement reveals, and for apparent escalations that are already neutralised. Resolves the effective permission set across all attached policies (Allow minus blanket Deny), then checks the cross-statement escalation combos (iam:PassRole + a compute-launch action, policy-rewrite-in-place, function-code hijack, self-attach admin, trust-policy rewrite + assume, credential minting for another identity), the wildcard grants (Action '*' on Resource '*', service-level wildcards, Allow+NotAction), and the trust-policy exposure. Its discipline is symmetric: it does NOT flag a PassRole combo killed by an explicit Deny, an Action '*' pinned to one bucket, an sts:AssumeRole whose target does not trust back, a mutation kit capped by a permissions-boundary Deny, or a cross-account assume sealed by an unsatisfiable Condition. Reports findings with severity and a fix, then names what a single principal's policies cannot answer (the privileges of a passed/assumed role, the permissions boundary, the org SCPs). Use when asked to audit an IAM policy, role, or user for escalation, over-broad grants, or "can this principal become admin." Vendor-neutral; runs offline against the policy JSON with no Anyshift account.
- As3-estate-calibration-auditorAudit an estate of AWS S3 buckets for the one bucket that is genuinely publicly or cross-account exposed, without over-flagging the many buckets that READ as exposed but are neutralised. Resolves each bucket's EFFECTIVE verdict by composing four layers (Block Public Access x bucket policy x bucket ACL x access points), never one layer alone, then rolls the per-bucket verdicts up into an estate verdict. Its discipline is symmetric: BPA (RestrictPublicBuckets / BlockPublicPolicy) neutralises a Principal '*' policy but NOT a cross-account grant; IgnorePublicAcls kills a public-group ACL grant but NOT a cross-account canonical-user grant; a narrowing Condition (org id, ExternalId, SourceIp, access-point delegation) scopes a Principal '*' so it is not public. On a needle estate it names the ONE live bucket as the primary finding; on a clean estate it reports NO live exposure and does not manufacture findings. Then it states what the bucket configs alone cannot answer (per-object ACLs, CloudFront/CDN fronting, the trusted principals' identity policies, account-level BPA dependency, data sensitivity). Use when asked to review an S3 bucket fleet for public exposure, cross-account access, or whether the estate is clean. Vendor-neutral; runs offline against describe-bucket / get-bucket-policy / get-bucket-acl / list-access-points JSON with no Anyshift account.
- Asg-deceptive-reachability-auditorAudit a fleet of AWS security groups for the multi-hop lateral-movement path that no single ingress rule reveals. Builds a directed reachability graph from the SG-to-SG references (an ingress rule on SG B naming SG A means a host in A can reach B), adds an internet edge for every 0.0.0.0/0 rule, then composes those edges into the transitive closure from a named entry point (the internet, or a compromised host). Reports the shortest reachable path to the crown-jewel tier, the blast radius, and any pivot/hub SG that bridges otherwise-isolated regions, each ranked by severity with a fix. Its discipline is symmetric: on a segmented or orphaned fleet where the chain does NOT reach the crown jewel, it reports clean and names the boundary instead of fabricating a path. Then it states what the SG graph alone cannot answer (live host membership, route tables, NACLs, app-layer auth). Use when asked to review a security-group fleet for lateral movement, blast radius, or whether the internet can reach a sensitive tier. Vendor-neutral; runs offline against describe-security-groups + describe-instances JSON with no Anyshift account.
- Asqs-queue-auditorAudit a single AWS SQS queue's configuration for the misconfigurations that silently drop or re-deliver messages while every attribute reads as fine. Parses the GetQueueAttributes output (and the referenced dead-letter queue), checks the redrive path (DLQ present, maxReceiveCount band, DLQ-vs-source retention ordering), the message lifecycle (poison messages aging out before they reach the DLQ, default visibility timeout, short retention), and exposure (open resource policy, encryption at rest, FIFO dedup contract). Reports findings with severity and a recommendation, then names the boundary: the questions a single queue's config cannot answer (consumer processing time, live behaviour, the IAM union, the producers and consumers on either side). Use when asked to review, harden, or sanity-check an SQS queue, or to explain why messages are going missing. Vendor-neutral; runs offline against the queue attributes with no Anyshift account.