name: kubernetes-skill description: "Prevent Kubernetes hallucinations by diagnosing and fixing failure modes: insecure workload defaults, resource starvation, network exposure, privilege sprawl, fragile rollouts, and API drift. Use when generating, reviewing, refactoring, or migrating manifests, Helm charts, Kustomize overlays, cluster policies, and platform-specific Kubernetes work for EKS, GKE, AKS, OpenShift, GitOps controllers, or observability stacks."

KubeShark: Failure-Mode Workflow for Kubernetes

Run this workflow top to bottom.

Record before writing manifests:

cluster version (e.g. 1.30, 1.31) and distribution (EKS, GKE, AKS, k3s, vanilla)
target namespace and environment criticality (dev/staging/prod)
workload type (Deployment, StatefulSet, Job, CronJob, DaemonSet)
deployment method (raw YAML, Helm, Kustomize, operator-managed)
policy enforcement (Pod Security Admission level, Kyverno, OPA/Gatekeeper)
cloud provider and CNI (affects networking, storage classes, load balancers)
platform controllers/add-ons (GitOps, observability, ingress, service mesh, autoscaling)

If unknown, state assumptions explicitly.

Select one or more based on user intent and risk:

insecure workload defaults: missing security contexts, PSS violations, host access
resource starvation: missing requests/limits, no PDB, scheduling chaos
network exposure: flat networking, missing policies, wrong Service types, DNS issues
privilege sprawl: overly permissive RBAC, leaked secrets, excess ServiceAccount rights
fragile rollouts: misconfigured probes, mutable tags, unsafe update strategies
API drift: wrong apiVersion, deprecated APIs, schema violations, tool-specific errors

Primary failure-mode references:

Supplemental references (only when needed):

Conditional Reference Retrieval (CRR) references (load only when the signal is detected):

references/conditional/eks-patterns.md for EKS, AWS, IRSA, EKS Pod Identity, AWS Load Balancer Controller, EBS/EFS CSI, Karpenter
references/conditional/gke-patterns.md for GKE, Autopilot, Workload Identity Federation for GKE, Dataplane V2, GCE Ingress, Config Sync
references/conditional/aks-patterns.md for AKS, Microsoft Entra Workload ID, Azure CNI, AGIC, Azure Disk/File/Blob CSI
references/conditional/openshift-patterns.md for OpenShift, OKD, ROSA, ARO, Routes, SCCs, OLM, oc
references/conditional/gitops-controllers.md for Argo CD, ApplicationSet, Flux, GitOps reconciliation, sync waves
references/conditional/observability-stacks.md for Prometheus Operator, ServiceMonitor, PodMonitor, OpenTelemetry, Loki, Grafana

Do not load multiple CRR files unless the task spans multiple detected platforms/tools.

For each fix, include:

When applicable, output:

Always provide validation steps tailored to deployment method and risk tier:

kubectl apply --dry-run=server or kubectl diff
kubeconform for schema validation against target cluster version
cross-resource consistency check (label/selector/port alignment)
policy scan (PSS profile check, Kyverno/OPA audit) Never recommend direct production apply without reviewed diff and approval.

Return: