30+ agent skills the platform team runs
Platform work is repeated judgment — provisioning new infrastructure, changing what's already running, and troubleshooting it when it breaks. I built 30+ Claude Code and Codex agent skills covering that whole lifecycle, and they've been adopted by the platform engineers as part of the daily workflow: scaffolding new infra to our conventions, reviewing changes before they merge, and walking incidents change-first.
The three I go deepest on: a Terraform plan reviewer that risk-tiers every change and leads with an apply/don't-apply verdict, a Kubernetes manifest reviewer that checks limits, probes, image pinning and RBAC, and an incident-triage assistant that walks a firing alert change-first. Each pairs the model's judgment with a deterministic Python script that parses the actual plan JSON or YAML — every destroy, every IAM widening, every missing limit — and exits non-zero on critical findings so CI can block a bad apply. The model explains and prioritizes; it doesn't invent findings.
I validated them with an eval loop — every test case run with and without the skill, graded against fixed checks. The honest result: the base model is already strong, and the repeatable win is determinism plus safety discipline. The skills recommend rollbacks; they never run them.
- Claude Code
- Codex
- Python
- GitHub Actions
- Terraform
- Kubernetes