1. Incident Post-Mortem Report Writer
Write comprehensive post-incident reviews with root cause analysis, timeline, impact, and lessons learned.
Copy-paste prompt
You are a senior site reliability engineer writing a blameless post-incident review for [service/system]. Use the following incident details: [paste timeline, alerts, logs, customer impact, mitigation steps, and follow-up notes]. Create a comprehensive report with: (1) incident summary, (2) severity and duration, (3) customer and internal impact, (4) detailed timeline from detection to resolution, (5) root cause analysis, (6) contributing factors, (7) what went well, (8) what slowed response, (9) lessons learned, and (10) action items with owner, priority, and due date. Keep the tone factual, blameless, and suitable for engineering leadership.
Specific examples:
- Turn a Slack incident timeline into a blameless post-incident review for Confluence.
- Summarize a P1 database failover with root cause, user impact, and corrective actions.
- Convert rough incident notes into an executive-ready review with owners and due dates.
2. On-Call Handoff Document Generator
Create handoff documents covering current incidents, escalation paths, and critical service status.
Copy-paste prompt
You are an SRE preparing an on-call handoff for the incoming engineer. Convert these notes into a clear handoff document: [paste active incidents, recent alerts, open tickets, risky deploys, service health notes, and people to contact]. Include: (1) current overall status, (2) active incidents with severity, owner, next action, and last update time, (3) critical services and their health, (4) alerts to watch and why, (5) noisy or known-flaky alerts, (6) pending deploys or maintenance windows, (7) escalation paths by service, and (8) unresolved risks for the next shift. Write it for an engineer reading quickly at the start of an on-call rotation.
Specific examples:
- Create a Friday-to-Monday handoff covering active incidents, risky deploys, and noisy alerts.
- Summarize current service status for the next primary and secondary on-call engineers.
- Document escalation paths for critical services before a holiday coverage rotation.
3. Runbook Documentation Updater
Generate or update runbooks with new troubleshooting procedures and automation steps.
Copy-paste prompt
You are a site reliability engineer updating an operational runbook for [service/procedure]. Existing runbook content: [paste current runbook or notes]. New troubleshooting steps, commands, automation, or lessons learned: [paste updates]. Rewrite the runbook so it includes: purpose, when to use it, prerequisites and access required, safety checks, step-by-step troubleshooting flow, exact commands or dashboard links where available, expected outputs, decision points, automation steps, rollback guidance, validation checks, and escalation criteria. Optimize for clarity during a live incident.
Specific examples:
- Update a runbook after a new cache-recovery command is added to your tooling.
- Generate a first-draft troubleshooting guide for elevated 5xx errors on an API service.
- Add verification and rollback steps to an existing Kubernetes restart procedure.
Nexus Vault
Need more than the free prompts?
Get 200 business prompts, instant download, and no subscription.
4. Alert Configuration Advisor
Design or refine alert thresholds and notification rules based on service criticality.
Copy-paste prompt
You are an SRE reviewing alert configuration for [service]. Service criticality: [tier / user-facing impact / SLOs]. Current alerts and thresholds: [paste alert rules, metrics, paging policy, and recent alert history]. Recommend improved alert thresholds and notification rules. For each recommendation include: metric, proposed threshold, window, severity, paging versus ticket/channel notification, rationale tied to user impact or SLO burn rate, expected false-positive risk, and any alerts that should be removed, grouped, or converted to dashboards. Prioritize actionable alerts over symptom noise.
Specific examples:
- Review latency alerts for a tier-0 checkout service and suggest cleaner thresholds.
- Design paging versus ticket-only rules for batch jobs, APIs, and background workers.
- Reduce alert fatigue by grouping noisy symptoms under one actionable service alert.
5. Incident Commander Brief Writer
Compose status briefs for incident commanders during active outages.
Copy-paste prompt
You are assisting the incident commander during an active outage. Use these current incident notes: [paste timeline, affected services, customer impact, mitigation attempts, owners, risks, and next checkpoints]. Write a concise incident commander brief with: (1) current status in one sentence, (2) severity and business impact, (3) affected services and user symptoms, (4) confirmed facts versus unknowns, (5) actions completed, (6) actions in progress with owners, (7) next decision point, (8) escalation needs, and (9) suggested internal update for the next 15- or 30-minute checkpoint. Keep it calm, precise, and easy to read aloud on an incident bridge.
Specific examples:
- Draft a 15-minute incident commander update for leadership during a regional outage.
- Summarize current impact, mitigation status, and next decision point for a P1 bridge.
- Create a customer-support-ready status brief without exposing internal implementation detail.