Investigate a production incident
Give Claude Code the symptom and when it started; it checks logs, recent deploys and config changes, then names the most likely cause with evidence.
Task: Investigate a production incident · Other tasks
Fill in the details
Optional. Deploys, config edits, migrations or traffic changes around that time.
Optional. Paste what you already have. Remove secrets, tokens and customer data first.
Copy your prompt
This site doesn't run Claude or show model output. Results depend on your input and the model you use.
You've edited the prompt, so changes to the fields aren't applied.
Replace what you've entered for this task with the example? This can't be undone.
What you type is kept in this tab's session storage so a reload doesn't lose it. Use "Clear this task" to remove it. Browsers can restore session data when they reopen tabs, so closing a tab isn't a guaranteed way to erase it.
When to use this template
- Something is broken in production, you have a symptom and a rough start time, and you want the investigation organized while you handle communication.
- Claude Code can reach the evidence: log files, the deploy history in git, configuration in the repository, or command line tools for your platform.
- You want an answer that shows its evidence, so you can judge the cause quickly instead of trusting a confident guess.
When not to use it
- The incident needs an immediate action you already know, such as rolling back a deploy that clearly caused it. Do that first and investigate afterwards.
- Claude Code has no access to logs, deploys or configuration. Paste the evidence into the logs field, or use the error debugging template with the stack trace.
- The logs contain customer data or secrets you are not allowed to share with an AI tool. Redact them, or follow your incident policy instead.
Why this structure
- Pasted logs go first inside
<logs>tags, and the rules say to treat them as data. Log lines can contain user input, so the tags keep them apart from your instructions. - The symptom, start time and known changes are separate lines, because an incident is usually explained by what changed shortly before it started.
- Naming logs, recent deploys and config changes points Claude at three common places to look for a cause, and keeps the investigation from wandering through the whole codebase.
- The read-only rule keeps an investigation from becoming an unreviewed change to a live system. Claude proposes mitigations; you decide and run them.
- The report asks for evidence with timestamps and for alternatives with a way to rule each in or out, so you can check the reasoning rather than accept a single story.
- Point 4 asks what could not be checked, which is often the most useful part: it tells you exactly which access or data would settle the question.
Example input (fictional)
- What is going wrong
The checkout endpoint returns 500 errors for about 20% of requests.
- When it started
around 14:10 UTC today
- Recent changes you know about
Deployed v4.12 at 13:55. A database migration ran at 14:00.
- Logs or error output
14:11:02 ERROR checkout: payment client timeout after 3000ms 14:11:02 ERROR checkout: payment client timeout after 3000ms 14:11:05 WARN pool: connection pool exhausted (size=10)
Follow-ups to send Claude
- Assume your most likely cause is right. Write the exact steps to mitigate it, in order, with a check after each step that shows whether it worked.
- Build a timeline of the incident from the logs and the deploy history, one line per event, in UTC.
- Draft a short incident summary for the team: what happened, impact, cause, what we did, and follow-up actions.
Common mistakes
- Pasting only the last error line. The lines just before the first error, and their timestamps, usually say more about the cause.
- Leaving out the start time and recent changes. Without them, Claude has to search everything instead of the window where the cause most likely sits.
- Letting Claude apply a mitigation during the investigation. Keep the read-only rule, and apply changes yourself once the cause is clear.
Related templates
- Debug an error: Turn an error message and the code around it into a ranked list of likely causes, a check for each, and a fix you can try.
- Optimize code to a measurable target: Give Claude Code a metric and a target value, so it measures first, changes one thing at a time and stops when the number is reached.
- Turn meeting notes into action items: Turn rough meeting notes into a clear list of action items with owners and due dates, plus decisions made and questions left open.
Sources
- Claude Code docs: Prompt library (Anthropic documentation)
- Claude Code best practices: Avoid common failure patterns (Anthropic documentation)
- Prompting best practices: Structure prompts with XML tags (Anthropic documentation)