Where AI helps, and where it must not lead
A clear split between the incident work AI should accelerate and the decisions that must stay human.
Not every part of an incident is a good fit for AI. Sorting the helpful uses from the dangerous ones is the first real skill.
Where AI helps, because the work is reading and summarizing at speed:
- Compressing alert storms into a short, ranked summary.
- Pulling recent deploys, related incidents, and relevant dashboards into one view.
- Generating several root cause hypotheses to investigate, not one answer to trust.
- Drafting a status update or a timeline for a human to review.
Where AI must not lead, because the work carries irreversible consequences or final accountability:
- Declaring root cause. A confident hypothesis is a starting point, never a conclusion.
- Running production commands unattended. Restarts, rollbacks, scaling, and config changes need a human who understands the blast radius.
- Changing severity or paging decisions on its own.
- Sending customer-facing communication without review.
The pattern is simple. AI is trusted with reading, correlating, and drafting. Humans stay in charge of deciding, acting on production, and communicating externally. When a tool offers to act automatically, ask what happens if it is wrong at three in the morning with no one watching. If the answer is bad, keep that step human and let AI prepare it, not perform it.
List the automated actions your current tooling can take during an incident. For each, write the worst case if it fires incorrectly, and decide whether it should require human approval.
Enjoying the free lessons? Get an email when we publish new courses and updates — no spam, unsubscribe anytime.
Discussion (0)
Ask a question or share what worked for you. Comments are reviewed before they appear.
No comments yet. Be the first to start the discussion!