Observability and SRE
Debugging and operations with AI
The system reports a fault, but reconstructing what happened takes time. Logs, measurements and previous experience are scattered.
Discuss this workshopAn example exercise
Illustrative task. We choose the actual exercise around your team’s work and tools.
Start with sanitised data from a past incident or a prepared failure. Check the AI’s hypotheses against logs and measurements.
How we work
-
Reconstruct the timeline from the available data.
-
Use AI to propose causes, then verify them.
-
Define the next action, its checks and rollback conditions.
What stays with your team.
- Reusable investigation steps and queries.
- A documented investigation of the selected case.
- An improved alert or operational runbook.
Making time for the work
Case review and practical sessions using prepared or sanitised data. Live production intervention is outside the exercise.
We agree session count and length, group size, preparation, follow-up and fees before starting. Environment preparation and practice get their own time allocation.
What do we need to start?
Usable logs, metrics or traces, a basic system description and agreed data handling boundaries.