Assess reliability judgment by asking how an engineer would understand and contain a failure before redesigning the system. A useful interview makes the sequence of decisions visible: what to check, what to protect, who to inform, and what to improve afterward.
Match the scenario to the ownership expected in the role. Do not require specialist infrastructure knowledge for a job that will not use it.
Give candidates a concrete starting point
Illustrative scenario: A fictional application accepts user requests, stores them, and processes them in a background worker. Users report missing results after a release. Some requests may have completed even though the interface still shows them as pending.
Provide a simple diagram, a short release summary, and a few sample observations. Make clear that the candidate may ask for additional information.
Do not use a current customer incident or expose operational details from your own systems.
Explore decisions in order
| Step | Question |
|---|---|
| Understand impact | What would you establish about affected users and requests? |
| Gather evidence | Which observations would help distinguish possible causes? |
| Contain harm | What could you pause, reverse, or limit while investigating? |
| Recover | How would you avoid duplicating work during recovery? |
| Communicate | What should teammates and users know while facts are incomplete? |
| Prevent recurrence | What would you change after the immediate issue is resolved? |
There is no universal correct sequence of commands. Look for a reasoned approach that updates as evidence arrives.
Introduce one new fact
Tell the candidate that retrying a request can trigger an external action twice. Ask how this changes the recovery plan.
Useful evidence includes recognizing the new consequence, checking which requests already completed, and explaining uncertainty before taking an irreversible action. Do not reward confident guesses that ignore missing information.
If the candidate proposes a technical mechanism, ask how it would be verified. Keep the depth appropriate to the role.
Distinguish diagnosis from prevention
A candidate may be good at identifying a likely cause but weak at organizing safe recovery. Another may propose extensive monitoring without first addressing users who are affected.
Record each dimension separately. Ask which improvements are urgent and which can wait. A smaller team needs someone who can explain the tradeoff between immediate risk and maintenance effort.
Use the result with past evidence
This is an interview exercise, not proof that someone will handle every production incident well. Pair it with a project deep dive that explores actual operational responsibility.
Write the conclusion in terms of observed behavior: “Checked for duplicate effects before suggesting retries” is more useful than “strong systems thinker.”