An AI evaluation hiring exercise should test how a candidate decides whether a system is useful and dependable. Ask them to define acceptable behavior, choose representative cases, inspect failures, and explain how they would compare a change.
A compelling demo is the starting point for the exercise, not the passing condition.
Use a small fictional workflow
Illustrative exercise: A fictional support assistant reads an approved knowledge base and drafts replies. It should acknowledge missing information and route requests it cannot handle. It must not invent account-specific facts.
Give candidates a short product description, a sample knowledge base, and several example conversations. Include an ordinary request, an ambiguous request, an unsupported request, and a case where the available information conflicts.
Use invented data throughout.
Ask for an evaluation plan
Have the candidate describe:
- What counts as a useful response for this workflow.
- Which failures matter most and why.
- Which checks can be mechanical and which require judgment.
- How they would inspect a failure before changing the system.
- What evidence would justify releasing a revised version.
Anthropic's guide to agent evaluations distinguishes tasks, repeated trials, grading, and the resulting outcome. That distinction is useful here: a plausible conversation and a successful task are not necessarily the same thing.
This interview exercise is an editorial adaptation, not a validated hiring test from Anthropic.
Score the reasoning
| Area | Evidence to look for |
|---|---|
| Success definition | Connects evaluation to the user's task |
| Case selection | Covers meaningful situations and failures |
| Grading | Explains what a check can and cannot establish |
| Diagnosis | Separates different failure causes |
| Release judgment | Recognizes tradeoffs and remaining uncertainty |
Ask candidates to explain a disagreement between two possible grading methods. The goal is to see whether they investigate the disagreement, not whether they insist their preferred method is always correct.
Change one constraint
Tell the candidate that a proposed improvement produces better answers but takes longer. Ask what they would measure and who should help decide whether the tradeoff is acceptable.
Alternatively, introduce a new customer workflow and ask which parts of the evaluation set need revision. This exposes whether the plan was built around the product's purpose or around convenient examples.
Keep the scope honest
The candidate does not need to build an entire evaluation platform during the interview. A concise plan and discussion can reveal the relevant decisions.
Use a separate engineering assessment for implementation depth. Before opening the search, clarify whether you need a research scientist or an AI engineer.