← Back to the blog

Interviewing

An AI evaluation exercise for engineering candidates

Assess how candidates define acceptable AI behavior, choose test cases, inspect failures, and evaluate changes to an AI product.

AI hiringInterviewing

An AI evaluation hiring exercise should test how a candidate decides whether a system is useful and dependable. Ask them to define acceptable behavior, choose representative cases, inspect failures, and explain how they would compare a change.

A compelling demo is the starting point for the exercise, not the passing condition.

Use a small fictional workflow

Illustrative exercise: A fictional support assistant reads an approved knowledge base and drafts replies. It should acknowledge missing information and route requests it cannot handle. It must not invent account-specific facts.

Give candidates a short product description, a sample knowledge base, and several example conversations. Include an ordinary request, an ambiguous request, an unsupported request, and a case where the available information conflicts.

Use invented data throughout.

Ask for an evaluation plan

Have the candidate describe:

  1. What counts as a useful response for this workflow.
  2. Which failures matter most and why.
  3. Which checks can be mechanical and which require judgment.
  4. How they would inspect a failure before changing the system.
  5. What evidence would justify releasing a revised version.

Anthropic's guide to agent evaluations distinguishes tasks, repeated trials, grading, and the resulting outcome. That distinction is useful here: a plausible conversation and a successful task are not necessarily the same thing.

This interview exercise is an editorial adaptation, not a validated hiring test from Anthropic.

Score the reasoning

AreaEvidence to look for
Success definitionConnects evaluation to the user's task
Case selectionCovers meaningful situations and failures
GradingExplains what a check can and cannot establish
DiagnosisSeparates different failure causes
Release judgmentRecognizes tradeoffs and remaining uncertainty

Ask candidates to explain a disagreement between two possible grading methods. The goal is to see whether they investigate the disagreement, not whether they insist their preferred method is always correct.

Change one constraint

Tell the candidate that a proposed improvement produces better answers but takes longer. Ask what they would measure and who should help decide whether the tradeoff is acceptable.

Alternatively, introduce a new customer workflow and ask which parts of the evaluation set need revision. This exposes whether the plan was built around the product's purpose or around convenient examples.

Keep the scope honest

The candidate does not need to build an entire evaluation platform during the interview. A concise plan and discussion can reveal the relevant decisions.

Use a separate engineering assessment for implementation depth. Before opening the search, clarify whether you need a research scientist or an AI engineer.

Put this guide to work

Hire people who build like founders.

Share the role, the outcomes this person should own, and your hiring constraints. Refery brings specialist recruiters and trusted referrals behind one brief.