Choose a research scientist when the central work is investigating an unresolved research question. Choose an AI engineering mandate when the central work is making AI capabilities function reliably inside a product. Titles vary, so define the work and expected evidence before selecting a label.
Some roles combine both. If yours does, explain which responsibility takes priority when experiments and product delivery compete.
Identify the uncertainty you are hiring to resolve
A useful starting question is: “If this hire succeeds, what will we know or be able to do that we cannot do today?”
| Dominant problem | Work to describe in the brief |
|---|---|
| An approach has not demonstrated the required capability | Hypothesis formation, experiments, comparison, and research interpretation |
| A capability exists but does not work reliably in your workflow | Integration, evaluation, failure analysis, and iteration |
| A prototype works but is difficult to operate | Deployment, observability, cost management, and system ownership |
| The team lacks a reliable way to judge quality | Evaluation design, representative cases, and review procedures |
These are mandate categories, not universal definitions of job titles.
Specify the output and its user
For research work, describe the decision an experiment should support. For product work, describe the user workflow and the operational constraints.
Illustrative example: A fictional document assistant already produces plausible answers in a demo. The next task is to handle missing information, measure errors, and connect to a permissions-aware product workflow. That brief should make evaluation and engineering ownership explicit. “Push the frontier of AI” would obscure the immediate need.
Conversely, an engineering delivery interview alone would not establish someone's ability to design and interpret a new research program.
Match the interview to the mandate
For a research-heavy role, explore how the candidate forms hypotheses, chooses baselines, interprets inconclusive results, and decides what to investigate next.
For an engineering-heavy role, explore how the candidate defines acceptable behavior, diagnoses failures, integrates dependencies, and handles operation after launch.
For a hybrid role, assess both areas separately. Do not let depth in one silently compensate for a missing requirement in the other.
Resolve ownership before opening the search
Name who decides research priorities, who owns the product milestone, and who maintains the resulting system. Explain what infrastructure and collaborators are available.
Avoid combining research, full-stack product development, customer deployments, and infrastructure into one undifferentiated list. If all are required, identify the first outcome and what support will arrive.
Use the AI evaluation hiring exercise for a product-oriented assessment. For a broader search, see Refery's AI engineer hiring page.