Use this guide to prepare for AI Engineer interviews, with a focus on openai apis, llm fundamentals, gpt. Explain your reasoning and connect it to experience you can substantiate.
These preparation themes come from the questions in this role’s bank. They help you organise your examples; individual employers may assess different things.
OpenAI APIs
LLM fundamentals
GPT
LangGraph
Communicating to executives
Claude
A useful preparation sequence
Choose your experience level and the round you expect.
Answer one question in your own words before opening its guide.
Compare your reasoning, evidence and trade-offs; adapt the answer to your experience.
Practise the follow-up, then revisit one answer you want to improve.
Representative questions and answer guidance
Open any question to read its answer. The complete guidance is included on this page.
Technical · Mid-level
1. How would you build a reliable application around an OpenAI API call?
Answer guide
Treat the model call as an unreliable external dependency. Ask for a structured output where possible and validate it in code, since output can be wrong or malformed. Record request IDs, latency, token usage and cost, and set budgets and alerts. Build a fallback for when the service is unavailable, and evaluate quality on a fixed set of realistic cases before and after any prompt or model change. API details and model names change over time, so check the current documentation for your account rather than hard-coding assumptions. Do not send confidential data unless policy allows it.
What this question explores
Production thinking around an LLM API: security, failure handling, cost control and evaluation.
Common mistakes
Calling the API directly from the browser with the key exposed.
Trusting model output as correct without validation, timeouts or a fallback.
Practise a follow-up
How would you handle a rate-limit response without creating duplicate actions?
How would you evaluate a prompt change before releasing it to all users?
2. Why can an LLM produce a confident but unsupported answer?
Answer guide
A language model generates text by predicting a likely continuation of its context, which is why it sounds fluent even when a fact is missing or wrong. It has no built-in step that verifies each claim against a trusted source, and a confident tone is just style. To reduce unsupported answers, supply retrieved, authoritative material and ask it to answer only from that, allow it to say it does not know, and require citations to the source text. For high-stakes uses, add human review, then measure how often answers are supported on a test set that includes questions with no answer.
What this question explores
Whether you can explain hallucination in simple terms and name practical ways to ground and measure the assistant's replies.
Common mistakes
They say the model is lying or deliberately deceiving, instead of explaining that it predicts plausible text.
They claim that adding a line such as be accurate in the prompt removes the problem completely.
Practise a follow-up
How does retrieval help reduce unsupported answers?
How would you build a test set to measure how often the assistant makes things up?
3. How would you compare two GPT models for a document classification task?
Answer guide
Build a fixed, labelled evaluation set that represents real documents, including hard and ambiguous cases, and run both models on it with the same prompt, instructions and output schema so the comparison is fair. Look at the errors themselves, not just the score, because two models can differ in the kinds of mistake. Run repeat trials if outputs vary, and re-check on new data periodically as inputs drift. Model names and behaviour change, so record the exact version used. Choose the cheaper model if quality is equal within the noise, and document the decision for the team.
What this question explores
Ability to run a fair, repeatable model comparison that includes cost, latency and error types.
Common mistakes
Comparing on a few hand-picked examples and choosing by impression.
Using different prompts or schemas for each model, then attributing differences to the model.
Practise a follow-up
What sample size or variation checks would make you trust the difference between two models?
How would you detect drift after the chosen model is in production?
4. How would you prevent a tool-using LangGraph workflow from repeating a failed action indefinitely?
Answer guide
Model the workflow as an explicit graph with a state object that records attempts, errors and the last action, and add conditional edges that decide whether to retry, take another path, or stop. Make side-effecting tools idempotent so a repeat does not cause duplicates, and persist checkpoints so you can resume or inspect a run. Track loops through metrics such as steps per run and retry counts, and alert on outliers. I would explain the design to the team as a state machine with a clear exit for every path.
What this question explores
Reliable control flow for agents: state, limits, idempotent actions and escalation rather than open-ended looping.
Common mistakes
Letting the model decide when to stop, with no step cap or retry limit.
Retrying tools that change data without idempotency, creating duplicate actions.
Practise a follow-up
How would checkpoints help you resume or debug a run that failed halfway?
How do you decide when a failure should go to a human instead of another retry?
5. How would you explain an inconclusive A/B result to an executive who wanted a clear yes or no?
Answer guide
I lead with the decision, not the statistics. I say what we learned: the change did not show a measurable improvement, and the data rule out a lift larger than a certain amount, using the confidence interval in business terms. I explain that no evidence of an effect is different from evidence of no effect, and what it would cost to resolve it with more traffic or time. Then I give options: ship if the change is cheap and harmless, drop it, or run a better powered follow-up, with a recommendation and its risks. I avoid jargon and keep the p-value to an appendix.
What this question explores
Whether you can turn uncertainty into a decision-focused message that a non-technical leader can act on.
Common mistakes
Opening with p-values and test names instead of what the business should do.
Telling the executive the change has no effect, when the data merely lack precision.
Practise a follow-up
How do you handle an executive who insists on a yes or no?
6. How would you constrain a Claude-based extraction workflow?
Answer guide
Define exactly what you want extracted as a schema, with field names, types and allowed values, and say in the instructions what to do when a value is missing, so the model returns null rather than guessing. Show a few representative examples, including tricky ones, and keep the source text clearly separated from your instructions. Measure field-level accuracy on a labelled set, and re-measure whenever the prompt or model version changes. Remember that document text can contain instructions, so treat it as untrusted.
What this question explores
Whether you can make LLM extraction dependable using schemas, examples, validation and measurement.
Common mistakes
Relying on the prompt alone and not validating the returned fields in code.
Letting the model guess missing values instead of allowing null or a not-found result.
Practise a follow-up
How would you handle a document that contains text trying to change the model's instructions?
Which metrics would you track for extraction quality, and at what level?
7. What checks belong around a Gemini function-calling integration?
Answer guide
Function calling lets the model propose a tool call, but your application decides whether to run it. Validate every argument in code, check the user's authorization against your own rules rather than trusting the model, and require confirmation for actions with side effects or cost. Treat tool results and retrieved text as untrusted, since they can contain injected instructions. Keep secrets out of prompts. API features differ by version, so check the current Gemini documentation when implementing, and add tests for denied and malformed calls.
What this question explores
A security-minded approach to tool use, where the application, not the model, enforces validation and authorization.
Common mistakes
Letting the model's proposed arguments run directly against a database or API without checks.
Giving the model broad tools and relying on the prompt to keep it in bounds.
Practise a follow-up
What extra control would you add before a function call that changes data or spends money?
How would you test that the integration refuses unsafe or malformed tool calls?
8. What would you assess before self-hosting a Llama model?
Answer guide
Start with the model's license terms, to confirm your intended use is allowed, and with a task-based evaluation to see whether it is good enough at your size and language needs. Add safety work that a hosted provider might otherwise handle, such as content filters, prompt-injection defences, logging and access control. Compare total cost including engineering time and idle hardware against a managed API at your expected volume. Self-hosting can help with data control and customisation but adds operations. Decide with a small pilot that measures quality, speed and cost.
What this question explores
A total-cost view of self-hosting an open-weight model: licence, capacity, operations, safety and comparison with hosted options.
Common mistakes
Assuming self-hosting is cheaper because there is no per-token fee.
Skipping the licence review or a task-specific evaluation before committing hardware.
Practise a follow-up
How would you estimate GPU memory needs and how does quantisation change the trade-off?
What operational work do you take on that a managed API would have covered?
9. How would you test a Mistral model for multilingual support?
Answer guide
Build an evaluation set for each language you support, using real task types rather than translated English only, with native or fluent reviewers who can judge fluency, terminology and tone. Include mixed-language text, names, numbers and dates, and domain terms that translate badly. Keep prompts the same across languages first, then try localised prompts to see if they help. Low-resource languages usually need extra attention. Set a minimum quality bar per language before launch, and monitor feedback after release. Check the provider's current model documentation for language coverage.
What this question explores
Evidence-based multilingual evaluation with per-language results and human review.
Common mistakes
Testing only in English or with machine-translated prompts and assuming other languages will match.
Reporting a single average score that hides weak languages.
Practise a follow-up
How would you get reliable human judgement for languages your team does not speak?
Why might token usage and cost differ across languages for the same content?
10. What operational boundaries would you set before using a DeepSeek model in production?
Answer guide
Before production use, confirm exactly how the model is delivered: hosted service or self-run weights, where data is processed and stored, what the provider's terms say about retention and training, and whether that meets your legal, security and regional obligations. Check the licence for your use. Run task-specific evaluations, including safety and prompt-injection tests, and have a fallback and an exit plan in case the service or terms change. Involve security and legal early, and start with low-risk use cases before wider rollout.
What this question explores
Risk-aware adoption of a model: data handling, licence, access limits, evaluation and an exit plan.
Common mistakes
Adopting a model because benchmark results look good without checking data handling or terms.
Allowing broad access and sending sensitive data with no rate limits, logging or review.
Practise a follow-up
What questions would you ask about data retention before sending customer text to a model service?
What would your exit plan look like if you had to replace this model quickly?
11. When does LangChain help, and what should remain explicit in application code?
Answer guide
LangChain gives ready-made pieces for calling models, formatting prompts, using tools and retrieval, and chaining steps, which speeds up a prototype and standardises integrations. The risk is that abstractions hide behaviour and change between versions, making debugging and upgrades harder. Use the framework for plumbing where it saves effort, and wrap it behind your own interface so you can swap it out. Pin versions and read release notes. Ask whether a plain function and an API call would do the job with less code before adding a framework layer.
What this question explores
Judgement about framework benefits versus lock-in and hidden behaviour, with critical logic kept explicit.
Common mistakes
Letting the framework's chains hold authorisation and business rules that nobody can inspect.
Adopting the framework for a single simple model call, adding complexity for no gain.
Practise a follow-up
How would you isolate LangChain behind your own interface to make it replaceable?
How do you debug a chain that gives a wrong answer and shows no error?
12. How would you ingest changing documents with LlamaIndex without serving stale answers?
Answer guide
Give every source document a stable ID and a version or content hash, and record which nodes and embeddings came from it. If you change the parser, chunking or embedding model, plan a full reindex, since mixed settings hurt retrieval. Store the source timestamp in metadata so answers can cite it or prefer the newest. Run scheduled reconciliation that compares the source and the index. For confirmation, ask questions whose answers changed and check the latest content is returned, and monitor the age of indexed data as a metric. Check the LlamaIndex documentation for the current update approach.
What this question explores
Data-freshness design for a retrieval index: identity, versioning, deletion and reindexing.
Common mistakes
Only adding new documents and never deleting or replacing old versions.
Changing chunking or embedding settings without reindexing, leaving mixed data in the index.
Practise a follow-up
How would you handle a deleted source document so it can never be cited again?
When would you choose a full rebuild over incremental updates?
Choose one answer containing an example or practical sequence. Explain what you would actually do, what you would check and when you would ask for help. Keep claims about your experience honest.
For technical or regulated work, check current documentation and applicable local requirements alongside this practice material.