Use this guide to prepare for Cloud or DevOps Engineer interviews, with a focus on automation, availability, change. Explain your reasoning and connect it to experience you can substantiate.
These preparation themes come from the questions in this role’s bank. They help you organise your examples; individual employers may assess different things.
Automation
Availability
Change
Incident
Docker
AWS
A useful preparation sequence
Choose your experience level and the round you expect.
Answer one question in your own words before opening its guide.
Compare your reasoning, evidence and trade-offs; adapt the answer to your experience.
Practise the follow-up, then revisit one answer you want to improve.
Representative questions and answer guidance
Open any question to read its answer. The complete guidance is included on this page.
Behavioural · Mid-level
1. Describe an automation that reduced toil without increasing operational risk.
Answer guide
Choose a real task that ate up time, for example manual server patching or ticket handling. Start by measuring the toil, such as hours per week and how often it caused mistakes. Explain how you thought through the ways the automation could fail. Then describe the safeguards, like testing first, doing small batches, approvals, logging and an easy rollback or stop switch. Share the result in hours saved and fewer errors, and what you learned.
2. How would you design a service to survive the loss of one availability zone?
Answer guide
Spread the service across at least two zones, or preferably three, so one loss leaves enough capacity. Use a load balancer with health checks that quickly stop sending traffic to a failed zone. For data, choose replication across zones and decide how you handle consistency, for example a failover database with clear behaviour. Keep enough spare capacity to take the load. Finally, test it by practising a zone failure in a safe way, and fix gaps you find.
Key points
Cover redundancy, data consistency, health checks and failover tests.
3. How would you set deployment guardrails for teams releasing independently?
Answer guide
Aim for freedom with safety. Give each team clear ownership of their service, including who is on call. Set common standards, such as automated tests, small releases, and canary deployments where a small share of users get the change first. Require good monitoring, with alerts that can stop or roll back a release. Set recovery standards, like fast rollback and runbooks. Build these into shared pipelines so they are easy, and review incidents to improve the rules.
Key points
Define ownership, canaries, telemetry and recovery standards.
4. A release succeeds but latency rises sharply. What do you inspect and change?
Answer guide
Treat it as a release problem until proven otherwise. Compare current latency with the baseline before the release, and see which endpoints or users are affected. Use traces and logs to find where time is being spent, such as a slow database call or a dependency. Check changes in config, traffic and downstream services. Set a clear trigger for rollback, for example if latency stays above target for a set time. If in doubt, roll back first, then investigate.
Key points
Compare baseline, traces, dependencies and rollback trigger.
5. How is a Docker image different from a container?
Answer guide
An image is a read-only, layered package containing the filesystem, application and metadata needed to start a program, built from a Dockerfile. Many containers can be started from one image, much as many programs can run from one installed application. Changes written inside a container are lost when it is removed unless you use volumes. Images are stored in registries and identified by tags and digests. A simple example is building a web app image once and running three containers from it. Use docker ps and docker images to see the difference.
What this question explores
Clear grasp of the template versus instance relationship between images and containers.
Common mistakes
Using the words image and container as if they meant the same thing.
Expecting data written inside a container to persist after the container is deleted.
Practise a follow-up
How do layers and caching affect how quickly you can rebuild an image?
How would you keep database files when a container is replaced?
6. How would you make a web service highly available on AWS across a zone failure?
Answer guide
Run at least two, ideally three, instances of a stateless service in different availability zones behind a load balancer with health checks, using an auto scaling group or a managed container service to replace failed instances. Make sure capacity is enough that the remaining zones can carry the load if one is lost. Plan for the dependencies too, such as DNS, queues and third-party APIs. The trade-off is higher cost and complexity, so match the design to the required availability. Prove it by actually simulating a zone failure in a game day and measuring recovery time.
What this question explores
Practical design of zone-level resilience: redundancy, health checks, state placement and rehearsed failover.
Common mistakes
Placing all instances in one availability zone and calling it highly available.
Leaving the database as a single instance while making the web tier redundant.
Practise a follow-up
How much spare capacity would you keep so a zone loss does not overload the rest?
What is the difference between multi-AZ and multi-region, and when would you need the second?
7. How would you forecast capacity for a seasonal demand peak?
Answer guide
Start with past data, such as last year's peak, growth trends and how much of the current capacity is used. Estimate expected demand using business plans, campaigns and a safety margin. Decide how much headroom you want. Run load tests to find where the system breaks, and fix bottlenecks first. Use autoscaling where possible, and pre-warm what cannot scale quickly. Weigh the cost of extra capacity against the risk of an outage. Review results after the peak.
Key points
Use growth, utilization, headroom, load tests and cost tradeoffs.
8. How would you lead adoption of Ansible across several teams that currently use their own scripts?
Answer guide
I would begin with outcomes the business cares about, such as fewer incidents, faster provisioning and audit readiness, and pick two or three high-value use cases for quick wins. A small platform team sets standards, a shared collection of approved roles, and a central controller with access control. Then I train champions in each team and migrate scripts gradually rather than forcing a big switch. I track measures like time to deploy, failed change rate and hours saved, and share them. Resistance is usually about fear of lost ownership, so I keep teams contributing to the shared content and give them credit.
What this question explores
Whether you can lead change by linking automation to business outcomes, people and governance, not just tooling.
Common mistakes
Mandating a tool without addressing team concerns, training or showing measurable benefits early.
Building a central team that owns everything, creating a bottleneck instead of enabling contributions.
Practise a follow-up
How would you measure the return on this programme?
9. What would you check before moving a service to Azure Kubernetes Service?
Answer guide
Start with whether Kubernetes is needed at all, since a simpler managed option may serve a small service better. Also decide who operates it, because the cluster adds platform work such as patching, policy and on-call. Check that the application is containerised properly, is stateless or has storage planned, and has health probes and resource requests. Run a pilot with a non-critical workload and compare cost, effort and reliability before moving anything important. Confirm current options in the Azure documentation.
What this question explores
Whether you weigh operational readiness and alternatives, not only technical feasibility, before adopting Kubernetes.
Common mistakes
Assuming Kubernetes is the default answer without comparing simpler hosting options.
Focusing only on deploying containers and skipping identity, networking, upgrades and ownership.
Practise a follow-up
What would make you choose a simpler service over AKS for this workload?
How would you plan cluster upgrades without disturbing running applications?
10. When would Cloud Run be a better fit than managing a Kubernetes cluster?
Answer guide
Cloud Run suits a stateless, containerised service that handles requests or events and can scale with demand, including down to zero. You hand over a container and the platform manages servers, scaling and much of the networking, so there is no cluster to patch or size. Trade-offs to weigh are cold starts, request time limits, concurrency settings and less control over the environment. Costs also differ by traffic shape. Before choosing, I would list the service's requirements against the platform limits in current documentation and run a small load test.
What this question explores
Ability to match a workload to a managed serverless container platform and to know where Kubernetes is still justified.
Common mistakes
Choosing Kubernetes for a simple stateless API because it is what everyone uses.
Ignoring cold starts, timeouts or concurrency behaviour when moving to Cloud Run.
Practise a follow-up
How would you reduce the impact of cold starts on a latency-sensitive endpoint?
What kind of workload would push you back to Kubernetes?
11. What is the difference between a Deployment and a Service in Kubernetes?
Answer guide
A Deployment describes the desired state for identical Pods, such as the image and replica count, and manages rolling updates and rollbacks by replacing Pods gradually. A Service gives clients a stable name and virtual address and spreads traffic across the Pods its label selector matches, because Pod IPs change whenever Pods are replaced. The two connect only through matching labels, so a typo in a selector gives a Service with no endpoints. Use ClusterIP for internal traffic and a LoadBalancer or ingress for outside traffic; when troubleshooting, compare the selector with the Pod labels.
What this question explores
Whether you separate workload management from networking in Kubernetes and know how the two link.
Common mistakes
Thinking a Service is what creates or restarts the Pods.
Not knowing that label selectors link a Service to Pods, so mismatched labels route no traffic.
Practise a follow-up
What happens to traffic during a rolling update if readiness probes are missing?
When would you use a ClusterIP service and when a LoadBalancer or ingress?
12. Why is Terraform state sensitive, and how should a team manage it?
Answer guide
Terraform state records which real resources correspond to your configuration, and it can contain sensitive values such as passwords or keys in plain text, so anyone who can read it may learn secrets. It also drives every plan and apply, so losing or corrupting it can cause duplicate or orphaned infrastructure. Store it in a remote backend with encryption, strict access control, versioning or backups, and locking to stop two people applying at once. Keep it out of version control. Restrict who may run apply, and run it through a reviewed pipeline.
What this question explores
Understanding of state as a critical and sensitive asset, and practices for storing and sharing it safely.
Common mistakes
Committing the state file to the Git repository along with the code.
Using local state on one laptop so teammates cannot safely collaborate or recover it.
Practise a follow-up
What does state locking prevent, and what happens if a lock is left behind?
How would you split state for many environments to limit the impact of a mistake?
Choose one answer containing an example or practical sequence. Explain what you would actually do, what you would check and when you would ask for help. Keep claims about your experience honest.
For technical or regulated work, check current documentation and applicable local requirements alongside this practice material.