Skip to content
Interviewpedia™

Topic preparation guide

Cloud and DevOps interview questions and answers

Prepare to connect cloud infrastructure and delivery tooling to a reliable operating outcome. The questions cover containers, infrastructure state, release controls and observability across several platforms; they are not a vendor-certification syllabus.

What interviewers are assessing

  • Reproducible delivery: explain how the same reviewed artifact moves through environments.
  • Operational safety: reason about access, secrets, state, rollback and recovery.
  • Evidence: use service behaviour and useful telemetry rather than a green pipeline alone.

How to approach your answer

  1. Clarify the workload, deployment constraints and the service objective.
  2. Explain build, configuration and deployment boundaries, including approvals and secret handling.
  3. Describe validation, rollback feasibility and the telemetry that would confirm a successful release.

Promote the same build through test and production.

Illustrative approach: I would build and test an immutable artifact once, give it an identifiable version and promote that artifact through environments. Configuration and secrets would be supplied through controlled environment-specific mechanisms rather than rebuilding different binaries. Each deployment would record the artifact and configuration versions, apply the relevant checks and run service-facing validation. I would also define a rollback or forward-recovery path before promotion. A successful pipeline is useful evidence, but it does not prove that customers can use the service.

Mistakes to avoid

  • Rebuilding a different artifact for each environment.
  • Treating infrastructure state or pipeline credentials as ordinary project files.
  • Using high-cardinality metric labels without considering cost and reliability.

Questions and answer guidance

Start with the level closest to your experience. Each question links to its exact practice exercise; the answer is also available here without opening the app.

Foundations

Start with the concepts and explain them using a small example.

Technical · Fresher

1. How is a Docker image different from a container?

Read the answer guide

An image is a read-only, layered package containing the filesystem, application and metadata needed to start a program, built from a Dockerfile. Many containers can be started from one image, much as many programs can run from one installed application. Changes written inside a container are lost when it is removed unless you use volumes. Images are stored in registries and identified by tags and digests. A simple example is building a web app image once and running three containers from it. Use docker ps and docker images to see the difference.

What the interviewer is assessing

Clear grasp of the template versus instance relationship between images and containers.

Common mistakes

  • Using the words image and container as if they meant the same thing.
  • Expecting data written inside a container to persist after the container is deleted.

Practise a follow-up

  • How do layers and caching affect how quickly you can rebuild an image?
  • How would you keep database files when a container is replaced?
Practise this question →
Technical · Fresher

2. What is the difference between merge and rebase in Git?

Read the answer guide

Merge joins two branches by creating a merge commit that preserves the real history, including when work happened in parallel. Rebase takes your commits and replays them on top of another base, giving a straight, tidy history but creating new commits with new IDs, so it rewrites history. That is fine for your own local branch before sharing, but rebasing commits others have already pulled forces them to fix their copies. Neither is always better; agree a policy. If a rebase goes wrong, git reflog usually lets you recover the previous state.

What the interviewer is assessing

Basic Git understanding of history rewriting and when rebase is safe compared with merge.

Common mistakes

  • Rebasing a branch that teammates have already pulled and pushed.
  • Thinking merge and rebase produce the same commit history.

Practise a follow-up

  • What is git reflog and how could it rescue you after a bad rebase?
  • What does squash merging do to history, and when would you prefer it?
Practise this question →
Technical · Fresher

3. What are logs, metrics and traces, and how do they differ in observability?

Read the answer guide

Logs are timestamped records of discrete events, such as an error message with context. Metrics are numeric measurements aggregated over time, like request count or CPU usage, and are cheap to store and good for trends and alerting. Traces follow a single request as it moves through services, broken into spans with timings, so I can see where latency or failures occur. I use metrics to notice a problem, traces to locate the slow or failing service, and logs to understand the exact cause. They work best together when linked by shared identifiers.

What the interviewer is assessing

Whether you understand that each signal answers a different question and that they are used together rather than as alternatives.

Common mistakes

  • Candidates say logs are enough and ignore how costly searching logs is for trends.
  • Candidates describe the three signals but cannot say when to use each one.

Practise a follow-up

  • Which signal would you check first when an alert fires, and why?
  • How would you link a log line to a trace?
Practise this question →

Applied decisions

Show how you would apply the idea to a constraint, disagreement or failure.

Technical · Mid-level

4. How would you make a web service highly available on AWS across a zone failure?

Read the answer guide

Run at least two, ideally three, instances of a stateless service in different availability zones behind a load balancer with health checks, using an auto scaling group or a managed container service to replace failed instances. Make sure capacity is enough that the remaining zones can carry the load if one is lost. Plan for the dependencies too, such as DNS, queues and third-party APIs. The trade-off is higher cost and complexity, so match the design to the required availability. Prove it by actually simulating a zone failure in a game day and measuring recovery time.

What the interviewer is assessing

Practical design of zone-level resilience: redundancy, health checks, state placement and rehearsed failover.

Common mistakes

  • Placing all instances in one availability zone and calling it highly available.
  • Leaving the database as a single instance while making the web tier redundant.

Practise a follow-up

  • How much spare capacity would you keep so a zone loss does not overload the rest?
  • What is the difference between multi-AZ and multi-region, and when would you need the second?
Practise this question →
Technical · Mid-level

5. What would you check before moving a service to Azure Kubernetes Service?

Read the answer guide

Start with whether Kubernetes is needed at all, since a simpler managed option may serve a small service better. Also decide who operates it, because the cluster adds platform work such as patching, policy and on-call. Check that the application is containerised properly, is stateless or has storage planned, and has health probes and resource requests. Run a pilot with a non-critical workload and compare cost, effort and reliability before moving anything important. Confirm current options in the Azure documentation.

What the interviewer is assessing

Whether you weigh operational readiness and alternatives, not only technical feasibility, before adopting Kubernetes.

Common mistakes

  • Assuming Kubernetes is the default answer without comparing simpler hosting options.
  • Focusing only on deploying containers and skipping identity, networking, upgrades and ownership.

Practise a follow-up

  • What would make you choose a simpler service over AKS for this workload?
  • How would you plan cluster upgrades without disturbing running applications?
Practise this question →
Technical · Mid-level

6. When would Cloud Run be a better fit than managing a Kubernetes cluster?

Read the answer guide

Cloud Run suits a stateless, containerised service that handles requests or events and can scale with demand, including down to zero. You hand over a container and the platform manages servers, scaling and much of the networking, so there is no cluster to patch or size. Trade-offs to weigh are cold starts, request time limits, concurrency settings and less control over the environment. Costs also differ by traffic shape. Before choosing, I would list the service's requirements against the platform limits in current documentation and run a small load test.

What the interviewer is assessing

Ability to match a workload to a managed serverless container platform and to know where Kubernetes is still justified.

Common mistakes

  • Choosing Kubernetes for a simple stateless API because it is what everyone uses.
  • Ignoring cold starts, timeouts or concurrency behaviour when moving to Cloud Run.

Practise a follow-up

  • How would you reduce the impact of cold starts on a latency-sensitive endpoint?
  • What kind of workload would push you back to Kubernetes?
Practise this question →
Technical · Mid-level

7. What is the difference between a Deployment and a Service in Kubernetes?

Read the answer guide

A Deployment describes the desired state for identical Pods, such as the image and replica count, and manages rolling updates and rollbacks by replacing Pods gradually. A Service gives clients a stable name and virtual address and spreads traffic across the Pods its label selector matches, because Pod IPs change whenever Pods are replaced. The two connect only through matching labels, so a typo in a selector gives a Service with no endpoints. Use ClusterIP for internal traffic and a LoadBalancer or ingress for outside traffic; when troubleshooting, compare the selector with the Pod labels.

What the interviewer is assessing

Whether you separate workload management from networking in Kubernetes and know how the two link.

Common mistakes

  • Thinking a Service is what creates or restarts the Pods.
  • Not knowing that label selectors link a Service to Pods, so mismatched labels route no traffic.

Practise a follow-up

  • What happens to traffic during a rolling update if readiness probes are missing?
  • When would you use a ClusterIP service and when a LoadBalancer or ingress?
Practise this question →
Technical · Mid-level

8. Why is Terraform state sensitive, and how should a team manage it?

Read the answer guide

Terraform state records which real resources correspond to your configuration, and it can contain sensitive values such as passwords or keys in plain text, so anyone who can read it may learn secrets. It also drives every plan and apply, so losing or corrupting it can cause duplicate or orphaned infrastructure. Store it in a remote backend with encryption, strict access control, versioning or backups, and locking to stop two people applying at once. Keep it out of version control. Restrict who may run apply, and run it through a reviewed pipeline.

What the interviewer is assessing

Understanding of state as a critical and sensitive asset, and practices for storing and sharing it safely.

Common mistakes

  • Committing the state file to the Git repository along with the code.
  • Using local state on one laptop so teammates cannot safely collaborate or recover it.

Practise a follow-up

  • What does state locking prevent, and what happens if a lock is left behind?
  • How would you split state for many environments to limit the impact of a mistake?
Practise this question →
Technical · Mid-level

9. How would you protect a repository's main branch while keeping delivery fast?

Read the answer guide

Protect the branch with rules that need a pull request, at least one review, and passing automated checks such as tests and linting before merge, and stop direct pushes and force pushes. To keep it fast, make pull requests small, automate the slow checks, use code owners so the right people review, and set review expectations such as same-day turnaround. Too many mandatory checks or slow pipelines push teams to work around controls. Track lead time and change failure rate to see whether protection helps, and adjust the rules with the team.

What the interviewer is assessing

Ability to balance governance and delivery speed, with controls that people will follow.

Common mistakes

  • Locking everything down with slow manual approvals so people look for workarounds.
  • Protecting the branch on paper but allowing admins to push directly without any record.

Practise a follow-up

  • How do you handle a genuine emergency fix without losing traceability?
  • Which metrics would show that your branch rules are slowing delivery?
Practise this question →
Technical · Mid-level

10. How would you keep a Jenkins pipeline reproducible and secure?

Read the answer guide

Store the pipeline as a Jenkinsfile in the repository so changes are reviewed and versioned like code. For security, keep secrets in the credential store and inject them with credential bindings so they are masked in logs, give jobs the least permissions, keep the controller free of builds, and run untrusted pull request builds in isolated agents without secrets. Update Jenkins and plugins regularly, since plugin vulnerabilities are a common risk. Done well, two builds of the same commit produce the same artifact, and a new agent can be created from scratch quickly.

What the interviewer is assessing

Ability to make CI reproducible and secure: pipeline as code, pinned inputs, isolated agents and managed credentials.

Common mistakes

  • Configuring jobs by hand in the web interface so nobody can reproduce or review them.
  • Putting passwords in job parameters or scripts, or running untrusted code with production credentials.

Practise a follow-up

  • Why should the Jenkins controller not run build jobs itself?
  • How would you handle plugin updates without breaking many pipelines?
Practise this question →
Technical · Mid-level

11. How would you promote the same build artifact through test and production in Azure Pipelines?

Read the answer guide

Build once and promote the same artifact, so what you tested is exactly what reaches production. The build stage produces a versioned, immutable artifact, such as a container image tagged with the commit or build number, and stores it in a registry or artifact feed. Attach approvals and checks to the production environment, and keep deployment history for audit. Rebuilding for each environment risks subtle differences between test and production. Roll back by redeploying an earlier known artifact, and rehearse that in a lower environment.

What the interviewer is assessing

Understanding of build-once, deploy-many promotion with configuration kept outside the artifact.

Common mistakes

  • Rebuilding the application for each environment so production runs code that was never tested.
  • Baking environment settings or secrets into the artifact.

Practise a follow-up

  • How do you supply different configuration and secrets per environment in Azure Pipelines?
  • How would you roll back a bad production release using stored artifacts?
Practise this question →
Technical · Mid-level

12. How do Prometheus labels affect a high-cardinality metric?

Read the answer guide

In Prometheus each unique combination of metric name and label values is its own time series, so labels multiply. Use labels for small, bounded dimensions such as method, status class, service and region. Put high-detail identifiers in logs or traces rather than in metrics, and normalise paths to route templates. To control this, review new metrics before release, watch the series count and set limits where available. If you really need per-customer detail, aggregate the top customers only.

What the interviewer is assessing

Understanding of why label cardinality drives Prometheus cost and how to choose labels responsibly.

Common mistakes

  • Adding user ID or request ID as a label because it looks useful for debugging.
  • Thinking labels are free and only the number of metric names matters.

Practise a follow-up

  • How would you find which metric is creating the most time series?
  • When would you use logs or traces instead of a metric label?
Practise this question →

A 30-minute practice plan

  1. 10 minutes: Explain images, containers and a simple deployment path.
  2. 10 minutes: Practise a state or credentials scenario and describe the control you would apply.
  3. 10 minutes: Design release validation that checks customer-visible behaviour and recovery.

Answer before reading the guide. Use feedback to improve the substance, then rehearse a follow-up without memorising the wording.

Further reading

Use these primary references to check concepts and current platform behaviour alongside the practice bank.

Make it your language

Language settings are saved on this device only.

Core interface translations are available. Some extended guidance and legal text remain in English.

Public guides remain in English where a translation is unavailable.

Voice availability depends on your browser and device. You can always type instead.

Open Library in your language