Skip to content
Interviewpedia™

Topic preparation guide

System design interview questions and answers

Build a design answer from requirements, constraints and a simple working architecture. This set progresses from scaling and caching fundamentals to capacity estimation and a URL-shortener design, without assuming every service needs microservices.

What interviewers are assessing

  • Requirements and estimates: make traffic, latency, availability and data assumptions explicit.
  • Trade-offs: explain why a component is needed and what new failure modes it introduces.
  • Evolution: start with a workable design, then show what you would measure before adding complexity.

How to approach your answer

  1. Clarify users, core operations and constraints. State a rough load estimate and which assumptions need validation.
  2. Sketch the simplest request and data paths. Discuss consistency, errors, security and operational visibility.
  3. Identify the first likely bottleneck and explain a measured scaling step, including how you would test failure behaviour.

Design a URL shortener.

Illustrative approach: I would clarify whether users can choose aliases, how long links remain valid and whether redirects need analytics. A first design stores a unique short code and destination, exposes a creation endpoint and serves redirects. I would enforce uniqueness in storage and validate destinations. If reads dominate, a cache can reduce repeated lookups, but I would define invalidation and expiry behaviour. Before adding sharding or multi-region writes, I would estimate traffic and explain the availability, consistency and abuse-prevention needs that justify them.

Mistakes to avoid

  • Drawing components before clarifying requirements.
  • Claiming that a cache or queue automatically guarantees reliability.
  • Ignoring data consistency, abuse controls or failure recovery.

Questions and answer guidance

Start with the level closest to your experience. Each question links to its exact practice exercise; the answer is also available here without opening the app.

Foundations

Start with the concepts and explain them using a small example.

Technical · Fresher

1. How would you begin when asked to design a system you have never built before?

Read the answer guide

I would not start drawing boxes. First I ask clarifying questions to pin down the core features, the users, and what is out of scope. Then I ask about scale, such as daily active users and read to write ratio, and about non functional needs like latency, availability and consistency. I write these down as agreed requirements. Next I sketch a simple high level design, walk through one main request end to end, and only then go deeper into the parts that matter most, such as the database or the cache. I state trade-offs aloud as I go.

What the interviewer is assessing

Whether you gather requirements and structure your thinking before jumping into solutions, which is what real design work demands.

Common mistakes

  • Jumping straight to a favourite technology stack without asking what the system must actually do.
  • Drawing a huge diagram quickly and never explaining why each component is there.

Practise a follow-up

  • What requirements would you clarify first for a ride booking app?
  • How do you decide which part of the design to go deep on?
Practise this question →
Technical · Fresher

2. What is the difference between vertical and horizontal scaling, and when would you pick each?

Read the answer guide

Vertical scaling means giving one machine more CPU, memory or disk. It is simple because the application does not change, but it has a hardware ceiling, costs rise steeply, and the machine remains a single point of failure. Horizontal scaling means adding more machines and spreading the work across them. It scales further and improves availability, but it needs a load balancer and usually stateless services or partitioned data, which adds complexity. For a small internal tool I would scale vertically first. For a public service expecting growth and needing high availability, I would design for horizontal scaling from the start.

What the interviewer is assessing

Whether you understand the basic scaling options and can match them to the situation rather than defaulting to one.

Common mistakes

  • Saying horizontal scaling is always better without mentioning the added complexity of state and coordination.
  • Forgetting that a single big machine is still a single point of failure.

Practise a follow-up

  • What makes a web service easy to scale horizontally?
  • Where does a database fit into this when you add more servers?
Practise this question →
Technical · Fresher

3. What does a load balancer do, and why do we place one in front of servers?

Read the answer guide

A load balancer receives incoming requests and distributes them across a pool of backend servers. This spreads the load so that no single server is overwhelmed, and it lets us add or remove servers without clients noticing. It also runs health checks and stops sending traffic to servers that fail, which improves availability. Common algorithms are round robin, least connections and hashing on a client attribute. Some load balancers also terminate TLS and can route by URL path. The load balancer itself must not become a single point of failure, so it is usually deployed as a redundant pair or as a managed service.

What the interviewer is assessing

Whether you know the purpose of a load balancer and its role in availability, not just in spreading traffic.

Common mistakes

  • Describing only traffic distribution and ignoring health checks and failover.
  • Forgetting that the load balancer can itself become a single point of failure.

Practise a follow-up

  • How does round robin differ from least connections?
  • What happens to user sessions if requests hit different servers?
Practise this question →
Technical · Fresher

4. What is caching, and why does it improve the performance of a system?

Read the answer guide

Caching means keeping a copy of frequently used data in a faster storage layer, usually memory, so that repeated requests avoid the slower original source such as a database or a remote service. This reduces latency and also reduces load on the backend, which lets the system handle more traffic. Caches work well when data is read often and changes rarely. The cost is that cached data can become stale, so we need an expiry time or an invalidation strategy. We also have to decide what to evict when the cache is full, and a common policy is least recently used.

What the interviewer is assessing

Whether you grasp what caching buys you and the main cost it introduces, which is stale data.

Common mistakes

  • Treating caching as free and never mentioning stale data or invalidation.
  • Caching everything instead of data that is read often and expensive to compute.

Practise a follow-up

  • Where in a web application can you place a cache?
  • What does a cache hit ratio tell you?
Practise this question →
Technical · Fresher

5. How would you choose between a relational database and a NoSQL database for a new application?

Read the answer guide

I start from the data and the access patterns. A relational database suits structured data with relationships, where we need transactions, joins and a strict schema, such as orders and payments. NoSQL databases come in different kinds, like key value, document, wide column and graph, and are chosen when the access pattern is simple and known, the schema changes often, or the scale needs easy horizontal partitioning. Many NoSQL stores trade away joins or full transactions for scale and flexibility, though this varies by product. If I am unsure, I default to a relational database because it is well understood and flexible.

What the interviewer is assessing

Whether you choose storage based on data shape and access patterns instead of fashion.

Common mistakes

  • Claiming NoSQL is always faster or more scalable than a relational database.
  • Choosing a database by popularity without describing the queries the application will run.

Practise a follow-up

  • Which kind of NoSQL store would you use for a shopping cart?
  • Can a relational database scale beyond one machine?
Practise this question →
Technical · Fresher

6. What is the difference between latency and throughput in a system?

Read the answer guide

Latency is the time one request takes from sending to receiving a response, usually measured in milliseconds. Throughput is how many requests or how much data the system handles in a given time, for example requests per second. They are related but different: a system can have high throughput and still have slow individual requests, for example by batching work. When talking about latency I prefer percentiles such as the 95th or 99th rather than averages, because a few slow requests hurt users and the average hides them. Under heavy load, latency usually rises sharply as queues build up.

What the interviewer is assessing

Whether you can define core performance terms precisely and know why averages can mislead.

Common mistakes

  • Using the two words interchangeably as if they meant the same thing.
  • Quoting only average latency and ignoring tail percentiles.

Practise a follow-up

  • Why do we look at p99 latency instead of the average?
  • Can you improve throughput without improving latency?
Practise this question →
Technical · Fresher

7. What is a message queue, and why would a system use one between two services?

Read the answer guide

A message queue is a buffer that holds messages sent by a producer until a consumer processes them. It decouples the two services, so the producer does not wait for the consumer and neither has to be available at the same moment. It smooths traffic spikes because messages wait in the queue instead of overwhelming the consumer. It also allows several consumers to share the work and supports retries when processing fails. The trade-offs are extra infrastructure, eventual rather than immediate results, and the need to handle duplicates and ordering. Typical uses include sending emails, resizing images and processing orders in the background.

What the interviewer is assessing

Whether you understand asynchronous decoupling and the costs that come with queues.

Common mistakes

  • Describing a queue only as a faster way to call a service.
  • Ignoring that consumers may receive the same message more than once.

Practise a follow-up

  • What happens to a message that keeps failing?
  • When would you avoid a queue and use a direct call?
Practise this question →
Technical · Fresher

8. Why do public APIs use rate limiting, and what happens when a client exceeds the limit?

Read the answer guide

Rate limiting caps how many requests a client can make in a period. It protects the service from overload, whether from a buggy client loop or a deliberate attack, keeps usage fair between customers, and helps control cost. It also supports paid tiers with different limits. When a client exceeds the limit, the server normally rejects the request with HTTP status 429, Too Many Requests, and often includes a header telling the client when to retry. Well written clients then back off and retry later. The limit is usually keyed by an API key, a user or an IP address, depending on the need.

What the interviewer is assessing

Whether you know why rate limiting exists and how a typical API communicates it to clients.

Common mistakes

  • Saying rate limiting is only for stopping hackers and not for fairness and protection.
  • Not knowing the standard response is HTTP 429 and that clients should back off.

Practise a follow-up

  • What would you use as the key for limiting requests?
  • How should a client behave after receiving a 429?
Practise this question →

Applied decisions

Show how you would apply the idea to a constraint, disagreement or failure.

Technical · Mid-level

9. How do you do a quick capacity estimate for a service that expects ten million daily users?

Read the answer guide

I turn users into rates. If each user makes about twenty requests a day, that is two hundred million requests, and dividing by roughly eighty-six thousand seconds gives an average near two to three thousand requests per second. Peaks are usually a few times the average, so I plan for that multiple. For storage I multiply records per day by record size and by retention. I also estimate read to write ratio, bandwidth and cache size. I round numbers aggressively, state assumptions aloud, and use the result to decide whether one database or several servers are even needed.

What the interviewer is assessing

Whether you can reason from users to requests, storage and peak load with simple arithmetic and stated assumptions.

Common mistakes

  • Using the average rate and forgetting that peak traffic can be several times higher.
  • Hiding behind exact numbers instead of stating rounded, clearly labelled assumptions.

Practise a follow-up

  • How would you estimate storage for five years of data?
  • How does the read to write ratio change your design?
Practise this question →
Technical · Mid-level

10. Can you explain cache-aside, write-through and write-back caching, and when each is appropriate?

Read the answer guide

In cache-aside the application checks the cache first, on a miss reads the database and fills the cache. It is simple and tolerates cache failure, but data can go stale. In write-through every write goes to the cache and the database together, so reads are fresh, at the cost of slower writes. In write-back the write lands in the cache and is flushed to the database later, which is fast but risks losing data if the cache node dies before flushing. I use cache-aside for most read-heavy services, write-through when freshness matters, and write-back only when some loss is tolerable.

What the interviewer is assessing

Whether you know the main caching write policies and can state their consistency and durability trade-offs.

Common mistakes

  • Recommending write-back without mentioning the risk of losing unflushed writes.
  • Describing the three patterns as identical because all of them use a cache.

Practise a follow-up

  • How would you invalidate a cache entry when the row changes?
  • What eviction policy would you start with?
Practise this question →
Technical · Mid-level

11. What is a cache stampede, and how would you protect a database from one?

Read the answer guide

A stampede happens when a popular cache entry expires and many requests miss at once, all hitting the database to rebuild the same value. The database can be overloaded, which makes rebuilding even slower. Defences include request coalescing, where only one request rebuilds while others wait or get the old value, a short lock per key, and adding random jitter to expiry times so entries do not expire together. For very hot keys I refresh in the background before expiry and serve slightly stale data meanwhile. Pre-warming the cache after a deploy or restart also helps, because a cold cache invites the same problem.

What the interviewer is assessing

Whether you recognise a classic caching failure mode and know several practical mitigations.

Common mistakes

  • Setting the same expiry on every key so they all expire together.
  • Suggesting a bigger database as the only fix instead of controlling duplicate rebuilds.

Practise a follow-up

  • What are the downsides of serving stale data while refreshing?
  • How would you find the hot keys in production?
Practise this question →

Senior judgement

Explain trade-offs, wider consequences and the evidence behind your decision.

Technical · Senior

12. How would you design a URL shortener?

Read the answer guide

Clarify the requirements first: expected traffic, link lifetime, custom aliases and analytics. The core is a write path that creates a short key and a much heavier read path that redirects. Generate keys with a counter encoded in base 62, or with random keys checked for collisions, and store the key-to-URL mapping in a key-value or relational store. Put a cache in front of reads, choose a permanent or temporary redirect depending on whether you need click analytics, and add rate limiting and expiry. Then discuss scaling: partitioning the store and making key generation work across many servers.

What the interviewer is assessing

Structured system design: requirements, data model, trade-offs and scaling.

Common mistakes

  • Jumping to technology before clarifying requirements
  • Ignoring that reads far outnumber writes
  • Not discussing key collisions

Practise a follow-up

  • How would you handle a link that suddenly goes viral?
  • How would you prevent abuse?
Practise this question →

A 30-minute practice plan

  1. 10 minutes: Practise a two-minute requirements clarification and a rough traffic estimate.
  2. 10 minutes: Draw a small request path and explain one storage and one cache decision.
  3. 10 minutes: Answer the failure follow-up: what happens when a dependency times out?

Answer before reading the guide. Use feedback to improve the substance, then rehearse a follow-up without memorising the wording.

Further reading

Use these primary references to check concepts and current platform behaviour alongside the practice bank.

Make it your language

Language settings are saved on this device only.

Core interface translations are available. Some extended guidance and legal text remain in English.

Public guides remain in English where a translation is unavailable.

Voice availability depends on your browser and device. You can always type instead.

Open Library in your language