Skip to content
Interviewpedia™

Role preparation guide

Solution Architect / Design Lead interview preparation

Use this guide to prepare for Solution Architect / Design Lead interviews, with a focus on problem framing, solution design, options assessment. Explain your reasoning and connect it to experience you can substantiate.

Technical roundSituation roundManagerial round

What to prepare

These preparation themes come from the questions in this role’s bank. They help you organise your examples; individual employers may assess different things.

  • Problem framing
  • Solution Design
  • Options assessment
  • Nonfunctional requirements
  • Integration choice
  • Data ownership

A useful preparation sequence

  1. Choose your experience level and the round you expect.
  2. Answer one question in your own words before opening its guide.
  3. Compare your reasoning, evidence and trade-offs; adapt the answer to your experience.
  4. Practise the follow-up, then revisit one answer you want to improve.

Representative questions and answer guidance

Open any question to read its answer. The complete guidance is included on this page.

Technical · Mid-level

1. What must be understood before drawing a target architecture?

Answer guide

Drawing a target architecture too early risks solving the wrong problem. I would first clarify the business outcomes and the users' journeys, then the constraints such as budget, timeline, regulation and skills. I need to understand the current state, the existing interfaces and the data involved, along with volumes and the quality attributes that matter, like availability, security and response time. Who has decision authority also matters, because a design that nobody can approve stalls. With these, the architecture can be justified against clear needs instead of being an attractive diagram.

What this question explores

Whether you gather outcomes, constraints and current-state facts before proposing a design.

Common mistakes

  • Starting with a favourite technology and drawing boxes around it.
  • Ignoring the current interfaces and data volumes until integration begins.

Practise a follow-up

  • Which quality attributes would you ask about first, and why?
  • How do you document assumptions that you cannot confirm yet?
Practise this question →
Managerial · Senior

2. How would you prevent architecture governance from becoming a bottleneck for routine changes?

Answer guide

The aim is proportionate governance. I would set review thresholds by impact, reversibility and whether a change crosses shared boundaries, so a small reversible change within one team's area needs no board. Reusable decisions, such as approved patterns, and delegated authority to trusted teams let routine changes move quickly. Deeper review stays for material risks, like new external integrations or irreversible data choices. I would also check whether governance works, by tracking lead time for decisions and whether reviews prevented problems, rather than counting approvals. If the board becomes the queue everyone waits in, people will route around it, which is worse.

What this question explores

Whether you can design proportionate governance with impact-based thresholds, delegation and reusable decisions, and monitor whether it improves outcomes instead of just counting approvals.

Common mistakes

  • They require full architecture board review for every change, so routine work queues up and teams bypass governance.
  • They remove review entirely for speed, so material risks and shared-boundary changes go unexamined.

Practise a follow-up

  • How would you decide which teams earn delegated authority, and how would you withdraw it?
  • What measures would show governance is helping rather than slowing delivery?
Practise this question →
Situation · Senior

3. Build, buy and extend are all plausible. How would you compare them?

Answer guide

I would agree the decision criteria with stakeholders before scoring the options: functional fit, integration effort, security and compliance, total lifecycle cost, time to value, skills and capability, vendor lock-in and exit cost. Each option is assessed against them with evidence such as a proof of concept, not just opinion, and assumptions are written down. Extend is often the middle path, but it can hide accumulated customisation costs. I present the trade-offs and a recommendation, and note what would change it, so the decision can be revisited if assumptions prove wrong.

What this question explores

Whether you run a transparent, criteria-based comparison that records assumptions and exit costs.

Common mistakes

  • Choosing build or buy on preference or licence cost alone.
  • Scoring options without agreeing the criteria or weights beforehand.

Practise a follow-up

  • How would you weigh short-term speed against long-term lock-in?
  • What would you include in a proof of concept for the buy option?
Practise this question →
Technical · Fresher

4. Before designing a claims-processing solution, which business facts would you establish?

Answer guide

I would start with the business, not the technology. For claims processing that means the claimant journey from first notice to payment, the decision rules adjusters apply, expected volumes and peaks, service expectations, legal or policy constraints, and the systems already holding claim data. I also list uncertainties that could change the architecture, such as how many claims need manual review. These inputs let me explain every design choice and later check the outcome against them. Starting with a preferred technology diagram feels quick, but it hides assumptions and makes the design hard to defend when the facts turn out different.

What this question explores

Whether you begin from claimant journeys, rules, volumes and constraints, and flag uncertainties, instead of jumping straight to a favourite technology diagram.

Common mistakes

  • They open with a technology stack or architecture diagram before asking anything about how claims are actually handled.
  • They list generic requirements but never name the uncertainties that could change the architecture or how they will be resolved.

Practise a follow-up

  • Which of these inputs would you confirm first if time were short?
  • What would you do if the business could not give reliable claim volume figures?
Practise this question →
Technical · Mid-level

5. How do you make 'high availability' testable?

Answer guide

High availability means nothing until it is measurable. I would define the service boundary, then an availability objective such as a percentage over a stated window, and the way it is measured, for example successful requests over total. Then I list the failure modes to survive, like an instance, a zone or a dependency failing, with recovery time and data loss targets. Dependencies' own availability limits what is achievable, so I state assumptions. A test plan then uses failover drills or fault injection to show the design meets the objective, and monitoring tracks it in production.

What this question explores

Whether you convert a vague quality wish into an objective, a measurement method and failure tests.

Common mistakes

  • Writing the requirement as 99.99 percent uptime with no boundary, window or measurement.
  • Ignoring dependencies whose availability caps the whole service.

Practise a follow-up

  • How do error budgets help balance reliability and feature delivery?
  • What is the difference between recovery time and recovery point objectives?
Practise this question →
Technical · Senior

6. How do you choose synchronous API versus asynchronous event integration?

Answer guide

I start with what the business needs: does the caller need an answer now, or can the work complete later? Synchronous APIs are simpler to reason about and suit immediate responses, but they couple availability, so one slow service hurts the caller. Asynchronous events decouple and absorb bursts, but add eventual consistency, ordering questions, retries and duplicate handling. Whichever I choose, I define who owns the data and make consumers idempotent so a repeated message is harmless. Often a hybrid works: a synchronous acknowledgement followed by asynchronous processing.

What this question explores

Whether you choose an integration style from consistency, coupling and failure needs and design for retries.

Common mistakes

  • Choosing events because they are modern, without handling ordering, duplicates or eventual consistency.
  • Using synchronous calls for everything and chaining services so one failure cascades.

Practise a follow-up

  • What is idempotency and how would you implement it for a payment message?
  • How would you monitor an event-driven flow for stuck or lost messages?
Practise this question →
Situation · Senior

7. Two services both want to own the customer master record. What do you decide?

Answer guide

Two owners of one record leads to conflicting updates and endless reconciliation. I would decide a single system of record for each attribute, and a clear authority about who may update what. Other services keep read-only copies fed by published events or an API contract, with rules for how conflicts and delays are handled and a scheduled reconciliation to detect drift. Dual writes without a consistency design are the danger, since a failure between them leaves the systems disagreeing. I would take the decision to the business data owner, document it, and explain what each team gains.

What this question explores

Whether you enforce single ownership with clear contracts and reconciliation rather than dual writes.

Common mistakes

  • Letting both services update the record and hoping they will stay in sync.
  • Settling the dispute by technology preference with no business data owner involved.

Practise a follow-up

  • How would you migrate ownership from one service to another safely?
  • What would you do about a customer attribute that both teams truly need to change?
Practise this question →
Technical · Mid-level

8. Where do authentication and authorization decisions belong in a solution design?

Answer guide

I begin by drawing trust boundaries and deciding who the identity provider is, so authentication is done once, by a trusted source, and services rely on its tokens. Authorisation is the separate question of what an authenticated identity may do, defined through roles and entitlements that follow least privilege. Each service must enforce its own authorisation at its boundary, not assume the caller has checked. I also cover token lifetime, secret handling, audit logging of sensitive actions, and how access is granted and removed. Centralising identity but distributing enforcement is a common pattern.

What this question explores

Whether you separate identity from permission and enforce both at every trust boundary.

Common mistakes

  • Treating authentication and authorisation as one thing, so logging in grants everything.
  • Enforcing permissions only in the user interface and trusting every internal call.

Practise a follow-up

  • Where would you validate a token in a system with several services?
  • How would you manage service-to-service credentials?
Practise this question →
Situation · Senior

9. Your proposed solution depends on a third-party API with intermittent outages. How do you design for it?

Answer guide

I would assume the API will be slow or down and design for it. Every call gets a timeout, retries are bounded and use backoff so we do not amplify the outage, and a circuit breaker stops calls to a failing service. For work that can wait, a queue holds requests until recovery, and operations are made idempotent so retries are safe. The user experience degrades gracefully, for example with cached data or a clear message. I add monitoring and alerts for the dependency's error rate, agree expectations with the provider, and rehearse the failure with fault injection.

What this question explores

Whether you design for dependency failure with layered resilience and observable recovery.

Common mistakes

  • Calling the third party with no timeout and retrying endlessly.
  • Treating the outage as the vendor's problem and having no fallback for users.

Practise a follow-up

  • How do you choose between queueing a request and failing fast?
  • What would you monitor to detect a degrading dependency early?
Practise this question →
Technical · Senior

10. Why should a design review include run cost and operations?

Answer guide

A design that is cheap to build can be expensive to live with, so the review should price the whole life of the service. Include hosting or licences, support staffing, monitoring, security patching, upgrades, training and the eventual cost of retiring it, not only build effort. For example, a bespoke integration may save a licence fee but need a specialist on call for years. I would ask each option for a three to five year cost range and its operating model, then show the sponsor which assumptions move the total most. Where run costs are unknown, I would treat that option as riskier rather than cheaper.

What this question explores

Whether you judge a design by lifetime ownership and supportability instead of by build cost and feature fit alone.

Common mistakes

  • Weak candidates treat the review as a feature check and leave support and licence costs to the operations team later.
  • They quote a single build estimate and never mention retirement, upgrades or the people needed to run the service.

Practise a follow-up

  • How would you estimate run cost when the supplier will not share its pricing model?
  • Who should own the run-cost budget once the project team disbands?
Practise this question →
Technical · Mid-level

11. What makes an architecture decision record useful?

Answer guide

A decision record earns its keep when someone joining two years later can see why a choice was made, not just what it was. Each record should hold the context and constraints, the options considered, the decision and who took it, the consequences accepted, and the conditions that would make us revisit it. Keep it short, dated and stored beside the code or design in a place that survives team changes. For example, a note saying we chose a managed queue because the team was small and volume was low, and to revisit if traffic grew tenfold, stops the same debate being repeated.

What this question explores

Whether you see decision records as a way to preserve reasoning and trade-offs for future teams, not as paperwork.

Common mistakes

  • They describe only the final decision and leave out the options rejected and the reasons, so the record cannot help later.
  • They store records in personal folders or slide decks that vanish when people move on.

Practise a follow-up

  • How do you stop decision records becoming a bureaucratic burden for delivery teams?
  • What would you do when a record is found to be based on a wrong assumption?
Practise this question →
Situation · Senior

12. A design meets features but the support team has no monitoring or runbook. Would you approve it?

Answer guide

I would not give a clean approval, because a feature-complete design that nobody can see or recover is a future outage. Operational acceptance is part of done, so I would ask for evidence: alerts tied to real symptoms, usable logs, a named owner, a support model, a tested recovery path and a view of capacity. Then I would decide between holding the release and granting a bounded exception. An exception needs an owner, an expiry date and a plan with a date, agreed by the business and support head, so the risk is accepted knowingly. If the service is low-impact, a lighter runbook may be enough to proceed.

What this question explores

Whether you hold operational readiness as a real release criterion and handle exceptions with ownership and an expiry date.

Common mistakes

  • They approve on functional completeness and promise that monitoring and runbooks will be added after go-live.
  • They refuse outright with no route to a time-boxed exception, which makes the assurance role look obstructive.

Practise a follow-up

  • Who should have the authority to grant an operational readiness exception?
  • What is the minimum runbook you would accept for a low-risk internal service?
Practise this question →

Make the examples yours

Choose one answer containing an example or practical sequence. Explain what you would actually do, what you would check and when you would ask for help. Keep claims about your experience honest.

For technical or regulated work, check current documentation and applicable local requirements alongside this practice material.

Continue in the full Library

Make it your language

Language settings are saved on this device only.

Core interface translations are available. Some extended guidance and legal text remain in English.

Public guides remain in English where a translation is unavailable.

Voice availability depends on your browser and device. You can always type instead.

Open Library in your language