ITIL and incident management interview questions and answers
Prepare to explain service restoration, coordinated communication and the investigation that follows an incident. This set uses IT service management scenarios; actual priority models and emergency-change authority depend on the organisation’s operating model.
Impact and urgency: assess business consequences instead of relying on the loudest caller.
Restoration discipline: organise ownership, communication and safe interventions while keeping an incident record.
Learning and control: connect workarounds, problem investigation and corrective changes without delaying restoration.
How to approach your answer
Establish affected services, user impact, urgency and any immediate containment needs.
Describe command, technical ownership, update cadence and authorised recovery options. Preserve the facts as they change.
Validate service restoration with user-facing checks, then link follow-up investigation and corrective actions to accountable owners.
A revenue-critical service fails during deployment.
Illustrative approach: I would confirm the impact and activate the relevant incident process, assigning coordination and communication responsibilities. The technical team would evaluate safe recovery options, including rollback if it is feasible, while changes remain controlled. Stakeholders would receive a factual update and the next update time, not an invented restoration estimate. I would verify recovery from the user’s perspective before closing the incident. The cause investigation and preventive work would continue through linked problem and change records.
Mistakes to avoid
Holding back incident logging until the cause is understood.
Treating a workaround as permanent resolution.
Letting several teams make uncoordinated production changes.
Questions and answer guidance
Start with the level closest to your experience. Each question links to its exact practice exercise; the answer is also available here without opening the app.
Foundations
Start with the concepts and explain them using a small example.
Technical · Fresher
1. Why are incident and problem records distinct?
Read the answer guide
They have different aims, so they need different records. An incident is an interruption or reduced quality of a service, and its aim is to restore normal operation quickly, even with a workaround. A problem is the underlying cause of one or more incidents, and its aim is to find and remove that cause so it does not recur. If we merge them, teams either patch endlessly or slow down restoration to investigate. I link incidents to the problem so that known errors, workarounds and the eventual fix stay traceable.
What the interviewer is assessing
Whether you can explain that incidents restore service while problems remove causes, and why the records stay linked.
Common mistakes
They say a problem is just a serious incident, which mixes up severity with root cause work.
They investigate root cause while service is down instead of restoring first.
Practise a follow-up
When would you raise a problem record and not simply close the incident?
What is a known error and how does it help incident handling?
2. Is an employee requesting software access an incident?
Read the answer guide
Usually it is a service request, not an incident. An incident means something that should work has failed, whereas a request is a standard ask for something new, like access or software, that the service is designed to fulfil. It follows an approval and fulfilment path with its own targets. There is an exception: if a user is entitled to access and cannot get it because of a fault, such as a broken group membership, that is an incident. Getting the classification right sends work to the correct team and keeps incident statistics honest.
What the interviewer is assessing
Whether you can classify a simple case correctly and explain the exception when a fault blocks entitled access.
Common mistakes
They log every user contact as an incident, which inflates incident numbers and hides real failures.
They ignore approval steps and treat access as instantly granted.
Practise a follow-up
What would change your answer if the user was entitled but the access failed?
Why does correct classification matter for reporting?
Show how you would apply the idea to a constraint, disagreement or failure.
Technical · Mid-level
3. How should incident priority be determined?
Read the answer guide
Priority should reflect consequence, so I use two dimensions: business impact, meaning how many people or how much money or safety is affected, and urgency, meaning how fast the harm grows or the deadline arrives. A matrix with clear definitions gives consistent results, and each level maps to a response target. Seniority alone does not raise priority, though a request from an executive may point to a real impact worth checking. Priority is also revisited, for instance upgraded when the outage spreads or downgraded once a workaround exists.
What the interviewer is assessing
Whether you set incident priority from impact and urgency using explicit criteria that can be revised, not from who is asking.
Common mistakes
They give the highest priority to whoever shouts loudest or holds the most senior job title.
They set priority once at logging and never revise it as the situation changes.
Practise a follow-up
How would you handle a senior executive who insists a low-impact issue is critical?
How often should incident priority definitions be reviewed and by whom?
4. How do you distinguish incident restoration from permanent problem resolution?
Read the answer guide
Restoring service and fixing the cause are different jobs. Incident management aims to bring the service back as fast as possible, using a restart, a rollback or a workaround, even if the cause is unknown. Once users are working again, the incident closes. Problem management then looks for the real cause using root cause analysis. It ends with a permanent fix or a known error record, and actions to stop a repeat, such as a change, a monitor or a process update. Track each action to completion.
5. A workaround restores service during warranty but the cause remains. What next?
Read the answer guide
A workaround restores service, but the problem stays. I would record the incident as resolved by workaround, then create or link a problem record to find the root cause and create a known error entry describing symptoms and the workaround, so support can act quickly next time. The permanent fix needs an owner, a target date and a risk assessment if it slips. I would also make the limits of the workaround clear to users, such as extra manual steps or reduced capacity. Warranty terms decide who pays for the fix.
What the interviewer is assessing
Whether you follow a workaround with problem management, a known error record, an owner and clear communication of its limits.
Common mistakes
They close the incident and forget the cause, so the same failure returns and support rediscovers the workaround.
They keep the workaround permanently without an owner or date for the real fix.
Practise a follow-up
What should a known error record contain to help a support analyst?
How do you decide the priority of a problem record when the workaround is stable?
6. How should incident restoration, underlying investigation and corrective changes remain connected?
Read the answer guide
I would link the related records, the evidence and the accountable decisions, without forcing every activity to share a single status. The service is restored first, while investigation and controlled correction continue as needed. Clear relationships prevent the restored incident from being mistaken for completion of all preventive work. For example, the incident can close once users are working, while the linked problem record stays open until the cause is fixed through a reviewed change. The limit is tool support, so even if the tool is weak I at least reference record numbers in each.
What the interviewer is assessing
Whether you can keep incident, problem and change connected through links and evidence without merging their lifecycles.
Common mistakes
They close the problem record when the incident closes, so the underlying cause is never investigated.
They keep one combined record for all activities, so a restored incident stays open for weeks.
Practise a follow-up
When would you open a problem record rather than just closing the incident?
How would you make sure a corrective change is traced back to the original incident?
7. A team delays logging incidents until diagnosis is complete to protect response metrics. How would you address it?
Read the answer guide
I would first validate the recording behaviour, by comparing logging times with other evidence such as call records or alerts, and look at the effect on evidence and customer experience. Then I agree meaningful measurement and accountability that starts from the intended event, namely when the user reported the problem. I improve incentives and controls rather than accepting altered timestamps as a genuine service improvement. The trade-off is a likely defensive reaction from the team, so I approach it as fixing a flawed measure and pressure, not as catching individuals, while still being clear that the practice stops.
What the interviewer is assessing
Whether you can address target gaming by checking the evidence, restoring honest measurement from the true start point and fixing the incentives.
Common mistakes
They accept the improved response figures as genuine because the team now meets its targets.
They punish the individuals involved without examining the target design that encouraged the behaviour.
Practise a follow-up
How would you verify when incidents were actually reported?
How would you redesign the target so it does not reward late logging?
8. An urgent incident requires a production change, but the usual approver is unavailable. What should the operating model provide?
Read the answer guide
The operating model should provide an agreed emergency decision and escalation route, such as a named deputy or an emergency approval group, with the evidence and accountability needed. Actions are recorded and a review follows afterwards. Urgency needs a workable authorised path rather than either uncontrolled changes or indefinite waiting for one named individual. For example, a deputy approver on call, with a short checklist covering impact and rollback, would let the change proceed safely. The trade-off is that emergency routes can be abused, so I report their use and review each case after the event.
What the interviewer is assessing
Whether you can describe an authorised emergency route with deputies, evidence and post-event review, so urgency neither bypasses control nor stalls.
Common mistakes
They make the production change anyway without approval because the incident is urgent, leaving no record.
They wait for the usual approver for hours because the process names only one person.
Practise a follow-up
What evidence would you require before approving an emergency change?
How would you review emergency changes after the event to detect misuse?
9. How would you assess recurring incident risk without being misled by ticket closure counts?
Read the answer guide
Closed tickets measure activity, not stability. I would look at how often the same failure recurs, which services it affects, how long restoration takes, and whether the underlying cause was ever fixed or only the symptom cleared. I check reopened tickets and duplicates, since a high closure count can come from a team repeatedly clearing the same fault. For example, five hundred closures might reduce to a handful of recurring causes. I would then ask whether problem management exists to remove root causes. Persistent exposure can sit quietly behind healthy-looking closure figures.
What the interviewer is assessing
Whether you judge recurring incident risk through recurrence, restoration time and root causes, and check reopenings and duplicates, rather than counting closed tickets.
Common mistakes
They treat a high ticket closure count as evidence that the service is well run and stable.
They count tickets by number only and never group them by cause, so recurring failures are invisible.
Practise a follow-up
How would you find the root causes behind a large set of recurring tickets?
What evidence would convince you that problem management is actually working?
10. Provider and customer report different monthly incident counts. What would you reconcile?
Read the answer guide
I would compare the source records, the time boundaries, the classification, the duplicates, the linked cases and the scope, to find out why the populations differ. For example, one side may count by creation date and the other by resolution date, or one may count linked incidents separately. Then I agree the governing definition and correct the evidence trail. I would not negotiate a middle number without understanding the difference, because that leaves the real cause unfixed. The trade-off is time spent, but a clear definition prevents the same dispute in later months.
What the interviewer is assessing
Whether you reconcile a reporting mismatch by tracing the differences in population and definition before agreeing a number.
Common mistakes
They split the difference and agree a middle number without finding why the counts differ.
They insist their own count is right without comparing the source records with the other party.
Practise a follow-up
Which differences in definition most commonly cause such mismatches?
How would you record the agreed definition so it applies to future reports?
Explain trade-offs, wider consequences and the evidence behind your decision.
Scenario · Senior
11. A revenue-critical service fails during a deployment. What do you do in the first 30 minutes?
Read the answer guide
In the first thirty minutes the goal is restoration and clear control. I would declare the correct severity so the right people are engaged, appoint an incident commander who coordinates while specialists work, and decide quickly whether to roll back the deployment or use a workaround. Communications go out early to leaders and affected users with impact and the time of the next update, even if the cause is unknown. A scribe records the timeline. Root cause analysis is left for the review after service returns, so investigation does not delay restoration.
What the interviewer is assessing
Whether you prioritise restoration, command structure and communication in a major incident, and keep a timeline for the review.
Common mistakes
They start hunting for root cause while the revenue service stays down instead of rolling back.
They leave communication until the fix is done, so leaders and customers hear nothing for an hour.
Practise a follow-up
What would you write in the first stakeholder update when the cause is unknown?
When would you decide to roll back instead of trying to fix forward?
12. Change, incident and project teams use different service names. How do you fix reporting?
Read the answer guide
Different names for the same service make integrated reporting impossible, so I would fix the root and not just the report. First comes a shared service taxonomy with unique identifiers, agreed with service owners and kept in one place. Then I map historical records from each process to those identifiers, accepting manual clean-up for important services. I assign a data steward and make the identifier mandatory in new tickets, changes and project records. Adoption is phased, starting with critical services, and I show value early, such as a combined incident and change view.
What the interviewer is assessing
Whether you solve inconsistent service naming with a shared taxonomy, identifiers, stewardship and phased adoption.
Common mistakes
They build a mapping spreadsheet for the report and leave the underlying records inconsistent.
They demand all teams change names overnight with no steward, mapping or phased approach.
Practise a follow-up
Who should own the service taxonomy and how do they handle new services?
How would you handle historical records that cannot be mapped confidently?