Courses Job Ready Program Fresher Trainings AI For Class 7 to 12 Corporate Training Placements Tutorials
Free Learning Resources

IT Tutorials & Interview Prep

Free guides, interview Q&As, and job responsibility breakdowns — curated by industry veterans to help you crack MNC interviews

237+
Tutorial Articles
16
Topic Categories
100%
Free to Read
← Back to ITIL & Service Desk Essentials

Problem Management

ITIL & Service Desk Essentials Last Updated: Sep 05, 2026

1. Definition of Problem Management

Definition

Problem Management is the ITSM (IT Service Management) practice responsible for reducing the likelihood and impact of incidents by identifying their actual and potential causes, and managing workarounds and known errors. It focuses on finding the ROOT CAUSE of one or more incidents so that the underlying issue is permanently fixed instead of the same problem happening again and again.

 

💡 Easy Hinglish Explanation

Agar ek hi problem baar-baar aa rahi hai (jaise roz email server down hona), toh Incident Management sirf usse turant theek karta hai temporarily. Lekin Problem Management yeh dekhta hai ki AAKHIR yeh baar baar kyun ho raha hai, aur uski jad (root cause) ko hamesha ke liye khatam karta hai.

 

Key Idea in One Line

Incident Management = “Fix it now” (symptom ka ilaj)   |   Problem Management = “Fix it forever” (root cause ka ilaj).

 

2. What is Problem Management?

Problem Management is a structured, proactive and reactive ITIL practice that deals with 'Problems' — where a Problem is defined as the underlying cause of one or more Incidents. It works closely with Incident Management, but its goal is completely different: instead of restoring service quickly, it aims to prevent incidents from happening in the first place, or at least reduce their impact when they do happen.

Core Objectives of Problem Management

  • Prevent problems and resulting incidents from happening.
  • Eliminate recurring incidents (same issue happening again and again).
  • Minimize the impact of incidents that cannot be prevented.
  • Maintain information about problems, workarounds, and known errors (in a Known Error Database, called KEDB).
  • Support the organization's continual service improvement.

 

📌 Day-to-Day Example

Suppose in a company, the office Wi-Fi disconnects every day at 3 PM. Every day, the IT helpdesk logs it as a new incident and restarts the router — this is Incident Management (quick fix). Problem Management steps in to ask WHY this happens daily, finds that the router overheats due to a faulty cooling fan, and gets it replaced permanently — so the daily disconnection never happens again.

 

3. Why is Problem Management Needed?

Without Problem Management, organizations keep firefighting the same issues repeatedly, wasting time, money and customer trust. Here is why it matters:

Business & Technical Reasons

  • Reduces the NUMBER of incidents by fixing root causes permanently.
  • Reduces DOWNTIME and improves overall service availability.
  • Saves cost — repeated incidents mean repeated effort, manpower and resources.
  • Improves customer/user satisfaction because the same complaint doesn't return.
  • Builds a Known Error Database (KEDB) — useful knowledge for faster future diagnosis.
  • Helps in better decision-making using data and trend analysis (proactive planning).

 

💡 Easy Hinglish Explanation

Agar company sirf Incident Management pe hi depend kare, toh IT team hamesha aag bujhati rahegi (firefighting) — lekin aag lagti kyun hai, yeh kabhi pata nahi chalega. Problem Management se company ek baar mein permanent solution nikaalti hai, jisse future mein time aur paisa dono bachta hai.

 

4. How Does Problem Management Work? (Process Flow)

Problem Management follows a structured, step-by-step process/workflow. Each step ensures the problem is properly identified, diagnosed, documented, and permanently resolved.

Fig 1: End-to-end Problem Management Process Flow

Step-by-Step Process Explained

Step 1: Problem Identification

The process begins when a problem is identified — either from repeated/major incidents, proactive monitoring/trend analysis, supplier notifications, or analysis of incident data. The goal here is simply to recognize that an underlying issue exists.

📌 Day-to-Day Example

Example: The Service Desk notices that 'VPN connection failed' has been logged 15 times in one week across different users — this pattern indicates a Problem.

 

Step 2: Problem Logging

Once identified, the problem is formally logged into the Problem Management tool/system with a unique Problem Record. Details captured include date, description, related incidents, affected services, and impact level.

📌 Day-to-Day Example

Example: A Problem Record 'PRB0002451 - Repeated VPN Failures' is created and linked to all 15 related incident tickets.

 

Step 3: Categorization & Prioritization

The problem is categorized (e.g., Network, Hardware, Software, Application) and prioritized based on urgency and impact (similar to incident prioritization), so that the most critical problems are worked on first.

📌 Day-to-Day Example

Example: VPN failure affecting 200+ remote employees is marked 'High Priority' compared to a minor UI glitch affecting 2 users.

 

Step 4: Investigation & Diagnosis

The Problem Management team investigates the issue using root cause analysis (RCA) techniques such as the '5 Whys', Fishbone (Ishikawa) diagram, Pareto analysis, or Kepner-Tregoe method to find the actual root cause.

📌 Day-to-Day Example

Example: Using the '5 Whys' technique, the team finds the VPN failures are caused by an outdated firmware version on the VPN gateway server.

 

Step 5: Raise a Known Error (KEDB Entry)

Once the root cause is identified (even if a permanent fix is not yet available), it is documented as a Known Error in the Known Error Database (KEDB), along with any available workaround, so the Service Desk can use it for faster future resolution.

📌 Day-to-Day Example

Example: Known Error 'KE00312 – VPN Gateway Firmware Bug' is logged with the workaround: 'Restart VPN gateway service until patch is applied.

 

Step 6: Resolution / Workaround Implementation

A permanent fix (or a temporary workaround, if permanent fix takes time) is planned and implemented, typically through a Change Request if it affects live systems, following Change Management process.

📌 Day-to-Day Example

Example: A Change Request is raised to upgrade the VPN gateway firmware during the next scheduled maintenance window.

 

Step 7: Problem Closure & Review

After the fix is confirmed successful and no further related incidents occur, the Problem Record is formally closed. A Post Implementation Review / lessons-learned review may also be done for major problems.

📌 Day-to-Day Example

Example: After firmware upgrade, no VPN failure incidents are reported for 30 days — the Problem Record PRB0002451 is closed.

 

⚠ Important Point

Problem Management does NOT close a problem just because a workaround exists. A problem should ideally be closed only after the ROOT CAUSE is permanently eliminated. If only a workaround exists, it stays as a 'Known Error' — still open, but manageable.

 

5. When is Problem Management Triggered?

Problem Management can start at different times depending on whether it is Reactive or Proactive:

  • When the same incident occurs repeatedly (recurring incidents).
  • When a single Major Incident happens (high impact, even if it happened only once).
  • When monitoring tools detect unusual trends/patterns before any incident occurs (proactive).
  • When a vendor/supplier reports a known defect in hardware/software being used.
  • During periodic trend analysis of the incident database by the Problem Manager.

 

6. Where is Problem Management Applied?

Problem Management is applied wherever IT services are delivered and incidents can recur. It is a core practice inside the ITSM framework and is commonly implemented using ITSM tools.

  • IT Service Desks and IT Operations teams (most common).
  • Data centers and Infrastructure teams (servers, networks, storage).
  • Software Development / Application Support teams (recurring application bugs).
  • Managed Service Providers (MSPs) handling client IT environments.
  • ITSM tools such as ServiceNow, Jira Service Management, BMC Remedy, Freshservice, etc., where Problem Records are logged and tracked.

 

7. Who is Involved in Problem Management?

RoleResponsibility
Problem ManagerOwns the overall Problem Management process; prioritizes problems, tracks progress, and ensures KEDB is maintained properly.
Service Desk / Incident ManagerIdentifies recurring incidents and raises them as potential problems; provides incident data for analysis.
Technical / Subject Matter Experts (SMEs)Perform root cause analysis and technical investigation; recommend permanent fixes.
Change ManagerEnsures that any permanent fix affecting live systems goes through a controlled Change Management process.
Service OwnerAccountable for the overall health of the service; approves priority and resourcing for major problems.
End Users / CustomersReport recurring issues; sometimes indirectly trigger a problem investigation through repeated complaints.

 

 

8. Important Concepts & Technical Terms

Problem

A Problem is the underlying cause of one or more Incidents. It may be identified after incidents have already occurred (reactive) or discovered proactively through monitoring and trend analysis, even before any incident happens.

💡 Easy Hinglish Explanation

Problem ek 'root cause' hai jo baar baar incidents create karta hai. Jab tak isko fix nahi karoge, incidents wapas aate rahenge.

 

📌 Day-to-Day Example

Slow database queries causing multiple application timeout incidents across different users.

 

Known Error

A Known Error is a problem that has been successfully analyzed and diagnosed, meaning its root cause is documented, but a permanent fix has not necessarily been implemented yet. A workaround may exist to reduce impact until the fix is applied.

💡 Easy Hinglish Explanation

Known Error matlab — humein pata hai problem kya hai aur kyun ho raha hai, bas permanent solution abhi lagaya nahi gaya, lekin temporary workaround available hai.

 

📌 Day-to-Day Example

A software bug causing report export failure is documented with the workaround 'export in CSV format instead of PDF until patch released.

 

Known Error Database (KEDB)

KEDB is a centralized repository/database that stores details of all known errors, including their symptoms, root cause, workaround and resolution status. It helps Service Desk agents quickly match new incidents to existing known errors for faster resolution.

💡 Easy Hinglish Explanation

KEDB ek 'reference book' jaisa hai jisme saare known problems aur unke workaround likhe hote hain, taaki agli baar wahi issue aaye toh turant solution mil jaye.

 

📌 Day-to-Day Example

A new agent sees a ticket about 'login page blank screen', searches KEDB, finds a matching known error, and applies the documented workaround immediately instead of re-investigating.

 

Workaround

A workaround is a temporary method of reducing or eliminating the impact of an incident or problem for which a full, permanent resolution is not yet available. It does not fix the root cause — it only manages the symptoms.

💡 Easy Hinglish Explanation

Workaround matlab jugaad — asli problem fix nahi hui, lekin user ko turant kaam chalu rakhne ka temporary tarika mil gaya.

 

📌 Day-to-Day Example

Restarting a server every night automatically to avoid memory leak crashes, until a permanent code fix is released.                                                          

 

Root Cause Analysis (RCA)

RCA is a systematic process/technique used to identify the fundamental, underlying cause of a problem rather than just addressing its symptoms. Common RCA techniques include the '5 Whys', Fishbone (Ishikawa) Diagram, and Pareto Analysis.

💡 Easy Hinglish Explanation

RCA ek technique hai jisse hum 'asli wajah' tak pahunchte hain, sirf upar upar se symptom treat nahi karte.

 

📌 Day-to-Day Example

Asking '5 Whys' for a website crash: Why did it crash? → Server ran out of memory → Why? → A memory leak in the code → Why? → An unclosed database connection → (root cause found).

 

Major Problem Review (MPR) / Post Implementation Review (PIR)

A formal review conducted after a major problem is resolved, to evaluate what went well, what could be improved, and to capture lessons learned for future prevention.

💡 Easy Hinglish Explanation

Bade problem ke baad team baithti hai aur discuss karti hai ki kya sahi hua, kya galat hua, aur agli baar aisa dobara na ho iske liye kya karna chahiye.

 

📌 Day-to-Day Example

After a 6-hour banking application outage, the team holds an MPR meeting to document lessons learned and update monitoring alerts.

 

Trend Analysis

The process of analyzing historical incident and problem data over time to identify recurring patterns, frequently affected components, or emerging risks — forming the foundation of Proactive Problem Management.

💡 Easy Hinglish Explanation

Trend analysis matlab purane data ko dekh kar pattern dhoondna, jisse pata chale ki koi issue baar baar toh nahi ho raha.

 

📌 Day-to-Day Example

IT team notices that 60% of laptop hardware incidents happen on devices older than 4 years, and plans proactive replacements.

 

Change Request (CR)

A formal request submitted to implement a permanent fix identified through Problem Management. It goes through the Change Management process for approval, scheduling, and risk assessment before implementation on live systems.

💡 Easy Hinglish Explanation

Jab problem ka permanent solution live system mein lagana ho, toh usse directly nahi kiya jata — pehle ek Change Request banti hai jisko approve karwaya jata hai.

 

📌 Day-to-Day Example

A Change Request is raised to patch a security vulnerability identified as the root cause of repeated server crashes.

 

 

9. Difference / Comparison Tables

9.1 Incident Management vs Problem Management

Fig 2: Relationship between Incident, Problem and Known Error

BasisIncident ManagementProblem Management
GoalRestore normal service as fast as possibleFind and eliminate the root cause permanently
NatureReactive (mostly)Both Reactive and Proactive
FocusSymptom / immediate impactUnderlying cause
Time SensitivityUrgent — speed matters mostThorough — accuracy matters more than speed
OutputService restored, incident closedRoot cause found, Known Error / permanent fix documented
ExampleRestart the crashed server to bring service backInvestigate why the server keeps crashing and fix it permanently

 

9.2 Problem vs Known Error vs Workaround

TermRoot Cause Known?Permanent Fix Applied?Status
ProblemNot yet knownNoOpen — under investigation
Known ErrorYes, documentedNot necessarilyOpen, but manageable via workaround
WorkaroundMay or may not be knownNo (temporary only)Active until permanent fix applied
Resolved ProblemYesYesClosed

 

9.3 Reactive vs Proactive Problem Management

Fig 3: Reactive vs Proactive approach comparison

BasisReactive Problem ManagementProactive Problem Management
TriggerStarts after incidents already occurredStarts before any incident occurs, via monitoring/trend analysis
ApproachResponding to existing pain pointsPreventing future pain points
Data SourceIncident records, repeated ticketsMonitoring tools, performance trends, capacity data
ExampleInvestigating why the payment gateway failed 5 times this weekNoticing disk usage trending upward and fixing it before it causes an outage

 

 

10. Scenario-Based Questions

These practical scenarios test your understanding of Problem Management concepts in real situations.

Q1. Users from three different departments report that the file server becomes extremely slow every day between 2 PM and 3 PM. The Service Desk has logged 20 such incidents this month. What should the team do?

Answer: This should be escalated to Problem Management, since it is a recurring pattern of incidents pointing to a common underlying cause, not a one-time issue.

Why/Reason: A single incident is handled by Incident Management, but when the SAME issue repeats multiple times (20 times here), it clearly indicates an underlying Problem that needs root cause analysis — possibly a scheduled backup job or heavy report generation running during that time window.

Q2. A critical banking application goes down for 4 hours, affecting thousands of customers. It has never happened before. Should this be treated as a Problem even though it occurred only once?

Answer: Yes. This qualifies as a Major Incident, and Major Incidents are typically followed by a mandatory Problem investigation (Major Problem Review), even if it happened only once.

Why/Reason: Problem Management is not only triggered by repetition — high-impact/major incidents are investigated immediately regardless of frequency, because the business impact and risk of recurrence is too high to ignore.

Q3. The Problem Management team has identified the root cause of an application crash — a memory leak in the code — but the development team says the permanent code fix will take 3 weeks. What should be done in the meantime?

Answer: A workaround should be implemented, such as scheduling automatic application/server restarts every night, and the issue should be documented as a Known Error in the KEDB with the workaround details.

Why/Reason: When a root cause is known but a permanent fix is not immediately possible, the problem becomes a Known Error. Workarounds reduce business impact while the permanent solution is developed.

Q4. A Service Desk agent receives a new incident and, while searching the Known Error Database, finds an exact match with a documented workaround. What should the agent do?

Answer: The agent should immediately apply the documented workaround from the KEDB to resolve the incident quickly, without needing to re-investigate the issue from scratch.

Why/Reason: This is exactly the purpose of maintaining a KEDB — to speed up incident resolution by reusing previously diagnosed root causes and workarounds, saving time and effort.

Q5. The IT monitoring dashboard shows that a particular server's CPU usage has been steadily increasing over the last 3 weeks, though no incident has occurred yet. Should Problem Management be involved at this stage?

Answer: Yes. This is a case for Proactive Problem Management — the trend should be investigated now, before it leads to an actual outage.

Why/Reason: Proactive Problem Management works on early warning signs from monitoring and trend analysis, aiming to prevent incidents before they even occur, rather than waiting for something to break.

Q6. After fixing the root cause of a recurring network outage problem, the Problem Manager wants to close the Problem Record. What must be verified first?

Answer: The team must confirm that the permanent fix has been successfully implemented (usually via a completed Change Request) and that no related incidents have recurred for a sufficient monitoring period before formally closing the Problem Record.

Why/Reason: Closing a problem prematurely, before confirming the fix works in the live environment, risks reopening the same problem later and reduces confidence in the Problem Management process.

Q7. A problem's root cause requires changing a critical piece of production code. Can the Problem Management team directly implement this fix on their own?

Answer: No. The fix must go through the formal Change Management process — a Change Request must be raised, reviewed, approved (often by a Change Advisory Board), and then scheduled for implementation.

Why/Reason: Problem Management identifies WHAT needs to be fixed and recommends the solution, but changes to live production systems must be controlled through Change Management to avoid uncontrolled risk and unintended side effects.

Q8. Two unrelated incidents — a printer failure and an email delay — are reported on the same day. Should these automatically be linked into a single Problem?

Answer: No. Unless investigation shows they share a common root cause, they should remain separate. Problems are only linked to incidents that stem from the same underlying issue.

Why/Reason: Incorrectly grouping unrelated incidents into one problem can mislead root cause analysis and waste investigation effort on the wrong direction.

 

11. Interview Questions

11.1 Basic Interview Questions

Q: What is Problem Management?

A: Problem Management is the ITSM practice that identifies the root cause of one or more incidents and works to prevent them from recurring, by managing problems, known errors, and workarounds.

Q: What is the difference between an Incident and a Problem?

A: An Incident is an unplanned interruption to a service, while a Problem is the underlying cause of one or more incidents. Incident Management restores service quickly; Problem Management finds and fixes the root cause.

Q: What is a Known Error?

A: A Known Error is a problem whose root cause has been identified and documented, with or without an available permanent fix, usually along with a workaround.

Q: What is KEDB?

A: KEDB stands for Known Error Database — a repository that stores information about known errors, their root causes, and workarounds, used to speed up future incident resolution.

Q: What is the difference between Reactive and Proactive Problem Management?

A: Reactive Problem Management is triggered after incidents occur, while Proactive Problem Management identifies potential problems before any incident happens, using monitoring and trend analysis.

Q: Name common Root Cause Analysis (RCA) techniques.

A: The '5 Whys' technique, Fishbone (Ishikawa) Diagram, Pareto Analysis, and Kepner-Tregoe Problem Analysis are commonly used RCA techniques.

11.2 Practical / Scenario-Based Interview Questions

Q: If the same incident occurs 5 times in a week, what would you do as a Problem Manager?

A: I would raise a Problem Record, link all 5 related incidents to it, prioritize it based on business impact, and start root cause analysis using techniques like the 5 Whys to find and fix the underlying cause.

Q: How do you decide the priority of a problem?

A: Priority is generally decided using Impact (how many users/services are affected, and how severely) combined with Urgency (how time-critical the fix is) — similar to incident prioritization, often shown as a priority matrix.

Q: Can a Problem be closed without a permanent fix?

A: Generally no — a problem should ideally stay open (or remain as a Known Error) until a permanent fix is verified. However, in some organizations, a problem may be closed if the business formally accepts the risk and workaround as a long-term solution.

Q: How does Problem Management support Continual Service Improvement (CSI)?

A: By performing trend analysis, major problem reviews, and root cause fixes, Problem Management generates lessons-learned data that feeds directly into Continual Service Improvement initiatives, helping the organization avoid repeat failures.

Q: What KPIs / metrics are used to measure Problem Management effectiveness?

A: Common metrics include: number of problems raised vs resolved, percentage of incidents linked to known problems, average time to diagnose a root cause, number of recurring incidents reduced, and number of proactive problems identified before incidents occurred.

Q: What would you do if two teams disagree on the root cause of a problem?

A: I would facilitate a structured RCA session (e.g., using a Fishbone diagram) with both teams, review supporting data/logs objectively, and if needed, escalate to the Problem Manager or a senior technical authority for a final data-backed decision.