Problem Management is ServiceNow's solution for finding and fixing the underlying causes of recurring IT issues. When your help desk keeps getting tickets about the same problems over and over, Problem Management helps you dig deeper to understand why these incidents keep happening and document permanent fixes or workarounds. IT problem managers, senior analysts, and infrastructure teams use it to break the cycle of firefighting the same issues repeatedly. Once Problem Management is running in ServiceNow, your organization moves from reactive incident response to proactive problem prevention. The system connects related incidents, tracks investigation progress, stores solutions in a searchable Known Error Database, and helps teams share knowledge about fixes and workarounds. This means fewer repeat incidents, faster resolution times when problems do occur, and a growing library of institutional knowledge that doesn't walk out the door when experienced staff leave.
Key Capabilities
Link Related Incidents to Problems
ServiceNow automatically suggests when multiple incidents might stem from the same underlying problem, or analysts can manually link them. This gives you a clear view of how widespread an issue really is and prevents teams from solving the same root cause multiple times in different places.
Track Investigation Progress
Problem records maintain a detailed timeline of investigation steps, findings, and assigned team members. Everyone involved can see what has been tried, what worked, what didn't, and what the next steps should be without having to dig through email chains or chat histories.
Document Workarounds and Permanent Fixes
Teams can record temporary workarounds to keep business running while working on permanent solutions. Once implemented, the system tracks whether the fix actually prevented recurrence, so you know if your solution really worked.
Build Known Error Database Articles
Proven solutions get stored as searchable knowledge articles that help desk agents and engineers can find quickly. This means the next time a similar issue appears, your team has documented steps to resolve it instead of starting from scratch.
Generate Problem Reports
ServiceNow creates reports showing which problems cause the most incidents, how long investigations take, and which areas of your infrastructure need attention. This helps managers prioritize where to invest time and resources for maximum impact.
Prevent Future Incidents
By tracking problems through to permanent resolution and measuring their success, the system helps reduce the total volume of incidents your team handles. Problems that get properly solved stop generating new tickets.
How It Works
Problem Management typically starts when someone notices a pattern of related incidents or when a major incident needs deeper investigation to prevent recurrence. A problem manager or senior analyst opens a problem record, links it to the relevant incidents, and assigns investigation tasks to appropriate technical teams. ServiceNow tracks all the research, testing, and solution attempts in one place, making it easy for multiple people to collaborate on complex issues. Once teams find a root cause, they document workarounds or permanent fixes, test the solution, and create knowledge articles so the next person who encounters this issue can resolve it quickly.
Who Uses It and How
Regional healthcare system
Their patient portal kept going down every few weeks, generating dozens of incident tickets each time. Problem Management helped them track all portal-related incidents back to a database connection pool that was too small for peak usage periods.
Result: Portal outages dropped from twice monthly to once every six months after they implemented the permanent fix.
Global manufacturing company
Factory floor systems at multiple plants were experiencing random slowdowns that couldn't be reproduced during maintenance windows. Problem investigators discovered that a recent network equipment firmware update was causing intermittent packet loss under heavy load.
Result: Rolling back the firmware and working with the vendor on a proper fix eliminated production disruptions across 12 manufacturing sites.
University IT department
Students couldn't access online course materials during peak study periods before exams. The problem investigation revealed that the learning management system wasn't scaling properly when thousands of students tried to download large files simultaneously.
Result: Implementing a content delivery network and load balancing eliminated the access issues and reduced help desk calls by 60% during finals week.
Financial services firm
Trading applications were experiencing brief but costly delays several times per day. Problem Management helped correlate these incidents with specific database maintenance jobs that were running during business hours and competing for system resources.
Result: Rescheduling maintenance tasks to off-hours eliminated trading delays and prevented an estimated $2 million in potential losses.
Sourdough: ServiceNow Monitoring and Analytics
A Chrome extension for ServiceNow Admins and Developers with essential tools, analytics, graphs and monitoring features.
Free to install. Pro $5/month after a 14-day no-card trial.
Pro requires the ServiceNow admin role. Upgrade inside the extension.
Implementation: What to Know
Getting Problem Management right requires buy-in from both your help desk team and senior technical staff, since problem investigation often needs deep system knowledge that frontline agents don't have. Plan for a 3-6 month rollout that starts with training problem managers and establishing clear escalation criteria for when incidents should become problems. You'll need your incident data to be reasonably clean and categorized first, since problem identification often depends on spotting patterns in incident history. Most implementations stall because organizations try to turn every incident into a problem instead of focusing on the recurring or high-impact issues that really need root cause analysis.
Common Use Cases
Investigate why email servers crash every Monday morning
The IT manager notices a pattern of email outages at the start of each work week and opens a problem record to investigate. Problem Management tracks the investigation as engineers discover that automated weekend backup jobs are consuming too much memory, leaving insufficient resources when users start logging in Monday morning.
Document a workaround for a vendor software bug
A critical business application has a known bug that the vendor won't fix for six months, but the support team found a configuration workaround. They create a problem record and Known Error Database article so that when this issue affects other users, agents can apply the workaround immediately instead of rediscovering the solution.
Track resolution of network performance complaints
Multiple departments are reporting slow network speeds, creating dozens of individual incident tickets. Problem Management links these incidents together and assigns network engineers to investigate, ultimately discovering that a misconfigured switch was creating bottlenecks affecting three office buildings.
Analyze recurring printer connectivity issues
Several different printer models keep losing network connections, and help desk agents are spending hours per week reconnecting them. Problem investigation reveals that the network team's recent security policy changes are causing the printers to lose their IP addresses, leading to a permanent configuration fix.
Prevent future database corruption incidents
After a major database corruption incident that took down customer-facing applications, the problem manager opens an investigation to understand the root cause. The analysis discovers inadequate disk space monitoring and backup validation procedures, leading to improved monitoring and backup processes that prevent future corruption.
Key Tables
Best Practices
- ✓Don't turn every incident into a problem - focus on recurring issues or ones that had major business impact, otherwise you'll overwhelm your investigation teams
- ✓Link problems to incidents as you discover connections, even if you didn't spot the pattern initially - this helps with accurate impact measurement and prevents duplicate investigations
- ✓Write Known Error Database articles like you're explaining the fix to someone who wasn't involved in the original investigation - include enough context and steps that future engineers can actually use them
- ✓Set realistic timeframes for problem resolution based on complexity, not business pressure - rushing investigations often leads to band-aid solutions that don't prevent recurrence
- ✓Train your incident management team to recognize patterns that should trigger problem investigations - they're usually the first to notice when the same type of issue keeps coming up
- ✓Review resolved problems quarterly to make sure your permanent fixes actually worked - some solutions that seemed successful initially still allow similar problems to develop later
Common Pitfalls
Creating too many problem records and overwhelming investigation teams with trivial issues
Establish clear criteria for what qualifies as a problem, typically recurring incidents or single incidents with major business impact. Not every incident needs a problem investigation.
Closing problems without creating Knowledge Database articles or workaround documentation
Make knowledge article creation a required step before closing problem records. Your future self will thank you when the same issue appears six months later.
Assigning problem investigations to people who don't have the technical depth or authority to implement solutions
Problem resolution often requires senior engineers or architects who can understand complex system interactions and get approval for infrastructure changes. Plan your assignments accordingly.
Never measuring whether your problem solutions actually prevented recurrence of incidents
Set up reports to track incident volumes by category after problems are resolved. If you're still getting the same types of incidents, your solutions might not be addressing the real root causes.
Keeping problems open indefinitely because permanent solutions require budget or vendor cooperation
Document workarounds and interim solutions, then close the problem with a clear plan for permanent resolution. You can always open a new problem later when resources become available.
Frequently Asked Questions
What's the difference between Problem Management and Incident Management?
Incident Management focuses on restoring service as quickly as possible when something breaks. Problem Management digs deeper to understand why it broke and prevent it from happening again. Think of incidents as treating symptoms and problems as curing the disease.
Do I need Problem Management if we already have good incident processes?
Yes, if you find yourself fixing the same types of issues repeatedly. Good incident management gets you back online quickly, but without problem management you'll keep having the same outages and user complaints. Problem Management is what moves you from reactive firefighting to proactive prevention.
Who should be assigned to investigate problems?
Problem investigation usually requires your most experienced technical staff who understand how different systems interact and have the authority to implement infrastructure changes. Junior technicians can help with data gathering, but complex root cause analysis needs senior engineers or architects.
How long should problem investigations take?
It varies widely depending on complexity, but most problems should have some progress or interim solution within two weeks. If investigations drag on for months without clear milestones, consider whether you have the right people assigned or if the scope is too broad.
What's a Known Error and when should I create one?
A Known Error is a problem where you understand the root cause and have a documented workaround or solution, but haven't implemented a permanent fix yet. Create Known Error Database articles when you have proven steps that others can follow to resolve or work around the issue.
Should every major incident automatically become a problem?
Not necessarily. Major incidents should get a post-incident review, but only become formal problems if there's a risk of recurrence or if the incident revealed systemic issues that need deeper investigation. One-off incidents caused by human error or external factors might not need problem management.
How do I know if my problem solutions are actually working?
Track incident volumes and types after implementing problem solutions. If you resolved a network connectivity problem but are still getting network tickets at the same rate, your solution probably didn't address the real root cause. Good problem management reduces future incident volume.
Related Modules
Test Your Knowledge
Quick 3-question quiz — see how your ServiceNow skills stack up.
A list view on a table with millions of records is slow. Best fix?
Select an answer to continue