The encyclopedia · Software & IT · Technical decision · 2013–2020
IBM predicted which servers would fail from their own incident tickets
IBM's PASIR classified incident tickets to find the outages that mattered, yielding about $1 billion a year in availability gains.
IBM
the move
A large IT operation produces a flood of incident tickets, and teams lose the signal that matters among the routine noise, so they fix what is loud instead of what is consequential.
IBM's Predictive Analytics for Server Incident Reduction reads the description and resolution of each ticket and classifies it, identifying the incidents that indicate real server outages and performance lags on a large population of machines.
Deployed to hundreds of IT environments since 2013, PASIR classified about 850,000 servers and drove action on the problematic ones, worth roughly $1 billion a year in increased availability and fewer incidents. It was a 2020 Franz Edelman Award finalist.
why it works
- Classifying tickets turns a noisy feed into a ranked list of what to fix first.
- Focusing on high-impact incidents spends scarce engineering time where it prevents the most downtime.
- The model learns from resolution text, so the classification improves as real fixes are recorded.
- Raising availability in one environment scales to hundreds, making the payoff compound.
what transfers
In a noisy stream of alarms, sorting the few serious signals from the many rest is a higher-leverage move than adding more monitoring.
what came after
PASIR was deployed across more than 360 IT environments and became a core IBM service-delivery analytics tool. The project was recognized as a 2020 Edelman finalist, and the method described in an INFORMS journal paper influenced how enterprises triage operational data.
references
spotted an error? The archive wants to know.