EN
Back to the archive

The encyclopedia · Software & IT · Technical decision · 2013–2020

IBM predicted which servers would fail from their own incident tickets

IBM's PASIR classified incident tickets to find the outages that mattered, yielding about $1 billion a year in availability gains.

IBM

the move

A large IT operation produces a flood of incident tickets, and teams lose the signal that matters among the routine noise, so they fix what is loud instead of what is consequential.

IBM's Predictive Analytics for Server Incident Reduction reads the description and resolution of each ticket and classifies it, identifying the incidents that indicate real server outages and performance lags on a large population of machines.

Deployed to hundreds of IT environments since 2013, PASIR classified about 850,000 servers and drove action on the problematic ones, worth roughly $1 billion a year in increased availability and fewer incidents. It was a 2020 Franz Edelman Award finalist.

why it works

  • Classifying tickets turns a noisy feed into a ranked list of what to fix first.
  • Focusing on high-impact incidents spends scarce engineering time where it prevents the most downtime.
  • The model learns from resolution text, so the classification improves as real fixes are recorded.
  • Raising availability in one environment scales to hundreds, making the payoff compound.
the payoffFind the high-impact incidents before they become outagesneat

what transfers

In a noisy stream of alarms, sorting the few serious signals from the many rest is a higher-leverage move than adding more monitoring.

what came after

PASIR was deployed across more than 360 IT environments and became a core IBM service-delivery analytics tool. The project was recognized as a 2020 Edelman finalist, and the method described in an INFORMS journal paper influenced how enterprises triage operational data.

references

spotted an error? The archive wants to know.

same kind of clever