Back to docs
Research Methods

Alarm Flooding: Why Your Research Alerts Stopped Meaning Anything (2026)

Two industries independently discovered that an alerting system has a fixed capacity measured in signals per person per hour. This guide translates the EEMUA 191 and ANSI/ISA-18.2 alarm management benchmarks into research operations, and shows how to rationalize an automated insight feed so the signals that survive still carry meaning.

An alert nobody acts on is not a harmless alert. It is a tax on every other alert in the system. Process control engineers learned this after a refinery explosion, hospital safety teams learned it independently after a run of patient deaths, and both arrived at the same uncomfortable rule: a human operator has a fixed attention budget, and every signal you add spends part of it. This guide takes the two alarm management standards that came out of those investigations, EEMUA 191 and ANSI/ISA-18.2, and translates their benchmarks into research operations, where the alarms are automated theme flags, sentiment drops, score movements and data quality warnings.

The short answer

An alerting system has a capacity limit measured in signals per person per hour, not per day. EEMUA 191 puts a healthy steady state at no more than one alarm every ten minutes per operator, which is six per hour. ANSI/ISA-18.2 defines a flood as ten or more annunciated alarms in any ten minute period per operator. Past that threshold the operator stops responding to individual alarms and starts responding to the queue, which means the next genuine signal gets handled no better than the noise around it.

The fix is not a better inbox, a smarter sort order, or an extra priority tier. It is alarm rationalization: deciding, signal by signal, whether anything in your process actually changes when it fires, and removing the ones where nothing does.

The number two industries found independently

Milford Haven: 275 alarms, 11 minutes, two operators

On 24 July 1994 the Texaco refinery at Milford Haven suffered an explosion that injured twenty six people and caused roughly 48 million pounds of damage. The control room record is the part worth studying. In the last eleven minutes before the explosion, the two operators on duty had to recognise, acknowledge and act on 275 alarms. Most of the 2040 alarms configured on the system were displayed as high priority, despite many of them being informative only.

The Health and Safety Executive summary of the incident records the consequence in a single flat sentence: Excessive number of alarms in emergency situation reduced effectiveness of operator response.

Do the arithmetic, because the arithmetic is the lesson. Eleven minutes is 660 seconds. Divided across 275 alarms, that is about 2.4 seconds per alarm, and that figure assumes an operator does nothing else at all. Split between two operators, each one faced roughly 12.5 alarms per minute, or 750 per hour. The EEMUA 191 healthy benchmark is six per hour. The control room was running at about 125 times the rate a human being is expected to absorb.

Hospitals: 85 to 99 percent of alarms need no action

Nearly twenty years later and in a completely unrelated industry, the Joint Commission published Sentinel Event Alert 50 on medical device alarm safety. It reported that between 85% and 99% of alarm signals do not require clinical intervention, and that 98 alarm-related events had been reported over a little more than three years, 80 of which resulted in patient death.

Nobody in the hospital world was reading refinery reports. They arrived at the same finding from their own data.

What the two have in common

Neither failure was caused by a missing alarm. In both cases the alarm that mattered was present, configured, and displayed. It simply arrived inside a stream of signals that had already exhausted the operator. This is the part that transfers directly to research: the theme your automated analysis correctly surfaced does not help anyone if it arrived as item 340 in a feed that nobody finished reading.

What counts as an alarm

The response test

ISA-18.2 defines an alarm as an audible or visible means of indicating to the operator an equipment malfunction, process deviation, or abnormal condition requiring a response. The 2016 revision added the word timely. That definition carries the entire discipline inside it. If no operator response exists, the signal is not an alarm. It is information, and information belongs on a dashboard that people consult when they choose to, not in a channel that interrupts them.

Apply the test honestly to your own research alerting and most of it fails. A weekly digest saying sentiment moved from 4.1 to 4.0 has no defined response. Nobody has agreed what anyone would do differently. It is not an alarm and it should not be delivered like one.

Why priority is a budget, not a label

The instinctive fix when a feed gets noisy is to add a priority field. Milford Haven had priorities. Most of its 2040 alarms were set to high.

Priority only works as a rationed quantity. If any configured signal is allowed to be high priority, then high priority is just the default value of a dropdown, and it carries no information. A workable rule borrowed from alarm management practice is to fix the distribution in advance, for example roughly 80% low, 15% medium, 5% high, and to treat a new high priority alert as something that must displace an existing one rather than join it.

The signals a research pipeline can legitimately alarm on

An alarm needs a defined limit to cross. This is exactly why the structure of your instrument determines what you can alert on at all. In Koji, a study built from structured questions produces signals with defined boundaries: a scale question has a numeric threshold, a single_choice or yes_no question has a proportion you can bound, a ranking question has an order you can watch for reversals. Free text does not come with limits, so an automatically extracted theme count drifting from 12 mentions to 15 has no natural alarm condition and should not generate one.

The benchmark rates, translated

Steady state: six signals per person per hour

EEMUA 191 states that in steady state operations there should be an average of no more than one alarm every ten minutes. For a research team that means a single owner of an insight feed should face roughly six items per hour that each demand a decision. Across a six hour working day that is about 36 decisions. Most automated flagging configurations produce several hundred, and the gap between 36 and several hundred is not a productivity problem to be solved with better filtering habits. It is a design defect in the alerting system.

After a shock: fewer than ten in the first ten minutes

EEMUA 191 also states that following a major plant upset there should be fewer than ten alarms displayed in the first ten minutes. This is the harder benchmark and the more useful one, because it governs exactly the moment when the alerting system is under load and most likely to fail.

The research equivalent is a bad launch, an outage, or a pricing change. That is when your feedback channels spike, when automated analysis produces the most flags, and when the team most needs the three signals that matter rather than the three hundred that correlate with them. If your alerting has no suppression behaviour for these moments, it is guaranteed to be least useful precisely when it is most needed.

Converting the benchmark into an alerting budget

Write the budget down as a number before you configure anything. One owner, six actionable signals per hour, 36 per day, with a hard ceiling of ten in any ten minute window. Then count what your current setup actually emits over a normal week. The comparison is usually decisive, and it reframes the conversation from which alerts are annoying to we are over capacity by a factor of eight and something has to be deleted.

The inversion: more alerts, fewer caught problems

The intuitive model is that alerting coverage accumulates. Add a rule for every failure you have seen, and you catch strictly more than you did before. Each rule can only help.

Past the flood threshold the sign flips. Adding an alert reduces the total number of problems the system catches, because the marginal alert lowers the response rate applied to every other alert including the ones that matter. A feed that fires on everything is informationally equivalent to a feed that fires on nothing, except that it costs more to run and it creates a documented record of vigilance that makes the failure harder to see afterwards.

Why every single alarm is defensible

Here is the property that makes alarm flooding so persistent, and it is worth stating plainly because it defeats the usual review process. Every individual alarm in a flood is defensible. Each one was added for a real reason, fires on a genuine condition, and has someone who will argue for keeping it. The defect is not located in any single component. It is a property of the set.

This is why reviewing alerts one at a time never fixes a flood. Asked should we keep this alert, the honest answer is almost always yes. Asked which six alerts per hour are we keeping, you get a different and far more useful conversation, because the budget forces the comparison that individual review never does.

Alarm rationalization for research signals

Rationalization is the formal step in ISA-18.2 where each proposed alarm is examined before it is allowed to exist. Four questions do most of the work.

Is there a response?

Name the specific action a person takes when this fires. If the answer is we would look into it, there is no response and this is not an alarm. Move it to a dashboard.

Is the response different from other alarms?

If two signals produce the same action, they are one alarm. Duplicate paths to an identical response are the most common source of volume in an automated feed, because each new data source tends to arrive with its own copy of the same rules.

Is there time to act?

If the condition resolves on its own, or becomes irreversible, before a human can realistically respond, the alarm is decoration. A flag raised on a study that closed yesterday cannot change anything about that study.

What happens if nobody acts?

This sets priority, and because priority is rationed, this question is a comparison rather than a rating. The alarm is high priority only if it is more consequential than whatever currently holds the last high priority slot.

Suppression, deadbands and shelving

Alarm management also supplies three mechanisms that translate cleanly. A deadband stops a signal that hovers at its threshold from firing repeatedly, which is the fix for a metric that crosses back and forth across a boundary. Suppression silences downstream alarms whose cause is a known upstream condition, so one outage produces one alarm rather than forty. Shelving lets an operator deliberately silence a known signal for a bounded period, with an automatic return, which is the honest version of the mute button people improvise anyway.

How Koji helps

Koji is built so that the signals worth alerting on are defined when the study is designed rather than discovered afterwards in a pile of free text. Because Koji studies are built from six structured question types, open_ended, scale, single_choice, multiple_choice, ranking and yes_no, the closed types arrive with the limits an alarm requires. You decide in advance that a scale question dropping below a stated value, or a yes_no proportion crossing a bound, is the condition that warrants interrupting someone.

The open_ended responses are still analysed. Koji runs thematic analysis across every transcript automatically, and AI-moderated interviews, including voice interviews, probe for the reasoning behind a response rather than stopping at the rating. The difference is where that output goes. Themes and quotes populate real-time reporting that a researcher consults deliberately, while the alerting channel is reserved for the small set of conditions with an agreed response. That separation is the whole of alarm management, and it is the reason a Koji insight feed stays readable as study volume grows rather than degrading into a flood.

Customizable AI consultants let you encode the rationalization rules themselves, so the consultant that reviews incoming results applies your stated thresholds rather than flagging anything that looks interesting. Traditional survey platforms like SurveyMonkey will happily email you on every response; the useful question is not how fast a tool can notify you, but whether it can be told when not to.

Common mistakes

Treating volume as thoroughness. A feed that produces 400 flags a week is usually evidence that nobody defined a response, not evidence of unusual diligence.

Adding a priority tier instead of removing alarms. This is the Milford Haven failure exactly, and it makes the display more complicated without reducing load.

Alerting on things with no threshold. Free text theme drift has no natural limit to cross, so any alarm built on it fires on ordinary variation.

Reviewing alerts individually. Every alarm survives individual review. Only a fixed budget forces the comparisons that shrink a flood.

Measuring the feed per day rather than per hour. Daily totals conceal the bursts, and bursts are where the floods live.

Frequently asked questions

What is alarm flooding?

Alarm flooding is a condition where alarms arrive faster than an operator can process them. ANSI/ISA-18.2 sets the threshold at ten or more annunciated alarms in any ten minute period per operator. The defining feature is that the operator switches from handling individual alarms to handling the queue, which means genuine signals receive no more attention than noise.

How many research alerts per person is too many?

Use the EEMUA 191 steady state benchmark as your starting budget: no more than one alarm every ten minutes per person, or six per hour, which is about 36 in a working day. Anything beyond that is over capacity, and the excess should be deleted or moved to a dashboard rather than filtered by the recipient.

Is alarm flooding the same thing as feedback prioritization?

No, and confusing them is why the problem persists. Prioritization and triage decide what to do with signals that have already been raised. Alarm rationalization decides whether a signal is permitted to be raised at all. A team can have excellent triage and still be flooded, because triage runs downstream of the defect.

Why does adding more alerts reduce the number of problems caught?

Because each additional alert lowers the response rate applied to every alert in the system. Below the flood threshold, coverage accumulates. Above it, the marginal alert degrades attention to all the others, so the total number of genuine problems detected falls even though the number of rules rose.

What does alarm rationalization mean in a research context?

It means examining each proposed automated signal before enabling it and asking four questions: is there a defined response, is that response different from other signals, is there time to act, and what happens if nobody acts. Signals that fail any of the first three are information rather than alarms and belong on a dashboard.

Can AI analysis make alarm flooding worse?

Yes, and this is the common failure in 2026. Automated analysis lowers the cost of generating a flag to nearly zero while leaving the cost of evaluating one unchanged. Unless the alerting budget is enforced explicitly, more capable analysis produces a larger flood rather than better detection.

Related Resources

Related Articles

AI Failure Mode Analysis: An FMEA Framework for AI Products (2026)

How to run Failure Mode and Effects Analysis (FMEA) on an AI product: the failure mode taxonomy, how to score severity, occurrence and detection when failures are probabilistic, and how user research supplies the numbers.

Common Cause vs Special Cause: When a Move in Your Research Metric Is Real

Most movement in a research metric is noise, and reacting to it makes the metric worse. How to build a process behaviour chart for NPS, satisfaction or completion rate, and the decision rule that tells you when to investigate.

How to Prioritize Customer Feedback: A Framework for Product Teams

A complete guide to triaging, scoring, and acting on customer feedback. Compare RICE, MoSCoW, Kano, and the Opportunity Solution Tree — and learn how AI-native research turns raw feedback into prioritized opportunities in minutes.

The Ironies of Automation: Why a Human Reviewer Cannot Catch Your AI's Analysis Errors (2026)

Adding a human to spot-check AI coding is the reflex fix. Bainbridge showed in 1983 why it backfires, and the arithmetic is worse than teams expect.

Product Feedback Triage: A Framework for Turning Noise Into a Prioritized Backlog

A practical framework for triaging product feedback at scale — capture, dedupe, tag, route, and validate every request before it ever reaches prioritization. Includes a triage workflow, a severity matrix, and an AI-native approach.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.