Safety9 min read

Your Safety Management System Is Measuring How Hard It Is to File a Report

Every safety management system produces a number that looks like a measure of safety and is not. Report volume is a measure of the distance between noticing something and saying so, and of what the person who noticed expects will happen next. Read as anything else, it misleads in the most comfortable direction available.
On this page

A quiet SMS is the most ambiguous signal you own

Reports are down this quarter. In most operations that slide gets a good reception. It reads as improvement, it fits the story the safety manager wants to tell the accountable executive, and nobody in the room has an obvious reason to argue.

It is the single most ambiguous data point a safety management system produces. A fall in occurrence reporting means one of two things. Either fewer hazards are occurring, or fewer people are willing to mention the ones that are. The number is identical in both cases. The required response is the opposite in each.

If hazards genuinely fell, the correct action is to understand which change produced it and protect that change. If willingness fell, the correct action is to find out what happened to the last few people who spoke up. Acting on the wrong reading is not merely wasted effort. An operation that celebrates a quiet system while its reporting culture erodes has removed its own early warning and congratulated itself for the silence.

The wider point is that report volume is a behavioural measurement, not an environmental one. It tells you about the people, the form and the aftermath. It tells you very little about the hazards.

Friction is the variable you are actually measuring

Between noticing something and filing a report sits a cost, and that cost is paid by a person who is usually tired, usually finishing a duty, and who gains nothing personally from paying it. Everything that raises the cost lowers the volume, and none of it has any relationship to how safe the operation is.

Length of the form

Every mandatory field that does not change the outcome is a small argument for not bothering. Forms grow by accretion and are almost never cut back.

Where it can be filed

A report that can only be made at a desk is a report made hours later, if at all, with the detail already fading. Distance in time is distance in accuracy.

Filed once or twice

When the same event has to be entered into a safety system and again into a maintenance or operational record, one of the two entries quietly becomes optional.

Categorise before describing

Asking a reporter to pick a hazard class and a severity before they can say what happened makes them do the safety department's job to earn the right to speak.

Why classifying before describing deters the reporter

The fourth of those deserves emphasis. Classification is analysis, and analysis is what the safety function exists to perform. Requiring it up front does two kinds of damage: it deters the reporter, and it contaminates the data with categories chosen by someone who was guessing. Description first, classification after, is not a nicety. It is the difference between a hazard report and a form.

Why double entry makes one of the two records optional

Double entry is worth its own attention because it rarely announces itself. The event exists in the safety record and in the operational record, and neither knows about the other — the same fragmentation problem that quietly taxes every other part of an operation, arriving in the one place where the tax is paid in silence.

The second report is a function of the first

Nobody files a report in a vacuum. They file it against a memory of what happened the last time, and against what they saw happen to colleagues who reported before them.

So the decisive variable in safety reporting is not the form. It is what visibly came back. Acknowledgement within a shift rather than a month. A named person who owns it. An outcome the reporter can see, even when that outcome is a considered decision to accept the risk and do nothing. A closing note that reaches the crew who raised it, and ideally the crews who did not.

When nothing comes back, the reporter draws an entirely accurate conclusion: reporting is unpaid administrative work that consumes their time and produces no observable change. That is not apathy or a poor safety culture in the abstract. It is a correct inference from the available evidence, and it will not be reversed by a poster, a stand-up briefing or a reminder in the monthly bulletin.

An operation that wants more reports has exactly one durable lever, which is to make reporting visibly worth doing. Everything else is exhortation.

Just culture is an operational property, not a poster

Every operator claims a just culture. The claim is cheap. What it costs is a set of concrete commitments about how a report is handled once it is in the building.

The questions that decide whether a report is safe to file

The substance is in a few unglamorous questions. Who reads the report first, and does that person hold line authority over the reporter. Whether the reporter's identity travels with the report by default or only when the reporter chooses. Whether a report can be used in a performance conversation. What the organisation does when the event is embarrassing, expensive, or involves someone senior — which is the only test that ever really counts.

Where the line between error and violation gets drawn

The distinction between error and violation is where most just culture policies fail in practice. The principle is well understood: honest error is examined for what produced it, while deliberate and reckless disregard of known rules is addressed. The difficulty is that the line has to be drawn identically every time. Drawn consistently, it is a rule people can operate under and plan around. Drawn according to who was involved or how badly the day ended, it becomes a lottery, and crews respond to lotteries by not entering.

Consistency here is worth more than sophistication. A blunt line applied the same way for everyone produces more honest reporting than a nuanced framework applied at the discretion of whoever is in the chair that week.

The aviation risk matrix and the pull of the amber middle

The severity-by-likelihood grid is one of the most useful communication devices in aviation risk management. It gives a mixed group of people a shared vocabulary for how bad and how often, and it makes an assessment legible to someone who was not in the room.

Why a matrix cell is not a measurement

The trouble starts when the communication device is mistaken for an arithmetic one. A matrix cell is an ordinal judgement dressed as a coordinate. Multiplying two subjective rankings produces a number that looks precise and carries no more information than the two guesses that made it. Operations then compare those numbers across events, rank them, average them by quarter, and build reports on a scale that was never metric.

What matters far more than precision is consistency between assessors. If two experienced people assess the same event and land in different cells, the matrix is measuring the assessor rather than the risk, and no amount of granularity fixes that. Calibration — a handful of worked examples, periodically re-run, that everyone assesses independently and then compares — buys more than a finer grid ever will.

How risk matrices drift to amber

The second failure is gravitational. Matrices pull towards the amber middle. The top-right cells demand action nobody has budget for, the bottom-left cells look dismissive of a colleague's concern, and the middle is defensible in both directions. Enough events assessed as tolerable-with-monitoring and the register becomes a place where things are recorded rather than decided. If almost nothing in your risk assessment history has ever been rated in a way that forced an operational change, the matrix is documenting the operation, not governing it.

Leading indicators describe a future you can still change

ICAO Annex 19 frames safety management around four pillars: safety policy and objectives, safety risk management, safety assurance, and safety promotion. Most operations build the first and the second competently, treat the fourth as a communications exercise, and quietly reduce the third to counting occurrences.

Occurrences are lagging indicators. They are easy to collect precisely because they announce themselves, and every one of them describes a past that is already fixed. An SMS built only on occurrence counts is a very well-documented rear-view mirror.

What a leading indicator looks like in practice

Leading safety performance indicators are the harder and more valuable half of safety assurance. They track conditions that tend to precede events rather than the events themselves: how often the schedule is being flown at the edge of duty limits, how long deferred defects sit open, how thin recency and currency margins have become across the roster, how much training is slipping to the right, how often a plan changed after the crew had already briefed it. None of these are safety events. All of them describe the conditions under which safety events become more likely.

They are also harder to agree on, harder to defend in a meeting, and impossible to collect if operational data lives somewhere the safety function cannot reach. That last constraint is why most programmes stop at lagging indicators, and it is a data problem long before it is a safety-management one.

A single report is often genuinely unremarkable. A crew found the ramp lighting poor at a particular stand. Somebody was handed a load sheet late. A frequency was congested at a handover point. Read alone, each is a note. Assessed alone, each lands in the amber middle and is monitored.

The ninth report of the same thing is a different object entirely. It is a pattern, and patterns are what justify spending money and changing procedures. The entire value of a safety database is that it can turn nine unremarkable notes into one argument.

Why free-text reports hide the pattern

This is where free-text-only records quietly fail. Nine people describe the same underlying condition in nine different vocabularies, filed under whatever category seemed closest, at nine different stations, across two quarters and possibly two safety managers. Nothing in the record links them. The pattern exists in the operation and not in the data, and it stays invisible until it produces an event large enough to trigger an investigation that finds all nine.

Making the ninth report findable is mostly about structure that a human did not have to supply under time pressure: consistent location, phase of flight, aircraft, and a taxonomy applied after the fact by someone whose job it is. It is also about being able to look at the record afterwards rather than only live — the same discipline that turns a movement board into a planning instrument turns a report queue into a trend.

Corrective actions are the part that decays

The most common failure in a mature SMS is not that hazards go unreported or unassessed. It is what happens after the assessment.

The pattern is consistent across operations of every size. An action is raised. An owner is assigned, often in their absence. A due date is set with more optimism than evidence. The date passes. It is extended once at a safety meeting, and after that it is not discussed again, because discussing it would require someone to say either that it is done or that it is not going to be. The register grows. Nobody reads it. At audit, it is reconciled in a fortnight of unpleasant work, and the cycle restarts.

Why corrective actions decay rather than close

Corrective action tracking decays for structural reasons rather than lazy ones. The action was assigned to a department rather than a person. It was never sized, so it competes badly against work that was. Nothing surfaces it between meetings. And crucially, nothing distinguishes an action that is genuinely progressing from one that has been open for eleven months, because the register renders both as a row with a date in the past.

This matters more than it appears, because the corrective action loop is the mechanism by which safety assurance stops being a document and becomes a practice. An operation that identifies hazards accurately, assesses them consistently, and then fails to close anything has built a very expensive observation system. It has also, without meaning to, sent the clearest possible message back to the people who reported.

Is your SMS working, or performing?

Both look similar in a binder. They diverge sharply under a small number of questions:

  • When report volume last fell, did anyone ask whether hazards had fallen or whether willingness had?
  • How long between a report being filed and a human being acknowledging it — measured, not estimated?
  • Can a reporter see what happened to their report, including when the answer was to accept the risk and act on nothing?
  • Does reporting require the reporter to classify the event before they are allowed to describe it?
  • Would the same event, reported by a junior first officer and by a senior captain, be handled identically?
  • When two assessors rate the same occurrence, how often do they land in the same cell — and has anyone ever checked?
  • How many risk assessments in the past year forced an operational change rather than continued monitoring?
  • Could you find the previous eight reports describing the same condition without knowing in advance that they existed?
  • What is the oldest open corrective action, and can anyone say what is actually blocking it?

A system that answers these well will not necessarily produce fewer reports. It will more likely produce more of them, and better ones, and a safety manager who is busier than they were — which is the correct shape for an operation that has made it easy and worthwhile to say something is wrong.

It is the assumption we approach safety reporting with in Aerotalon: the work that decides whether a safety management system functions happens at the moment somebody weighs up whether to bother, not in the register afterwards.

The quiet system is the one to worry about. Not because silence is proof of anything, but because it is the one state an SMS can reach without anyone having to make a decision, and the one state that looks, from a distance, exactly like success.

Bring us a reporting problem, not a feature list

If your report volume is falling, or your corrective action register has more open items than anyone can defend, that is a better conversation than a demo script. Bring it to Aerotalon and we will work through what the numbers are actually telling you.