Technology11 min read

The Question Is Not What AI Can Decide. It Is What You Can Sign For

Aviation is an industry organised around accountability. Somebody signs, and the signature has a name on it. That one property, more than any technical limitation, determines where artificial intelligence is genuinely useful in aviation operations and where it remains a well-produced demonstration.
On this page

AI in aviation begins with a signature

In most industries, a wrong automated decision is a cost. Somebody refunds an order, a campaign underperforms, a forecast is revised. The organisation absorbs it and the model gets retrained.

In this one, a wrong decision is a name on a document and, eventually, a regulator asking that person why. The release to service was signed. The flight was dispatched. The crew was rostered legal. Each of those is an assertion made by a human being who is accountable for it, and no reasonable reading of the rules lets that accountability transfer to a piece of software because the software was confident.

Why deliberate is not the same as slow

This is routinely mistaken for conservatism, usually by people selling something. It is not. It is the correct response to a consequence structure where the downside is not measured in revenue. An industry that has spent decades learning to be deliberate about the introduction of new failure modes is not being slow; it is being appropriately expensive to convince.

Any honest discussion of AI in flight operations has to start there, because that property does not soften as the models improve. A more capable model does not acquire a licence. It does not appear at the enquiry. The question worth asking of any capability is therefore not what it could decide, but what a named person could reasonably sign for having seen it.

A model is a function of what you feed it

The second thing to say is duller and matters more. Every machine learning system in aviation is a function of its inputs, and most operations cannot yet answer simple factual questions about their own past with confidence.

How many hours did that airframe fly last quarter, and does the answer from the schedule match the answer from the flight log and the answer from the maintenance record. How long did that turn actually take, as opposed to how long it was planned to take and how long somebody typed into a form afterwards. Which of last year's delays were the same cause under three different labels. These are not analytical questions. They are bookkeeping questions, and in a great many operations they still require somebody to open several systems and reconcile by hand.

Why prediction inherits the contradictions in your record

Prediction built on a record that disagrees with itself does not produce insight. It produces confident noise, faster and in greater volume than before, with a presentation layer that makes it look like knowledge. The model has no way to distinguish a real operational pattern from an artefact of how two departments happened to enter data differently in 2023.

So the actual precondition for anything useful is unfashionable: a coherent operational record, captured once, consistent across the functions that touch it. That is the single-source-of-truth problem, and it is worth solving on its own merits, because it pays for itself in reduced reconciliation work whether or not a model is ever pointed at the result. Aviation data quality is the precondition, not the afterthought. Operations that skip it and buy the model first usually discover the sequencing the hard way.

Attention is the scarce resource, not autonomy

Here is the reframing that makes the rest of this tractable. The scarce resource in an operations centre is not decision-making capacity. Experienced controllers, schedulers and planners are perfectly capable of making good decisions. What they do not have is the ability to look at everything at once.

A duty manager with ninety movements on the board cannot give each one considered attention. They triage, mostly by habit and experience, and the failures that hurt are usually not bad decisions but unexamined ones — the situation nobody looked at closely because ninety other things also needed looking at, and it did not appear unusual until it was.

That is where automation in flight operations earns its keep. The highest-value applications rank, filter and surface. “These four of today's ninety movements deserve a second look, and here is what makes each of them different from the other eighty-six.” The system has not decided anything. It has spent the operation's attention better, which in a day with finite attention is close to the whole game.

This is a continuation of what good operational displays already do rather than a break from them. A well-built movement board directs attention through geometry — the tight turn that looks tight without anyone doing arithmetic. Ranking is the same instinct applied where the signal is too subtle to be drawn. And it leaves the decision precisely where the accountability already sits, which is the only place it can go.

Where the value genuinely concentrates today

Strip away the demonstrations and a short list remains. These are real, they are available now, and each has a limit worth naming.

Patterns in your own history

Which turns were routinely optimistic rather than occasionally unlucky. Which stations absorb delay and which propagate it. Answerable from records you already hold.

The text nobody reads twice

Safety reports, defect narratives, handover notes, correspondence. Humans read each one once and remember almost none of it in aggregate. Machines are good at aggregate.

Anomalies against your baseline

Unusual compared with how this operation normally behaves, not compared with an industry average assembled from operations that fly nothing like yours.

First drafts, checked by a person

Summaries, briefings, the first version of a document. Generative AI is useful here because a human reviews the output before it counts for anything.

Notice what these have in common. Each takes a task that is currently done badly because it is done by tired people at volume, and each returns something a human evaluates before it has any effect. None of them is autonomous. The trade in every case is that the machine is allowed to be wrong, because being wrong is cheap when the next step is a person reading it.

Where each of these stops being useful

The limits are equally worth stating. Historical pattern-finding tells you what used to happen, which is a poor guide to an operation that has just changed fleet type or network. Text summarisation compresses, and compression loses things; the item a summary omits is invisible in a way the item a human skipped is not. Anomaly detection against a baseline flags novelty, and novelty is not the same as risk — a large share of what it surfaces will be uninteresting, which is exactly why the ranking has to be good enough that people keep looking.

Predictive maintenance, treated carefully

Predictive maintenance is the application most often cited and the one most often oversold, which is unfortunate, because in the right conditions it is the real thing.

Those conditions are specific. It works where there is dense recorded parameter data and enough similar airframes accumulating enough hours for a genuine pattern to separate from noise. Large fleets of common types generate that. The physics helps too: components that degrade gradually and leave a trace in recorded data before they fail are exactly the ones a model can learn.

Predictive maintenance in a small or mixed fleet

The picture in smaller and mixed fleets is considerably weaker, and the honest version of the pitch says so. Six aircraft of three types, with modest recorded parameter coverage, do not produce a population a model can learn a rare failure mode from. What such an operation gets instead is usually simpler and still worthwhile: better trend monitoring, better visibility of recurring defects, and a defect history that is actually searchable. That is not machine learning and does not need to be described as such.

A prediction you cannot act on in time

The point that gets skipped entirely is the operational one. A prediction changes nothing unless the maintenance programme and the schedule can act on it. If a component is flagged as likely to fail within thirty days but the approved programme has no provision for opportunistic replacement, the parts lead time is longer than the window, and the aircraft is committed for the whole month, the prediction has generated anxiety rather than value. The gap between knowing and being able to do something is where most predictive maintenance projects quietly stall, and it is a scheduling and planning problem rather than a modelling one.

Why “why” is a functional requirement

A recommendation an operator cannot interrogate will meet one of two fates, and both are expensive.

It will be ignored, because a professional who is accountable for the outcome will not act on an instruction whose basis they cannot examine, and after a few unexplained flags they stop reading them. Or it will be over-trusted, which is worse, because the person accepting it has no way to notice the case where it is wrong. Uninterrogable output does not get calibrated trust. It gets either zero or too much.

So explanation is not a nicety in this domain, and it is not primarily a compliance box. It is what makes appropriate reliance possible. “This turn is flagged because the last eleven attempts at this station in this weather window averaged twenty minutes over plan” is something a duty manager can evaluate in seconds against what they know. They can accept it, override it with a reason, or notice that the model is comparing against a period when the station was operating differently. All three are good outcomes. A bare risk score of 0.81 permits none of them.

This has a practical consequence for how AI aviation safety applications should be judged. The useful question is not how accurate the system is in aggregate. It is whether a competent operator, shown an individual output, can tell whether this particular one is right.

The failure modes aviation has already met

The risks worth planning for are not the dramatic ones. They are ordinary and they compound slowly.

  • Confident wrongness — a wrong output arrives with exactly the same fluency and formatting as a right one. There is no tell. Human error usually comes with hesitation attached; this does not.
  • Quiet degradation — the operation changes shape, new type, new network, new season, and the model keeps answering in the register of the operation it learned. Nothing breaks visibly. It simply becomes less right, and nobody is watching for that.
  • Over-trust after a good run — twenty accurate outputs in a row train the user to stop checking the twenty-first. The failure is not in the model; it is in the relationship the user has formed with it.
  • Skill fade — the person who no longer performs the task no longer maintains the judgement that let them supervise it. The oversight is nominal by the time it is needed.

What the flight deck already taught this industry

None of this is speculative for this industry. Aviation has lived through precisely this once already on the flight deck, and the human-factors tradition that emerged from it — automation complacency, mode confusion, the erosion of hand-flying currency, the difficulty of monitoring a system that is usually right — is one of the more mature bodies of practical knowledge anywhere about people supervising machines. The findings were expensive to acquire. They transfer almost directly to a controller supervising a ranking system, and an industry that has paid for this lesson once should not pay for it twice on the ground floor of the same building.

The AI you can see, and the AI you cannot

Almost all of the conversation about AI in business aviation and airline software concerns features a buyer can be shown. That is not a conspiracy; it is a consequence of how software is sold. A capability that can be put on a slide and clicked through in forty minutes is the capability that gets built and the capability that gets discussed, because both the seller and the buyer need something to point at.

The AI that never reaches the interface

A substantial share of the genuine value of artificial intelligence in operational software never reaches the interface. It sits in how the software is built, tested, monitored and supported. Review that catches a defect before it ships. Monitoring that surfaces a fault in an operator's environment before the operator notices and reports it. Support staff who arrive at the conversation already holding the relevant history rather than asking the customer to reconstruct it. None of that appears in a screenshot. All of it shows up as a product that breaks less often and gets fixed faster when it does.

This creates a genuine evaluation problem, and it runs in an unhelpful direction. The visible AI is the easiest to demonstrate and by some distance the easiest to fake — a plausible conversational panel over a thin capability is a weekend's work, and it demonstrates beautifully. The invisible AI is the harder thing to build, requires sustained engineering discipline rather than a feature decision, and cannot be shown at all. So the demonstration, as a selection mechanism, is systematically biased towards the less valuable half of the picture.

Observable proxies a buyer can actually check

The honest counterweight is that a buyer cannot audit an engineering practice directly, and any vendor can claim one. Nobody is going to review a supplier's internal process, and the claim is unfalsifiable in a sales meeting. What a buyer can do is ask about the observable proxies, which are the things that would be true if the practice existed. How quickly does a reported defect get diagnosed, as distinct from acknowledged. When something goes wrong in your environment, does the vendor tend to find it first or do you. When you raise something, does support arrive already understanding how your operation is configured, or does every ticket start from the beginning. Those answers are checkable against reference customers, and they are considerably more informative than anything on the feature list.

What good adoption looks like

The operations that get value from this do something unremarkable. They pick one task, narrow enough to describe in a sentence and painful enough that somebody complains about it. They agree in advance what better would look like and how they would know. They draw the human decision boundary explicitly, in writing, so that everyone knows which outputs are advisory and which are nothing at all. They keep it reversible. And they check, after a quarter, whether it actually helped.

The operations that get nothing tend to have started from the technology rather than the problem, bought breadth rather than depth, and set no criterion that could have been failed. There is a related question of cost discipline here: capability that nobody uses still appears on the invoice, and it is worth reading any AI line item the way you would read the rest of the total cost of ownership.

The skill that fades when the task goes away

There is one more question, and it deserves an honest answer rather than a reassuring one. What happens to the skill of the people who stop doing the task? If a scheduler no longer builds the roster from first principles, some part of the judgement that let them recognise a bad roster goes with it. That may be an acceptable trade — the industry has made it before, knowingly, in other places. It is only a problem when it is made accidentally, by an organisation that assumed the supervision would stay as good as it was on day one.

Diagnostic questions to ask about any AI capability

That leaves a short set of diagnostic questions worth putting to any AI capability in an operational product, in roughly this order:

  • What decision does this touch, and who signs for that decision today?
  • What is it actually replacing — a task nobody wants to do, or a judgement somebody is accountable for?
  • What record does it depend on, and does that record currently agree with itself?
  • Can an operator see why a given output was produced, well enough to disagree with it?
  • Who is accountable when an output that influenced a decision turns out to have been wrong, and is that written down anywhere?
  • What happens when it is wrong — who catches it, and how?
  • How would we notice it had degraded as the operation changed shape?
  • Does it reduce the number of things needing human attention, or add to them?
  • Can the supplier say clearly what they do not use this for? A vendor with no boundary has not thought about one.
  • What does the supplier do with it that we will never see — and what would we observe if that were true?
  • Could we turn it off next month without unpicking the operation?

These are practical questions rather than sceptical ones. A supplier who has thought seriously about the problem will find them easy and will probably have better answers than the questions deserve. A supplier who has not will reach for the demonstration.

We would rather be asked these questions at Aerotalon than not, including the last one.

The technology is real. A good deal of what is said about it is not. The distinction between the two, in this industry, comes down to the same thing it always has: whether the output is something a named person would put their signature next to, having understood what they were signing.

Start with the operation, not the technology

If you are trying to work out which parts of your operation would actually benefit from automation and which are being sold to you, bring the specifics — the fleet, the record you keep today, the decisions that cost you most. We are happy to have that conversation with Aerotalon whether or not it ends in software.