Skip to main content

Which Clinical Workflows Should AI Touch First?

9 September 2026 | David and Goliath

Quick answer

Sequence by reversibility, not volume. Start where the tool records rather than interprets, where a mistake is visible before it reaches a patient, and where checking the output is realistic for the person responsible for it. Volume is the wrong first filter because it puts the least tested tool closest to the most decisions.

  • Sort by whether the output could change a clinical decision unchecked
  • Recording work sits on the unregulated side of the TGA line, interpreting work does not
  • A mistake that is visible teaches you how the tool behaves
  • Volume first puts the least tested tool nearest the most decisions

Mentioned: Therapeutic Goods Administration, Ahpra, ARTG

Every practice that deploys AI well starts somewhere small and unglamorous. Every practice that deploys it badly starts with the thing that would save the most time. The difference is a sorting method, and it takes about an hour to apply to a whole practice.

What is the right first filter for a workflow?

Whether the output could change a clinical decision if nobody checked it. That question sorts a practice's work faster than any category scheme, because it asks about consequence rather than about task type.

Rostering, billing and correspondence fail that test harmlessly. A tool producing a differential diagnosis passes it immediately.

Everything else falls somewhere between, and where it falls tells you the order to deploy in.

Why is volume the wrong first filter?

Because it puts the least tested tool closest to the most decisions. The highest volume task is attractive precisely because it is where the hours are, which is also where an unnoticed failure repeats fastest.

Volume is the right filter for the second deployment, once you know how the tool behaves under your conditions.

Starting there means learning the tool's failure modes at scale rather than at leisure.

Does the regulator's line help with sequencing?

It does, and it is the clearest one available. The Therapeutic Goods Administration drew it on 30 January 2026: a digital scribe which only transcribes and translates is not a medical device, while one that analyses or interprets, for example by generating a diagnosis, differential diagnosis or treatment recommendation the practitioner did not state, is a medical device and must be in the ARTG (Source: TGA, digital scribes guidance, 30 January 2026).

That line is not a safety ranking, but it correlates usefully with one. Work where the tool records sits on the simpler side of it in both regulatory and clinical terms.

Sequencing along it means the early deployments are also the ones with the lighter compliance burden, which is a rare alignment.

What makes a mistake useful rather than dangerous?

Visibility before it reaches the patient. A summary a clinician reads before signing is a mistake you catch. A message sent automatically is a mistake you learn about later, from someone else.

Early deployments should be the ones where being wrong is recoverable and obvious, because that is how a practice builds an accurate sense of what the tool is actually good at.

Confidence built on invisible outputs is not confidence, it is absence of feedback.

Is checking the output realistic for the person responsible?

This is the test most sequencing plans skip. Ahpra holds the registered practitioner ultimately responsible for any AI used in their practice and says they cannot defer to an output without applying professional judgement (Source: Ahpra, Meeting your professional obligations when using Artificial Intelligence in healthcare).

That obligation is only meetable if checking is practical. A tool producing more output than anyone can realistically review has converted a responsibility into a formality.

So the question for each candidate workflow is not only whether a human is in the loop, but whether that human has the minutes.

What does a workable first deployment look like?

One workflow, one team, a defined period, and a named practitioner who reads the output. Small enough that the reading is real.

Documentation and summarising are the usual answer, because they keep a practitioner between the tool and the patient and because the failure mode is a bad draft rather than a bad decision.

Correspondence drafting works too, provided nothing sends without a human pressing send.

What should wait?

Anything producing a conclusion the practitioner did not reach, anything patient-facing without review, and anything where the output is too voluminous to check. Those are not permanent exclusions, they are second-phase work.

Triage and symptom assessment sit here for most practices. They are attractive and they are the ones where an unnoticed error reaches a patient directly.

Waiting is not caution for its own sake. It is sequencing so that the tool has a track record before it sits near a decision that matters.

How do you know when to move to the next workflow?

When you can describe how the tool fails. Not whether it fails, which is a given, but the shape of its errors: what it tends to get wrong, under what conditions, and what a wrong output looks like.

A practice that cannot answer that has not learned anything from the first deployment, whatever the time savings say.

That description is the actual output of a first phase, more than the hours saved.

Where does sequencing go wrong?

The most common failure is deploying by department rather than by workflow. A whole clinic gets a tool at once, and the variation in how each practitioner uses it means nobody can tell what the tool did.

The second is treating a pilot as exempt. Obligations attach the moment a tool touches a real patient, so a pilot is a deployment with a shorter name.

The third is sequencing on vendor recommendation. A supplier's suggested starting point optimises for demonstrable value, which is usually the highest volume workflow, which is the one to avoid first.

If you want the triage applied to your own workflow list rather than a generic framework, the AI Triage Assessment is written by a person within 48 hours. The governance it sits inside is in AI governance for Australian healthcare providers.

Sources: TGA, digital scribes guidance, 30 January 2026. Ahpra, Meeting your professional obligations when using Artificial Intelligence in healthcare.

Ready to move from reading to shipping?

Ten business days. Four modules. One agent live by the end.