AI Systems

When Should AI Make a Decision—and When Should a Person Step In?

Not every decision needs a person, and not every decision can be handed to a system. Six tests, three paths, and the thresholds that keep the difference visible.

BYBO Editorial10 min read

FigureThree paths for one decision
Routine and reversible
  1. ProceedInside written limits, logged and sampled
  2. AskA named reviewer approves before anything happens
  3. EscalateA specialist decides outside the system
Serious and hard to undo
Contents
  1. In brief
  2. What are you actually deciding when you automate a decision?
  3. Six tests that decide who decides
  4. Proceed, ask or escalate: giving each decision a path
  5. How do you set a threshold you can defend?
  6. What does a reviewer need in order to decide anything?
  7. How do you know the gate is not just a rubber stamp?
  8. When should the limits change, and who may switch it off?
  9. Where this has limits
  10. Questions
  11. Sources

In brief

  • Six tests settle most cases: reversibility, consequence, evidence quality, model confidence, regulation and how directly the customer feels it.
  • Turn the tests into three paths — proceed, ask, escalate — with a written threshold for moving between them.
  • A reviewer needs evidence in one place, authority to reject and enough time. Without those three, the gate is decoration.
  • Watch approval, edit and rejection rates. A gate that approves everything is a rubber stamp, not a control.

What are you actually deciding when you automate a decision?

The question is rarely whether AI should decide. It is which slice of a decision a system may finish on its own, which slice it should prepare for someone else to approve, and which slice it must not touch at all. A refund is not one decision: reading the complaint, matching the order, judging whether the policy applies and releasing the money are four, and they do not all belong to the same owner.

The NIST AI Risk Management Framework describes the range without prescribing a point on it. Human-AI configurations, it notes, can span from fully autonomous to fully manual: a system may decide on its own, defer the decision to a human expert, or serve as an additional opinion for a person who decides. Some systems need no oversight at all; others specifically require it. The framework also declines to set a risk tolerance for you, because tolerance depends on your context, your obligations and what a mistake would cost.

So this is a business judgement, made once per decision type and written down. The six tests below are the ones worth arguing about in a room with the person who owns the process today.

Six tests that decide who decides

Six tests for a single decision type
TestThe questionSend it to a person when
ReversibilityCan we undo this, and how easily?Undoing needs an apology, a refund or a filing
ConsequenceWhat is the worst plausible outcome?One mistake harms someone or costs seriously
EvidenceIs the information complete and current?Sources are missing, stale or contradict each other
ConfidenceHow sure is the system, and on what basis?The case is unusual or sits near the line
ObligationDoes law, contract or policy require a person?A regulator, a contract or your own policy says so
Customer impactWill the customer feel this directly?The message is sensitive or makes a promise

Two of the six are misread often enough to be worth a warning. Confidence is not correctness: a model can produce a fluent, well-formatted answer with a high score attached and still be wrong, particularly on a case unlike anything it has seen. And reversibility is a property of your process rather than of the model. A journal entry you can reverse before the month closes is more reversible than the same amount paid out through a bank file, and far more reversible than an email already read by a customer.

Run the tests on a decision type, not on a whole workflow. Most workflows contain one or two decisions that need a person and a dozen steps that do not, and treating them all the same is how teams end up either approving everything by hand or approving nothing at all.

Proceed, ask or escalate: giving each decision a path

Three paths cover almost everything, and every decision type should be assigned to one of them in writing.

  • Proceed. The system acts inside written limits and records what it did. The check is sampling and monitoring after the fact, not approval before it.
  • Ask. The system prepares the action with its evidence and a named reviewer approves, edits or rejects it before anything happens.
  • Escalate. The case leaves the system for a specialist, and the outcome comes back into the record so the pattern stays visible.

The mechanics of a good gate — the reviewer, the evidence, the fallback when nobody responds — are covered in our note on how to design the human decision into the workflow. What this article adds is the allocation: which decisions belong in which path, and what moves a case between them.

Choose the default deliberately. Sending everything to the ask path feels safe and rarely is. An approval queue that nobody clears delays the work, trains reviewers to click through in batches, and hides the few cases that genuinely needed a person. Automating a decision and reviewing a sample of the results is sometimes the more careful choice.

How do you set a threshold you can defend?

A usable threshold has three parts: a business limit expressed in your own units, an evidence rule, and a named exception list. The limit is an amount, a discount, a tenure or a quantity. The evidence rule says what must be present and agree before the automatic path is available. The exception list names the cases that always go to a person, whatever the numbers say.

Start narrower than you think you need, and widen on evidence. A run of cases at the current limit that reviewers approved without a single edit is a reason to raise it. A cluster of corrections, a new supplier, a changed policy or a festival-season spike in volume is a reason to lower it again. Record each change with a date and a reason, so nobody has to reconstruct why the limit is what it is.

What does a reviewer need in order to decide anything?

A gate is only as good as the person standing at it, and reviewers need four things: the evidence in one place, a clearly stated proposed action, real authority to reject it, and time. The first two are design work. The third is a management decision. The fourth is arithmetic that teams skip.

Count the review capacity before launch. Suppose a workflow sends 120 cases a day to one person, and an honest review takes two minutes. That is four hours of somebody’s day, every day, before holidays and sick leave. Either the volume reaching the gate has to fall, or a second reviewer has to exist, or the queue will be cleared by clicking. Reviewing is work, and it belongs in the operating cost of the system.

How do you know the gate is not just a rubber stamp?

People defer to confident machines, and a well-presented recommendation can crowd out the reviewer’s own reading of the case. NIST puts the risk carefully: results from human-AI interaction vary, and under some conditions the AI part of the interaction can amplify human biases, producing more biased decisions than either the person or the system alone. The same framework observes that data on how often, and why, people overrule a system in live use is worth collecting and analysing.

So collect it. These five measures make a rubber stamp visible within a month:

  • Approval, edit and rejection rates, per reviewer and per case type.
  • Median seconds spent on a case before the decision.
  • Reversals, complaints and corrections that arrive after approval.
  • Results on planted test cases with known errors in them.
  • What reviewers wrote in the reason field when they rejected something.

Read the numbers in both directions. If almost everything is approved within a few seconds, the gate is a formality. If reviewers edit more than half of what reaches them, the system is adding work rather than removing it, and the rules, the inputs or the scope need attention before anyone widens the automatic path.

When should the limits change, and who may switch it off?

Decisions age. A threshold set in a quiet month meets the festival season. A supplier changes its invoice layout, a policy is rewritten, a model version is updated by your provider without anyone on your team touching the system. Put the decision rules on the same review cycle as the system itself, and treat a change to a threshold as a change to the system: someone approves it, the record says why, and there is a way back.

Agree in advance who may narrow the automatic path or stop it. NIST asks for exactly that: mechanisms and assigned responsibilities to supersede, disengage or deactivate a system whose performance or outcomes do not match its intended use. In a business, that is a pause switch, a fallback to the manual process, and one named person allowed to use both without a meeting. The evaluations, logs and cost visibility behind those decisions are what our Infrastructure & Governance work puts in place, and what a review of a running system looks at first.

Two Indian reference points are worth knowing. The India AI Governance Guidelines, released by MeitY in November 2025, ask organisations to build human-in-the-loop mechanisms at critical decision points where appropriate, so that outputs can be reviewed, overridden or supplemented by human judgement before they cause harm. They also note that direct human oversight is ineffective in some settings, such as high-velocity algorithmic trading, where automated checks and system-level constraints do the work instead.

Where this has limits

  • These tests order a discussion; they do not produce a score. Two sensible people can place the same decision in different paths, which is why the choice should be written down and dated.
  • Confidence scores are not probabilities of being right. Treat them as one input among several, and never as the only condition for an automatic action.
  • A human review step does not by itself make a decision safe or lawful. If a regulator requires a qualified person, the requirement is about who decides, not about who clicks approve.
  • Thresholds tuned on a quiet period will misbehave in a peak one. Test them against your busiest month, not your calmest.

Frequently asked questions

Which decisions should AI never make on its own?

Anything irreversible, anything where a single error causes serious harm, and anything a law, contract or professional obligation reserves for a qualified person. In everyday operations that usually means releasing money outside agreed limits, changing master data such as bank details or a GSTIN, terminating a service, making a commitment to a customer, and any decision about a person’s credit, health, employment or admission.

How confident does a model need to be before it can act?

Confidence alone is the wrong gate. Pair it with conditions you can check: required fields present, sources agreeing with each other, the case falling inside a written business limit, and no exception flag raised. Then set the level from your own logs by finding where reviewers stopped changing anything, rather than from a number quoted by a vendor.

Is a human review step enough to make an AI system safe?

Only if the review is real. A reviewer needs the original evidence, a clear proposed action, authority to reject and the time to look. Without those, approval becomes a formality that adds delay and a false sense of control. Measure approval, edit and rejection rates and the seconds spent per case; if approval is near universal and instant, redesign the gate.

How do we stop reviewers approving everything automatically?

Reduce the volume reaching the gate so each case gets attention, show what is uncertain rather than only the recommendation, make rejecting as easy as approving, and rotate reviewers so nobody spends a whole day clicking. Add occasional test cases with known errors, and review the results with the team as a design problem rather than as a performance issue.

Should we start fully manual and automate decisions later?

Usually yes for consequential decisions. Run the system in preparation mode first: it gathers evidence and proposes an action while every case goes to a person. The approvals, edits and rejections from those weeks become the evidence for your first automatic path, and they show which cases were never suitable for one. It is slower to start and much easier to defend.

Where BYBO fits

Sources

  1. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1National Institute of Standards and Technology (NIST)
  2. India AI Governance Guidelines: Enabling Safe and Trusted AI InnovationMinistry of Electronics and Information Technology (MeitY), Government of India

General information for business readers, not legal, financial or regulatory advice. Examples are illustrative, not client work. Published 11 September 2026.

Find your first useful workflow.

Explore the Blueprint