Where AI Earns Its Place, and Where a Rule Is Safer
The useful question is not what AI can do. It is what uncertainty the organisation will accept in exchange for speed, scale or creative range.
Generative AI is probabilistic. It does not retrieve one predetermined answer the way a rules-based system does. It produces a likely response. That makes task selection a matter of commercial judgement rather than technical enthusiasm. This post takes one decision from the operating model and works through it: where AI adds value, and where a rule is safer.
The strongest use cases tolerate interpretation
AI earns its place where the work tolerates interpretation. Four conditions tend to hold at once:
- Unstructured input. Twelve call transcripts, a loose brief, a data room to triage, a long report to adapt for several audiences.
- More than one acceptable answer. Patterns, contradictions, alternative territories, first-pass variations.
- Reviewed before it matters. Someone who knows the source material owns the final output.
- Variation is useful. Range before judgement, not one mathematically correct synthesis.
The value came less from replacing a role than from changing the starting point. Not a blank page, an unfiltered document set or an empty search box, but a provisional answer to interrogate.
A provisional answer is useful when the next step is judgement. It is dangerous when the next step is publication, payment or exclusion.
The same verb can carry different accountability
Five task types recur, and each one turns from useful to risky at a different point.
- Synthesis. Well suited to reducing a large body of material into themes for review. Changes character when omissions could change a regulatory, legal or commercial interpretation.
- Ideation. Well suited to novelty and variation at the exploratory stage. Changes character when volume is mistaken for insight, or ideas go untested against customer, brand or market.
- Transformation. Well suited to changing tone, length, format or reading level of content that already exists. Changes character when it starts altering meaning, substantiation or emphasis.
- Extraction and classification. Well suited to stable categories, limited consequence per mistake and real scale. Changes character when the classification decides eligibility, treatment or compliance status.
- Personalisation. Well suited to reorganising approved material around known needs. Changes character when the system infers sensitive traits, makes unsupported claims or produces variants nobody can review.
The category never settles it on its own. "Summarisation" can mean internal notes from a meeting, or a set of risk disclosures reduced for a consumer audience. The technical action is similar. The accountability is not. Judge the consequence of the output, not the name of the technique.
When a deterministic rule beats a model
Repeatability is often more valuable than intelligence. The visibility of generative AI has created an odd bias: teams now ask how AI might solve a problem even where the outcome is fixed and variation has no value.
Many processes do not need intelligence. They need consistency. Applying a fee schedule, routing a request, checking mandatory fields, selecting the approved disclaimer, appending a reference to a filing. The same input should give the same output every time. A rule, template or automation is easier to test, explain and maintain.
Put a model on that work and the process gets less dependable. Identical cases can be treated differently, the verification burden goes up, and failure gets harder to diagnose. In a regulated firm, evidence of consistency can matter as much as the result.
When you already know the rule, encode the rule.
Cost of error sets the control, not the volume
Use cases get prioritised by frequency and time saved. Both matter. Neither says whether the task is suitable. Four dimensions do:
- Consequence. What happens when the answer is wrong?
- Detectability. Will anyone notice before it is used?
- Reversibility. Can it be corrected without lasting harm?
- Explainability. Can you reconstruct why it was produced and who approved it?
A five-minute task done thousands of times looks like an obvious automation target. If the error is hard to detect, touches a client or creates a misleading statement, the calculation changes. A low-volume task can be the better bet if its outputs are reversible and easy to inspect.
Oversight is not an approval box at the end. The reviewer needs knowledge, time and authority to challenge the output. One person waving through thousands of low-level classifications is the appearance of oversight. Thresholds, sampling and exception routing often do the job better. A team that cannot verify an output has not saved the verification time. It has deferred it.
Selecting the approach: five routes and eight questions
AI should not be the default. Most selection exercises begin one decision too late: a platform is bought, a pilot group is formed, and now the question is where the technology can be made to fit. Start earlier. Define the decision the workflow has to make, the evidence available, and the consequence of being wrong. Only then does technology become relevant.
There are five routes, not one:
- Deterministic automation. Known rule, fixed result.
- Predictive model. Pattern in historical data, provided the data is representative and the use is tested.
- Generative AI. Ambiguous input, useful variation.
- Manual judgement. Accountability or negotiation that cannot be reduced to the available information.
- Hybrid workflow. Mixed requirements. In regulated work this is usually the most credible route, and the human in it should carry real authority, not a routine approval click.
Eight questions stop business value and technical novelty collapsing into a single score:
- How variable are the inputs?
- One correct output, or a range of acceptable ones?
- What is the cost of error?
- How easily will an error be detected?
- Does it involve sensitive or personal data?
- Must the result be explained to a customer, regulator or reviewer?
- Does scale create enough value to justify the control burden?
- Will the organisation own the workflow after the pilot user moves on?
Value and risk are two axes, not one score. A frequent task can still be a poor candidate if mistakes are hard to detect. A sensitive task can stay viable if the input is minimised and the decision stays with a qualified person. A creative task can earn its place at low volume because the value is range, not minutes.
Once you have chosen the method, the next job is holding it in place, which we cover in governed workflows. The wider maturity picture is in activity is not capability.
To see the full screening matrix, download the full field note. If you want a plain read on which of the five routes fits, send us the use case and we will tell you whether it passes.