Can AI Write Your STR Narrative? What LLMs Do Well and Where They Create Risk
The Question Every Compliance Team Is Now Asking
Narrative writing is the slowest part of preparing a Suspicious Transaction Report and the part analysts like least. It is also the part that most obviously resembles what large language models are good at: taking messy inputs and producing fluent, structured prose. So the question arrives on its own, usually from someone outside compliance: why is an analyst still typing this by hand?
It is a fair question and it deserves a real answer rather than a reflexive no. There is no rule that a human must personally type every character of an STR narrative. The obligation is that the report be accurate and complete, and that the grounds for suspicion be the reporting entity's own — reached, held, and defensible. Those are different things, and the distinction is where the entire answer lives.
This post covers what LLMs genuinely do well in STR drafting, where they fail in ways that are hard to see until an examination, the confidentiality problem that most pilots discover far too late, and the division of labour that actually survives scrutiny.
The One-Line Version
A model can describe what happened. It cannot conclude that what happened is suspicious. Every strength and every limitation below is a consequence of that boundary.
What an STR Narrative Actually Has to Accomplish
Before judging whether a model can write one, it helps to be precise about the job. A strong STR narrative does four things at once, and they are not equally automatable.
It Describes
The factual chronology — who, what, when, how much, through which accounts and instruments, and how the activity connects across parties and time. This is reporting, and it is fully contained in your data.
It Explains Why That Is Unusual
The activity set against the expected pattern: this customer's stated purpose, occupation, and history, or the norms of their peer group. Partly in your data, partly in what staff know.
It States the Grounds
The specific facts that took this from odd to suspicious, and the reasoning connecting them to a money laundering or terrorist activity financing offence. This is judgment. It belongs to the reporting entity and nothing else.
It Records the Action Taken
What the reporting entity did in response — restrictions, exits, escalation, further monitoring. Factual, but drawn from decisions made outside the transaction record.
That first component is often more than half the word count and almost none of the risk. The third is a fraction of the word count and nearly all of the risk. Any sensible use of AI here follows that split rather than treating the narrative as one undifferentiated block of writing.
Four Things LLMs Genuinely Do Well Here
Compressing a Case File Into a Chronology
Given structured transaction data and case notes, models are reliably good at producing an ordered, readable account of what occurred. This is the descriptive half of the narrative, it is the bulk of the typing, and it is grounded entirely in records you already hold. Analysts editing a generated chronology work substantially faster than analysts starting from an empty box.
Enforcing Consistency Across Analysts
Ten analysts produce ten registers, ten levels of detail, and ten habits about what to include. A model applying the same structure every time raises the floor. The weakest narratives in most programs are not wrong — they are thin and idiosyncratic, and that is exactly the failure mode consistency fixes.
Checking for Omissions
Models are effective as a reviewer against a checklist: is the time period stated, are all involved accounts named, is the aggregate amount given, is the action taken recorded, does the narrative reference every transaction attached to the report. Used this way the model criticises a human's draft rather than producing one, which inverts the risk profile entirely.
Translating and Levelling Language
For teams operating in more than one language, or where analysts are writing in a second language, models materially improve clarity and reduce the chance that a well-reasoned suspicion reads as confused. This is an underrated and genuinely low-risk use.
Note the shape of that list. Every item is either descriptive work drawn from records, or quality control applied to a human's output. None of it involves the model deciding anything.
Five Things They Do Badly
Generic Grounds Language
Asked to explain why activity is suspicious, models produce plausible, fluent, non-specific reasoning — structuring, layering, inconsistent with stated purpose — that could be pasted into any report about anything. This is the single most damaging failure, because it looks finished. A narrative that names no specific fact is not grounds; it is vocabulary.
Confident Fabrication of Specifics
Models fill gaps. Given partial data they will smooth over a missing counterparty, infer a relationship that was never established, round a figure, or assert a sequence that the records do not support. In ordinary writing this is a nuisance. In a report filed to a regulator, an invented fact is a defect in a legal filing that your organisation signed.
No Access to Off-System Context
The decisive detail is frequently something never captured in a system: what the customer said when questioned, what the branch or support agent observed, what the relationship manager already suspected. The model cannot see any of it, so it produces a narrative that is complete with respect to the data and incomplete with respect to the truth — while reading as authoritative either way.
It Cannot Hold the Assessment
Reasonable grounds to suspect is a determination made by the reporting entity. A model has no standing to make it, no accountability for it, and no ability to be questioned about it later. When an examiner asks why this report was filed, the answer cannot route through software, and a narrative whose reasoning was authored by a model has no author.
Confidentiality and Where the Data Goes
STR-related information carries strict confidentiality obligations, and disclosing that a report has been or will be made is prohibited. Sending case material to a third-party model means customer identifiers, transaction detail, and the fact of a pending report leave your environment. Retention, training use, sub-processors, jurisdiction of processing, and who inside your own organisation gains visibility all become live compliance questions.
The Failure That Is Hardest to Catch
Fabrication gets found in review because someone checks it against the data. Generic grounds language does not, because there is nothing to check it against — it is not false, it is empty. It passes review, files cleanly, and only surfaces as a problem when someone assesses the quality of your reporting as a whole.
The Division of Labour That Holds Up
Once the narrative is broken into its components, the allocation is not really a judgment call.
| Narrative component | Source of truth | Draft by | Reviewed by |
|---|---|---|---|
| Factual chronology | Transaction and account records | Model, from structured data | Analyst, against source |
| Parties and relationships | KYC and account records | Model, from structured data | Analyst, against source |
| Deviation from expected pattern | Records plus staff knowledge | Model drafts the data-derived part | Analyst adds what is not in systems |
| Grounds for suspicion | Analyst's assessment | Analyst only | Compliance Officer |
| Action taken | Internal decisions | Analyst | Compliance Officer |
The rule of thumb: the model may state what is in the records; the analyst must state what it means. If a sentence would begin "the reporting entity suspects" or "these facts indicate", a person writes it.
A second rule is worth adopting alongside it. Every generated sentence must be traceable to a specific record. If an analyst cannot point at the transaction, the account, or the note that a sentence came from, that sentence comes out — not gets softened, comes out. This is the only reliable defence against smooth, confident filler.
What to Put in Policy Before Anyone Uses This
Most AI drafting in compliance teams starts informally — an analyst pastes a case summary into a chat tool to get past a blank page. By the time anyone writes a policy, the practice already exists. Getting ahead of that is cheaper than discovering it during an internal audit.
Name What Is Permitted and What Is Not
Be specific at the level of narrative component, not at the level of 'AI'. Permitted: drafting factual chronology from case data inside the approved system. Not permitted: generating grounds for suspicion, or pasting case material into any tool not on the approved list. A blanket ban gets ignored; a precise rule gets followed.
Fix Where the Data May Go
Approved tooling only, with contractual terms covering retention, exclusion from model training, sub-processors, and jurisdiction of processing. Document it. This is the question a privacy review will ask first and the one most pilots cannot answer.
Require Attribution to Source
Any generated statement must trace to a record in the case file. Make this an explicit review step rather than an assumption, because the whole point is that unsupported sentences do not look unsupported.
Record That a Human Wrote the Grounds
The audit trail should show which analyst authored the assessment and when — distinct from who assembled the report. If your system cannot distinguish those two acts, you cannot demonstrate the distinction later, and demonstrating it is the entire point.
Review Output Quality on a Cycle
Sample filed narratives specifically for genericness: do the grounds cite named facts from this case, or could this paragraph appear in any report? Model behaviour and analyst habits both drift, and this failure mode is invisible unless somebody looks for it deliberately.
Train Analysts on the Failure Modes, Not Just the Tool
An analyst who knows that models fabricate specifics and default to generic reasoning reviews differently from one who has only been shown how to click generate. The training that matters is about what to distrust.
An Honest Summary
Used well, an LLM removes a large share of the typing from STR preparation, makes narratives more consistent across a team, and catches omissions a tired reviewer would miss. Those are real gains and they are worth having.
Used badly, it produces reports that are longer, more fluent, more uniform, and less true — with grounds that assert suspicion without evidencing it, and occasional invented facts sitting inside a filing your organisation is accountable for. The dangerous property of that output is not that it is obviously wrong. It is that it reads better than what your analysts write themselves.
The difference between the two outcomes is almost entirely a matter of where you draw the line, and the line is the same one that governs the rest of the reporting workflow: automate everything around the decision, and leave the decision alone.
How Quantoflow Approaches This
Quantoflow uses structure rather than generation where the stakes are highest.
- Case data assembled and mapped so the descriptive groundwork is done from records rather than recalled or re-typed
- Structured narrative prompts that ask the analyst for pattern, specific facts, and grounds — guaranteeing coverage without supplying the reasoning
- Traceability to source records so every factual statement in a narrative can be tied back to the transaction or record it came from
- Separate audit capture of who assembled the report and who authored the assessment, timestamped as the work happens
- Data stays in your environment, so confidentiality obligations are a design constraint rather than a question raised after a pilot
The chronology gets easier to produce. The grounds stay yours to write, and stay defensible because of it.
Working Out Where AI Fits in Your Reporting Process?
If your team is already experimenting with generated narratives — formally or otherwise — the useful conversation is about where the line sits, not whether to have one.
Talk Through Your STR Workflow
Citations
- FINTRAC — Suspicious Transaction Reporting Guidance https://fintrac-canafe.canada.ca/guidance-directives/transaction-operation/Guide2/str-eng
- FINTRAC — Compliance Program Requirements https://fintrac-canafe.canada.ca/guidance-directives/compliance-conformite/Guide4/4-eng
- Proceeds of Crime (Money Laundering) and Terrorist Financing Act https://laws-lois.justice.gc.ca/eng/acts/P-24.501/