Garbage In, Audit Out? How AI Actually Handles Badly-Kept Books
"If the underlying data is poor, is the system effective?" is one of the most reasonable objections to any AI audit claim, and it deserves a direct answer rather than a dismissal. Most demos run on clean sample datasets, because clean data makes any tool look impressive. Real client books rarely look like that — a part-time accountant's inconsistent naming conventions, blank narrations on a chunk of vouchers, GSTINs that were typed in wrong two years ago and never corrected. The question isn't whether AI works on perfect data. Everything works on perfect data. The question is what happens on the books you actually get.
Why "garbage in, garbage out" isn't quite the right frame for audit AI
The classic worry is that an AI system fed messy data will confidently produce wrong conclusions, amplifying the mess rather than catching it. That's a real risk for a system designed to auto-conclude. It's a much smaller risk for a system designed to flag and suggest, with a human confirming or overriding every classification before it becomes part of the working papers. The right question to ask any vendor isn't "does your AI handle bad data perfectly" — no system does — it's "what does your AI do when it's not confident, and does it say so, or does it guess silently."
What a tiered classification approach actually looks like
On CORAA, ledger classification runs across three tiers rather than one black-box pass. Tier 1 is deterministic — Tally group, name keywords, and standard-group matching resolve the large majority of ledgers instantly, at high confidence, with no ambiguity involved. Ledgers Tier 1 can't settle — the ones with unclear naming, unusual structure, or genuinely ambiguous grouping — go to Tier 2, where an LLM reads a sample of the actual voucher narrations attached to that ledger and proposes a classification with its stated reason. Tier 3 checks how auditors on other engagements confirmed similar ledgers as a further signal. Every suggestion arrives with its reason attached, and none of it is treated as final — the auditor confirms or overrides, and only the confirmed value becomes part of the record.
That structure is the actual answer to "what happens with messy data": messy ledgers are exactly the ones that fail Tier 1's clean deterministic match and get routed into a slower, narration-reading, human-reviewed path — the system is designed to be less confident on the harder cases, not equally confident on everything.
What "built for real data" means in the red-flag detection itself
Journal-entry testing is designed around the reality of how books actually get entered, not an idealized version. Empty narration at a material value is itself one of the named red-flag categories the engine looks for — a blank narration isn't treated as a data-quality nuisance to route around, it's treated as a signal worth surfacing on its own, since a material entry with no explanation attached is often exactly the kind of thing worth a second look regardless of what it turns out to be. The same testing is built to handle blank or garbled GSTINs, amended and re-posted invoices, inconsistent account naming, and free-text narrations as the normal input, not exceptions that break the system — because that's what real Tally exports actually look like, not a curated demo dataset.
What good data-quality handling actually looks like in a demo
If you're evaluating a tool, don't let a vendor demo on their own clean sample file and extrapolate. Bring your messiest recent client — the one with the account naming nobody's ever cleaned up, the ledger a departed bookkeeper set up wrong three years ago — and watch what the tool actually does with it. Does it flag the ambiguous ledgers for review instead of confidently mis-classifying them? Does a blank narration on a large transaction get surfaced as worth looking at, or does it just get skipped because there's nothing to parse? That's the real test, not a demo optimized to look flawless.
Frequently Asked Questions
Does AI ledger classification work reliably on messy, inconsistently-named accounts?
The design goal isn't uniform confidence on everything — it's routing cleanly-structured ledgers through fast, deterministic matching, while genuinely ambiguous ones get a slower narration-reading pass and always require the auditor's confirmation before anything becomes final.
What happens when a voucher has no narration at all?
An empty narration at a material value is itself a named red-flag category the engine looks for — it's treated as a signal worth surfacing, not skipped or silently ignored because there's no text to parse.
Is journal-entry testing built for clean sample data or real Tally exports?
For real exports specifically — blank or garbled GSTINs, amended and re-posted entries, inconsistent account naming, and free-text narrations are the expected input the testing is built to handle, not edge cases that break it.
How should I actually evaluate whether a tool handles bad data well?
Bring your messiest real client's data to the demo, not a clean sample file — and watch specifically whether ambiguous cases get flagged for review or confidently (and silently) misclassified.
Related: Scrutiny module · Start a free trial