CORAA
Blog/AI in Audit

Can ChatGPT Do My Audit? Where the Useful Part Ends

For correspondence, research and drafting it earns its seat. For audit procedures it fails on four structural counts — not because the model is weak, but because a chat window cannot hold an engagement. Where the boundary sits, and how to make it a habit.

CCORAA Team18 August 202610 min

Can ChatGPT Do My Audit? Where the Useful Part Ends

The question comes up in almost every firm conversation, and it usually gets one of two unhelpful answers. Either a flat refusal that treats the whole category as reckless, or an enthusiasm that has clearly never been near a working paper.

Neither is right. The useful answer is a boundary: there is a real and substantial set of audit-adjacent work where a general assistant is excellent, and there is a hard edge past which it is the wrong instrument. Most of the confusion in firms comes from not knowing where that edge sits.

The Part That Genuinely Works

A general chat assistant is good — properly good — at work made of language rather than evidence:

  • Engagement correspondence: covering letters, information request lists, chasing a slow client
  • Turning a standard or a circular into a plain-language brief for the juniors
  • First-draft checklists, templates and formulas
  • Editing. Taking a paragraph written at 1 a.m. and making it read like daylight

A manager preparing a new engagement might need a covering letter, an information request list, and a summary of what changed in a recent standard. That is most of a morning's work, and it compresses to a short drafting-and-review session. The saving is real.

The pattern connecting all of it: none of that work touches the client's books, and none of it produces something that goes into the file as evidence. That is not a coincidence. It is the boundary itself, visible from the useful side.

What an Audit Needs That a Chat Window Cannot Hold

The failures below are structural. They are not gaps a better model closes next year, because they follow from what a chat interface is rather than from how capable it is.

It holds no engagement

An audit is a connected object. Materiality, the ledger, planned procedures, evidence collected, issues still open, and the reporting consequence all reference each other. Change materiality and a dozen downstream conclusions move.

A conversation holds one question at a time. It does not know your materiality, what you already tested, or what last week's exception turned out to be. Every answer is produced in isolation from the engagement it is supposedly about.

It cannot be held to its answer

A general model can produce something fluent, specific and wrong — a misquoted clause, an interpretation that does not survive contact with the section, a plausible figure. Ask twice and you may get two different answers.

The difficulty is not that errors occur. Every source has errors. It is that these errors carry no signal that they are errors. A search result shows you its source and lets you judge it; a generated paragraph arrives already dressed as a conclusion. And the questions where you are least equipped to catch a mistake are precisely the ones you asked because you did not already know.

For a covering letter that risk costs nothing. For a tax position or a statutory conclusion it is the whole exposure.

It cannot show the path

A working paper's value is the chain: source transaction → procedure applied → evidence obtained → judgement formed → conclusion reached. That chain is what a reviewer follows and what a regulator asks for.

A generated answer typically cannot reproduce it. What you have is an assertion and no way to defend it — which is not a working paper, whatever it looks like on the page.

It forgets

Engagement knowledge compounds. What the client explained about that recurring adjustment, why the estimate was accepted last year, which control was already found ineffective — a good firm carries that forward, and it is much of what makes the second-year audit better than the first.

A fresh conversation starts at zero. An audit that starts at zero every year is a firm that never gets the benefit of its own history.

The Confidentiality Question Is Not Optional

The three failures above are professional-quality problems. This one is a legal one, and it deserves separating.

When client data goes into a public chat product, it has left your control. The duty of confidentiality has no efficiency exception. And an Indian client's ledger will ordinarily contain personal data somewhere in it — payroll, the vendor master, employee advances — which brings the Digital Personal Data Protection Act, 2023 into play, with the obligations sitting on your firm rather than on the vendor.

The tier matters enormously here, and most firms have never checked which one they are on. A consumer subscription, a business tier with a data processing agreement, and an API relationship are three different contractual positions, and only two of them are defensible for client data. That analysis is worked through in where your client's data actually goes, and what to actually buy in the CA firm's AI procurement guide.

Making the Boundary a Habit

A policy nobody remembers under deadline pressure is not a policy. The boundary has to reduce to something a second-year article can apply without deliberating.

The most reliable version we have seen is a single check, applied before the tool is opened rather than after:

Will the output of this task end up in the file, or does the input come out of the client's books?

Either one being true sends the work to a controlled system. Both being false leaves it in the assistant's territory — letters, summaries, research, drafting.

The reason this works better than a longer policy is that it asks about the work, not about the technology. Nobody has to assess a model's capabilities or remember which product is approved. They only have to look at what is in front of them, which they can already do.

Firms that write a page of rules get compliance for a fortnight. Firms that give people one question get it indefinitely.

What This Actually Means

The framing that has caused the most damage is treating this as a question about trust — whether AI is reliable enough for professional work. That framing produces bad decisions in both directions: firms that ban it and lose the genuine hours, and firms that let it near a conclusion because it seemed impressive on a demo.

It is better understood as a question of fit. A general assistant is built to produce language on demand, with no state, no evidence chain and no contractual perimeter. Those are not defects; they are the design, and they are exactly right for drafting a letter. They happen to be disqualifying for a procedure that depends on real books, a defensible chain, and continuity across years.

Different instruments, different jobs. The error was never in using the assistant. It is in taking it past the point where its design stops matching the work.

Frequently Asked Questions

Can ChatGPT perform a statutory audit?

No. It holds no engagement context, cannot reproduce the chain from source transaction to conclusion, does not carry engagement history between sessions, and can generate confidently incorrect technical output. It remains genuinely useful for correspondence, research and drafting around the audit.

Is it a confidentiality breach to paste a client ledger into ChatGPT?

Treat it as client information leaving your control, which engages your professional duty of confidentiality. Because ledgers routinely contain personal data — payroll, vendor master, employee advances — DPDP Act 2023 obligations are also engaged, and they attach to your firm. The product tier and the contract behind it change the position materially, so establish both before any client data is involved.

What can auditors safely use ChatGPT for?

Work made of language rather than evidence: engagement correspondence, information request lists, plain-language summaries of standards, first-draft checklists and templates, and editing text you have already written. This is a substantial and immediately available set of savings.

Why is a wrong AI answer more dangerous than a wrong search result?

A search result exposes its source and invites judgement. A generated answer arrives fluent and specific, often with a citation attached, and carries no signal distinguishing a sound answer from an unsound one — and it is most persuasive on exactly the questions you asked because you could not answer them yourself.

What is the difference between a chatbot and an audit system?

A chatbot answers the question in front of it. An audit system holds the engagement — materiality, ledger, procedures, evidence, open issues, reporting effect — and can show the path from a source transaction to a conclusion. One is an assistant. The other is where the file and the opinion live.

Topics
can ChatGPT do my auditChatGPT for auditorsis ChatGPT safe for audit workChatGPT client data confidentiality CAAI hallucination audit riskChatGPT vs audit software
Share
← Back to all articles
Keep reading

More in ai in audit.

Built for India · DPDPA compliant

Ready to automate your audit work.

See how Coraa reduces audit engagement time by 60%, from ledger scrutiny to working papers, all from one Tally import.

Run one complete audit free