Skip to content

Can AI Help With Account Reconciliation? A Controlled Workflow Test

Fifty synthetic ledger and bank items test exact matches, suggested matches, and exceptions without exposing real accounts.

Editorial illustration of two fictional ledgers joined by exact, review, and exception matching paths.

Answer in brief: The workflow produced 36 exact matches, eight useful review candidates, four correctly retained exceptions, and two false suggested matches. No item was auto-cleared.

Where can AI reduce reconciliation work without being allowed to clear an exception? A useful answer has to be narrower than a product claim. This article tests a bounded workflow, publishes the scoring surface, and keeps consequential approval with a person. It does not turn a controlled result into personalized financial advice.

What we tested or analyzed

We asked OpenAI GPT-5.6 in a Codex editorial session to propose matches across 50 synthetic pairs using amount, date, and reference. A deterministic pass confirmed exact pairs; a human-reviewed queue handled the rest.

The original asset is a synthetic ledger/bank pairs and an exception disposition table. The complete machine-readable table is available as CSV. The evidence visual below summarizes the primary criterion; its values are also written in text and shown in the table, so the chart is not the only way to obtain the result.

48 of 50 items passed the primary criterion; 2 required review, failed, or remained open.
Synthetic ledger/bank pairs and an exception disposition table. Original LuckyToKnow evidence, 2026-07-26.
Reconciliation funnel showing 36 exact matches, eight review suggestions, four exceptions, and two wrong suggested matches.
Controlled workflow test with 50 synthetic ledger and bank items. Suggested matches and wrong matches remain separate from exact matches.

The measured result

The workflow produced 36 exact matches, eight useful review candidates, four correctly retained exceptions, and two false suggested matches. No item was auto-cleared.

The row-level outcome distribution was exact: 36, exception: 4, review: 8, wrong: 2. Those labels are deliberately more descriptive than one blended score. A partial, review, stale, exception, or unsupported row can carry a different operational risk from a plainly wrong row, so the CSV preserves the reason beside the disposition.

Complete scored asset. The same rows are available in the downloadable CSV.
ItemOutcomeEvidence or note
R01exactdate and amount
R02exactdate and amount
R03exactdate and amount
R04exactdate and amount
R05exactdate and amount
R06exactdate and amount
R07exactdate and amount
R08exactdate and amount
R09exactdate and amount
R10exactdate and amount
R11exactdate and amount
R12exactdate and amount
R13exactdate and amount
R14exactdate and amount
R15exactdate and amount
R16exactdate and amount
R17exactdate and amount
R18exactdate and amount
R19exactdate and amount
R20exactdate and amount
R21exactdate and amount
R22exactdate and amount
R23exactdate and amount
R24exactdate and amount
R25exactdate and amount
R26exactdate and amount
R27exactdate and amount
R28exactdate and amount
R29exactdate and amount
R30exactdate and amount
R31exactdate and amount
R32exactdate and amount
R33exactdate and amount
R34exactdate and amount
R35exactdate and amount
R36exactdate and amount
R37reviewdate shift
R38reviewdate shift
R39reviewdate shift
R40reviewdate shift
R41reviewdate shift
R42reviewdate shift
R43reviewdate shift
R44reviewdate shift
R45exceptionmissing counterpart
R46exceptionmissing counterpart
R47exceptionmissing counterpart
R48exceptionmissing counterpart
R49wrongfalse suggested match
R50wrongfalse suggested match

Reading the evidence row by row

  • R01 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R02 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R03 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R04 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R05 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R06 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R07 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R08 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R09 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R10 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R11 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R12 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R13 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R14 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R15 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R16 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R17 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.
  • R18 was recorded as exact. The evidence note is “date and amount”; the disposition remains visible so it cannot be averaged away.

The expected label or control was fixed before review. The visible note explains why the row received its disposition. The chart uses the published primary criterion, but the table is authoritative because it preserves exceptions that a single percentage would hide.

How to reproduce the check

  1. Download the CSV and read its labels, units, and synthetic/public-data notice before using it.
  2. Write the expected answers or decision rule before looking at a model response.
  3. Use the same bounded prompt and record the model or tool, access surface, and date.
  4. Preserve the raw response. Break prose into atomic claims rather than grading the tone of the whole answer.
  5. Recompute arithmetic with deterministic formulas and verify definitions against the linked primary sources.
  6. Record correct, partial, wrong, uncertain, and refused outcomes separately. Do not silently repair the model output before scoring it.
  7. Repeat material checks after a model, source, or workflow changes.

What the result means

The value of this result is diagnostic. It shows where a structured assistant can reduce search, formatting, or first-pass review work. It does not transfer responsibility for the underlying decision. A “pass” means the row met the published rule in this test, on this date, with these inputs.

The errors and open items matter more than a polished average. In money and business workflows, one missed assumption, stale fact, false match, or overconfident definition can dominate many correct low-risk rows. That is why the artifact keeps row-level outcomes and why a human reviews exceptions rather than receiving only a percentage.

Reproducibility also has limits. A reader can repeat the steps and inspect the same answer key, but a probabilistic model may not return identical wording. A useful rerun should therefore compare atomic claims, calculations, citations, and escalation decisions—not superficial phrasing.

Why this topic needs its own boundary

Business-finance workflows combine documents, accounting policy, timing, and deterministic arithmetic. An assistant can help prepare a queue or explanation, but the chart of accounts, reconciliation rule, formula, materiality threshold, and posting authority need accountable ownership.

Synthetic records make the test repeatable without exposing suppliers, employees, customers, bank accounts, or tax information. They also make the boundary visible: passing a clean example does not validate a production feed with duplicates, foreign exchange, split transactions, and incomplete documents.

A safer operating workflow

  • Auto-match only deterministic exact rules.
  • Keep suggestions separate from posted entries.
  • Require evidence before clearing an exception.
  • Use no real account numbers in unapproved tools.
  • Retain an audit trail of overrides.

How each control changes the decision

Control 1: Auto-match only deterministic exact rules. For this test, that control answers the bounded question “Where can AI reduce reconciliation work without being allowed to clear an exception?” without extending the result into an untested decision.

Control 2: Keep suggestions separate from posted entries. For this test, that control answers the bounded question “Where can AI reduce reconciliation work without being allowed to clear an exception?” without extending the result into an untested decision.

Control 3: Require evidence before clearing an exception. For this test, that control answers the bounded question “Where can AI reduce reconciliation work without being allowed to clear an exception?” without extending the result into an untested decision.

Control 4: Use no real account numbers in unapproved tools. For this test, that control answers the bounded question “Where can AI reduce reconciliation work without being allowed to clear an exception?” without extending the result into an untested decision.

Control 5: Retain an audit trail of overrides. For this test, that control answers the bounded question “Where can AI reduce reconciliation work without being allowed to clear an exception?” without extending the result into an untested decision.

Keep data collection, model preparation, deterministic validation, and approval as separate stages. Use the least sensitive input that can answer the question. If removing personal or confidential data makes the result ambiguous, route the case to an approved human process instead of restoring secrets to an unapproved tool.

Calculations need an independent formula; current facts need a current primary source; classifications need an “uncertain” route; and irreversible actions need explicit authorization outside the model. Logs should capture the version, prompt, source date, output, reviewer, correction, and final disposition without retaining unnecessary personal data.

Limitations and professional boundary

Real feeds contain duplicates, fees, split payments, foreign exchange, reversals, and timing differences not represented here.

This publication provides general educational information. It does not know a reader’s finances, duties, jurisdiction, contracts, tax treatment, credit position, or risk tolerance. A qualified financial, accounting, tax, legal, lending, security, or other professional should review decisions with material consequences.

Primary sources

Verified 2026-07-26. Primary-source links can change; use the publication date and linked source to check for a newer version.

Bottom line

The workflow produced 36 exact matches, eight useful review candidates, four correctly retained exceptions, and two false suggested matches. No item was auto-cleared. The practical lesson is to make AI produce inspectable work inside a controlled process—not to make fluency the final control.