Skip to content

Building a Cash-Flow Forecast With AI Without Trusting the Numbers Blindly

A synthetic 12-month case separates formula accuracy from assumption risk and tests three scenarios.

Editorial illustration of monthly cash-flow channels feeding a reservoir and three separate scenario basins.

Answer in brief: All 12 cash-balance formulas recomputed exactly. The first narrative missed a delayed-receivable assumption, proving that correct arithmetic can still support an incomplete story.

Can AI help explain a cash-flow forecast while deterministic formulas remain the authority? A useful answer has to be narrower than a product claim. This article tests a bounded workflow, publishes the scoring surface, and keeps consequential approval with a person. It does not turn a controlled result into personalized financial advice.

What we tested or analyzed

We used OpenAI GPT-5.6 in a Codex editorial session to explain a fictional forecast, then recomputed opening cash plus receipts minus payments for every month and compared three explicit scenarios.

The original asset is a twelve-month forecast, formulas, and base/downside/upside scenarios. The complete machine-readable table is available as CSV. The evidence visual below summarizes the primary criterion; its values are also written in text and shown in the table, so the chart is not the only way to obtain the result.

12 of 12 items passed the primary criterion; 1 required review, failed, or remained open.
Twelve-month forecast, formulas, and base/downside/upside scenarios. Original LuckyToKnow evidence, 2026-07-26.
Twelve-month scenario strip showing seven base months, three upside months, and two downside months, all with formula checks passed.
Controlled test using a synthetic 12-month case. Formula outcomes are verified; scenario labels are assumptions rather than predictions.

The measured result

All 12 cash-balance formulas recomputed exactly. The first narrative missed a delayed-receivable assumption, proving that correct arithmetic can still support an incomplete story.

The row-level outcome distribution was formula pass: 12. Those labels are deliberately more descriptive than one blended score. A partial, review, stale, exception, or unsupported row can carry a different operational risk from a plainly wrong row, so the CSV preserves the reason beside the disposition.

Complete scored asset. The same rows are available in the downloadable CSV.
ItemOutcomeEvidence or note
Janformula passbase
Febformula passbase
Marformula passbase
Aprformula passbase
Mayformula passupside
Junformula passupside
Julformula passdownside
Augformula passdownside
Sepformula passbase
Octformula passbase
Novformula passupside
Decformula passbase

Reading the evidence row by row

  • Jan was recorded as formula pass. The evidence note is “base”; the disposition remains visible so it cannot be averaged away.
  • Feb was recorded as formula pass. The evidence note is “base”; the disposition remains visible so it cannot be averaged away.
  • Mar was recorded as formula pass. The evidence note is “base”; the disposition remains visible so it cannot be averaged away.
  • Apr was recorded as formula pass. The evidence note is “base”; the disposition remains visible so it cannot be averaged away.
  • May was recorded as formula pass. The evidence note is “upside”; the disposition remains visible so it cannot be averaged away.
  • Jun was recorded as formula pass. The evidence note is “upside”; the disposition remains visible so it cannot be averaged away.
  • Jul was recorded as formula pass. The evidence note is “downside”; the disposition remains visible so it cannot be averaged away.
  • Aug was recorded as formula pass. The evidence note is “downside”; the disposition remains visible so it cannot be averaged away.
  • Sep was recorded as formula pass. The evidence note is “base”; the disposition remains visible so it cannot be averaged away.
  • Oct was recorded as formula pass. The evidence note is “base”; the disposition remains visible so it cannot be averaged away.
  • Nov was recorded as formula pass. The evidence note is “upside”; the disposition remains visible so it cannot be averaged away.
  • Dec was recorded as formula pass. The evidence note is “base”; the disposition remains visible so it cannot be averaged away.

The expected label or control was fixed before review. The visible note explains why the row received its disposition. The chart uses the published primary criterion, but the table is authoritative because it preserves exceptions that a single percentage would hide.

How to reproduce the check

  1. Download the CSV and read its labels, units, and synthetic/public-data notice before using it.
  2. Write the expected answers or decision rule before looking at a model response.
  3. Use the same bounded prompt and record the model or tool, access surface, and date.
  4. Preserve the raw response. Break prose into atomic claims rather than grading the tone of the whole answer.
  5. Recompute arithmetic with deterministic formulas and verify definitions against the linked primary sources.
  6. Record correct, partial, wrong, uncertain, and refused outcomes separately. Do not silently repair the model output before scoring it.
  7. Repeat material checks after a model, source, or workflow changes.

What the result means

The value of this result is diagnostic. It shows where a structured assistant can reduce search, formatting, or first-pass review work. It does not transfer responsibility for the underlying decision. A “pass” means the row met the published rule in this test, on this date, with these inputs.

The errors and open items matter more than a polished average. In money and business workflows, one missed assumption, stale fact, false match, or overconfident definition can dominate many correct low-risk rows. That is why the artifact keeps row-level outcomes and why a human reviews exceptions rather than receiving only a percentage.

Reproducibility also has limits. A reader can repeat the steps and inspect the same answer key, but a probabilistic model may not return identical wording. A useful rerun should therefore compare atomic claims, calculations, citations, and escalation decisions—not superficial phrasing.

Why this topic needs its own boundary

Business-finance workflows combine documents, accounting policy, timing, and deterministic arithmetic. An assistant can help prepare a queue or explanation, but the chart of accounts, reconciliation rule, formula, materiality threshold, and posting authority need accountable ownership.

Synthetic records make the test repeatable without exposing suppliers, employees, customers, bank accounts, or tax information. They also make the boundary visible: passing a clean example does not validate a production feed with duplicates, foreign exchange, split transactions, and incomplete documents.

A safer operating workflow

  • Lock formulas separately from narrative.
  • Name every timing assumption.
  • Run downside and delayed-payment scenarios.
  • Reconcile opening and closing cash.
  • Escalate financing or insolvency questions to qualified professionals.

How each control changes the decision

Control 1: Lock formulas separately from narrative. For this test, that control answers the bounded question “Can AI help explain a cash-flow forecast while deterministic formulas remain the authority?” without extending the result into an untested decision.

Control 2: Name every timing assumption. For this test, that control answers the bounded question “Can AI help explain a cash-flow forecast while deterministic formulas remain the authority?” without extending the result into an untested decision.

Control 3: Run downside and delayed-payment scenarios. For this test, that control answers the bounded question “Can AI help explain a cash-flow forecast while deterministic formulas remain the authority?” without extending the result into an untested decision.

Control 4: Reconcile opening and closing cash. For this test, that control answers the bounded question “Can AI help explain a cash-flow forecast while deterministic formulas remain the authority?” without extending the result into an untested decision.

Control 5: Escalate financing or insolvency questions to qualified professionals. For this test, that control answers the bounded question “Can AI help explain a cash-flow forecast while deterministic formulas remain the authority?” without extending the result into an untested decision.

Keep data collection, model preparation, deterministic validation, and approval as separate stages. Use the least sensitive input that can answer the question. If removing personal or confidential data makes the result ambiguous, route the case to an approved human process instead of restoring secrets to an unapproved tool.

Calculations need an independent formula; current facts need a current primary source; classifications need an “uncertain” route; and irreversible actions need explicit authorization outside the model. Logs should capture the version, prompt, source date, output, reviewer, correction, and final disposition without retaining unnecessary personal data.

Limitations and professional boundary

Forecasts depend on assumptions. This synthetic case has no tax, financing covenant, exchange-rate, or customer-default model.

This publication provides general educational information. It does not know a reader’s finances, duties, jurisdiction, contracts, tax treatment, credit position, or risk tolerance. A qualified financial, accounting, tax, legal, lending, security, or other professional should review decisions with material consequences.

Primary sources

Verified 2026-07-26. Primary-source links can change; use the publication date and linked source to check for a newer version.

Bottom line

All 12 cash-balance formulas recomputed exactly. The first narrative missed a delayed-receivable assumption, proving that correct arithmetic can still support an incomplete story. The practical lesson is to make AI produce inspectable work inside a controlled process—not to make fluency the final control.