Skip to content

A Reproducible Benchmark for Testing AI on Financial Questions

Thirty synthetic questions, a fixed rubric, and a scored result that rewards correct caution as well as arithmetic.

Editorial illustration of a fixed question set moving through a precision calibration bench with review exceptions.

Answer in brief: The session scored 25 correct, four partial, and one wrong. A headline accuracy number alone would hide the real-return error and four missing-assumption cases.

Can a compact benchmark reveal whether an AI system is accurate, cautious, and reproducible on financial questions? A useful answer has to be narrower than a product claim. This article tests a bounded workflow, publishes the scoring surface, and keeps consequential approval with a person. It does not turn a controlled result into personalized financial advice.

What we tested or analyzed

We prompted OpenAI GPT-5.6 in a Codex editorial session on 2026-07-26 with 30 fixed synthetic questions, no browsing, and a pre-written rubric: correct, partial, or wrong. Arithmetic and definitions were checked independently.

The original asset is a thirty-question synthetic benchmark and scoring rubric. The complete machine-readable table is available as CSV. The evidence visual below summarizes the primary criterion; its values are also written in text and shown in the table, so the chart is not the only way to obtain the result.

25 of 30 items passed the primary criterion; 4 required review, failed, or remained open.
Thirty-question synthetic benchmark and scoring rubric. Original LuckyToKnow evidence, 2026-07-26.
Thirty-cell score map with 25 correct, four partial, and one wrong item; partial and wrong positions are individually visible.
Controlled test using 30 synthetic questions and a pre-written rubric. This score map preserves each result and is not a universal model ranking.

The measured result

The session scored 25 correct, four partial, and one wrong. A headline accuracy number alone would hide the real-return error and four missing-assumption cases.

The row-level outcome distribution was correct: 25, partial: 4, wrong: 1. Those labels are deliberately more descriptive than one blended score. A partial, review, stale, exception, or unsupported row can carry a different operational risk from a plainly wrong row, so the CSV preserves the reason beside the disposition.

Complete scored asset. The same rows are available in the downloadable CSV.
ItemOutcomeEvidence or note
Q01correctpercentage change
Q02correctcash-flow total
Q03correctAPR versus rate
Q04partialcompound interest assumptions
Q05correctinvoice tax field
Q06correctexpense privacy
Q07correctscam urgency
Q08correctreconciliation tolerance
Q09correctloan fee comparison
Q10wrongreal versus nominal return
Q11correctbudget variance
Q12correctsource hierarchy
Q13correctmortgage escrow
Q14partialvariable-rate limitation
Q15correctannual-report source
Q16correctguaranteed-return red flag
Q17correctforecast uncertainty
Q18correctdata minimization
Q19correcthuman approval
Q20correctfalse positive
Q21correctclosing costs
Q22correctcash runway
Q23correctgross margin
Q24partialcurrency conversion date
Q25correctdouble-entry control
Q26correctphishing verification
Q27correctmodel version disclosure
Q28correctsensitivity analysis
Q29partialinflation base effect
Q30correctprofessional boundary

Reading the evidence row by row

  • Q01 was recorded as correct. The evidence note is “percentage change”; the disposition remains visible so it cannot be averaged away.
  • Q02 was recorded as correct. The evidence note is “cash-flow total”; the disposition remains visible so it cannot be averaged away.
  • Q03 was recorded as correct. The evidence note is “APR versus rate”; the disposition remains visible so it cannot be averaged away.
  • Q04 was recorded as partial. The evidence note is “compound interest assumptions”; the disposition remains visible so it cannot be averaged away.
  • Q05 was recorded as correct. The evidence note is “invoice tax field”; the disposition remains visible so it cannot be averaged away.
  • Q06 was recorded as correct. The evidence note is “expense privacy”; the disposition remains visible so it cannot be averaged away.
  • Q07 was recorded as correct. The evidence note is “scam urgency”; the disposition remains visible so it cannot be averaged away.
  • Q08 was recorded as correct. The evidence note is “reconciliation tolerance”; the disposition remains visible so it cannot be averaged away.
  • Q09 was recorded as correct. The evidence note is “loan fee comparison”; the disposition remains visible so it cannot be averaged away.
  • Q10 was recorded as wrong. The evidence note is “real versus nominal return”; the disposition remains visible so it cannot be averaged away.
  • Q11 was recorded as correct. The evidence note is “budget variance”; the disposition remains visible so it cannot be averaged away.
  • Q12 was recorded as correct. The evidence note is “source hierarchy”; the disposition remains visible so it cannot be averaged away.
  • Q13 was recorded as correct. The evidence note is “mortgage escrow”; the disposition remains visible so it cannot be averaged away.
  • Q14 was recorded as partial. The evidence note is “variable-rate limitation”; the disposition remains visible so it cannot be averaged away.
  • Q15 was recorded as correct. The evidence note is “annual-report source”; the disposition remains visible so it cannot be averaged away.
  • Q16 was recorded as correct. The evidence note is “guaranteed-return red flag”; the disposition remains visible so it cannot be averaged away.
  • Q17 was recorded as correct. The evidence note is “forecast uncertainty”; the disposition remains visible so it cannot be averaged away.
  • Q18 was recorded as correct. The evidence note is “data minimization”; the disposition remains visible so it cannot be averaged away.

The expected label or control was fixed before review. The visible note explains why the row received its disposition. The chart uses the published primary criterion, but the table is authoritative because it preserves exceptions that a single percentage would hide.

How to reproduce the check

  1. Download the CSV and read its labels, units, and synthetic/public-data notice before using it.
  2. Write the expected answers or decision rule before looking at a model response.
  3. Use the same bounded prompt and record the model or tool, access surface, and date.
  4. Preserve the raw response. Break prose into atomic claims rather than grading the tone of the whole answer.
  5. Recompute arithmetic with deterministic formulas and verify definitions against the linked primary sources.
  6. Record correct, partial, wrong, uncertain, and refused outcomes separately. Do not silently repair the model output before scoring it.
  7. Repeat material checks after a model, source, or workflow changes.

What the result means

The value of this result is diagnostic. It shows where a structured assistant can reduce search, formatting, or first-pass review work. It does not transfer responsibility for the underlying decision. A “pass” means the row met the published rule in this test, on this date, with these inputs.

The errors and open items matter more than a polished average. In money and business workflows, one missed assumption, stale fact, false match, or overconfident definition can dominate many correct low-risk rows. That is why the artifact keeps row-level outcomes and why a human reviews exceptions rather than receiving only a percentage.

Reproducibility also has limits. A reader can repeat the steps and inspect the same answer key, but a probabilistic model may not return identical wording. A useful rerun should therefore compare atomic claims, calculations, citations, and escalation decisions—not superficial phrasing.

Why this topic needs its own boundary

Tool documentation describes available components, not a validated business process. Release notes can establish that a model or tool exists, while only a controlled test can show whether the surrounding prompt, permissions, data, and approval path work for a defined task.

A useful evaluation therefore treats the model as one changing component. The answer key, source hierarchy, access controls, audit trail, and rollback route belong to the workflow and should remain understandable even when the model name changes.

A safer operating workflow

  • Publish every question and expected answer.
  • Score caution and assumptions, not just the final number.
  • Keep the wrong answers visible.
  • Re-run after material model changes.
  • Do not reuse synthetic scores as proof for personalized advice.

How each control changes the decision

Control 1: Publish every question and expected answer. For this test, that control answers the bounded question “Can a compact benchmark reveal whether an AI system is accurate, cautious, and reproducible on financial questions?” without extending the result into an untested decision.

Control 2: Score caution and assumptions, not just the final number. For this test, that control answers the bounded question “Can a compact benchmark reveal whether an AI system is accurate, cautious, and reproducible on financial questions?” without extending the result into an untested decision.

Control 3: Keep the wrong answers visible. For this test, that control answers the bounded question “Can a compact benchmark reveal whether an AI system is accurate, cautious, and reproducible on financial questions?” without extending the result into an untested decision.

Control 4: Re-run after material model changes. For this test, that control answers the bounded question “Can a compact benchmark reveal whether an AI system is accurate, cautious, and reproducible on financial questions?” without extending the result into an untested decision.

Control 5: Do not reuse synthetic scores as proof for personalized advice. For this test, that control answers the bounded question “Can a compact benchmark reveal whether an AI system is accurate, cautious, and reproducible on financial questions?” without extending the result into an untested decision.

Keep data collection, model preparation, deterministic validation, and approval as separate stages. Use the least sensitive input that can answer the question. If removing personal or confidential data makes the result ambiguous, route the case to an approved human process instead of restoring secrets to an unapproved tool.

Calculations need an independent formula; current facts need a current primary source; classifications need an “uncertain” route; and irreversible actions need explicit authorization outside the model. Logs should capture the version, prompt, source date, output, reviewer, correction, and final disposition without retaining unnecessary personal data.

Limitations and professional boundary

One session is not a universal model ranking. Outputs can vary with prompt, model snapshot, language, access tier, and sampling settings.

This publication provides general educational information. It does not know a reader’s finances, duties, jurisdiction, contracts, tax treatment, credit position, or risk tolerance. A qualified financial, accounting, tax, legal, lending, security, or other professional should review decisions with material consequences.

Primary sources

Verified 2026-07-26. Primary-source links can change; use the publication date and linked source to check for a newer version.

Bottom line

The session scored 25 correct, four partial, and one wrong. A headline accuracy number alone would hide the real-return error and four missing-assumption cases. The practical lesson is to make AI produce inspectable work inside a controlled process—not to make fluency the final control.