Skip to content

Can AI Spot a Financial Scam Message? A Controlled Test

Forty synthetic messages expose false negatives, false positives, and the danger of treating a classifier as a guarantee.

Editorial illustration of abstract message slips inspected under a forensic light with missed and review cases visible.

Answer in brief: The model produced 18 true positives, 17 true negatives, two false positives, and three false negatives: 87.5% accuracy, 85.7% recall, and 90% precision.

Can AI reliably flag common scam patterns without blocking every urgent legitimate message? A useful answer has to be narrower than a product claim. This article tests a bounded workflow, publishes the scoring surface, and keeps consequential approval with a person. It does not turn a controlled result into personalized financial advice.

What we tested or analyzed

We evaluated OpenAI GPT-5.6 in a Codex editorial session on 40 synthetic messages built from FTC-documented patterns. Labels were fixed first; no live victim messages, links, phone numbers, or personal data were used.

The original asset is a forty-message corpus, confusion matrix, and error analysis. The complete machine-readable table is available as CSV. The evidence visual below summarizes the primary criterion; its values are also written in text and shown in the table, so the chart is not the only way to obtain the result.

35 of 40 items passed the primary criterion; 5 required review, failed, or remained open.
Forty-message corpus, confusion matrix, and error analysis. Original LuckyToKnow evidence, 2026-07-26.
Confusion matrix showing 18 scam messages correctly flagged, two scams missed, 17 legitimate messages cleared, and three legitimate messages falsely flagged.
Controlled test using 40 synthetic messages. False negatives and false positives are separated because neither result makes a classifier a guarantee.

The measured result

The model produced 18 true positives, 17 true negatives, two false positives, and three false negatives: 87.5% accuracy, 85.7% recall, and 90% precision.

The row-level outcome distribution was legitimate: 19, scam: 21. Those labels are deliberately more descriptive than one blended score. A partial, review, stale, exception, or unsupported row can carry a different operational risk from a plainly wrong row, so the CSV preserves the reason beside the disposition.

Complete scored asset. The same rows are available in the downloadable CSV.
ItemOutcomeEvidence or note
M01scamscam
M02scamscam
M03scamscam
M04scamscam
M05scamscam
M06scamscam
M07scamscam
M08scamscam
M09scamscam
M10scamscam
M11scamscam
M12scamscam
M13scamscam
M14scamscam
M15scamscam
M16scamscam
M17scamscam
M18scamscam
M19scamlegitimate
M20scamlegitimate
M21scamlegitimate
M22legitimatescam
M23legitimatescam
M24legitimatelegitimate
M25legitimatelegitimate
M26legitimatelegitimate
M27legitimatelegitimate
M28legitimatelegitimate
M29legitimatelegitimate
M30legitimatelegitimate
M31legitimatelegitimate
M32legitimatelegitimate
M33legitimatelegitimate
M34legitimatelegitimate
M35legitimatelegitimate
M36legitimatelegitimate
M37legitimatelegitimate
M38legitimatelegitimate
M39legitimatelegitimate
M40legitimatelegitimate

Reading the evidence row by row

  • M01 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M02 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M03 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M04 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M05 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M06 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M07 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M08 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M09 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M10 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M11 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M12 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M13 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M14 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M15 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M16 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M17 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.
  • M18 was recorded as scam. The evidence note is “scam”; the disposition remains visible so it cannot be averaged away.

The expected label or control was fixed before review. The visible note explains why the row received its disposition. The chart uses the published primary criterion, but the table is authoritative because it preserves exceptions that a single percentage would hide.

How to reproduce the check

  1. Download the CSV and read its labels, units, and synthetic/public-data notice before using it.
  2. Write the expected answers or decision rule before looking at a model response.
  3. Use the same bounded prompt and record the model or tool, access surface, and date.
  4. Preserve the raw response. Break prose into atomic claims rather than grading the tone of the whole answer.
  5. Recompute arithmetic with deterministic formulas and verify definitions against the linked primary sources.
  6. Record correct, partial, wrong, uncertain, and refused outcomes separately. Do not silently repair the model output before scoring it.
  7. Repeat material checks after a model, source, or workflow changes.

What the result means

The value of this result is diagnostic. It shows where a structured assistant can reduce search, formatting, or first-pass review work. It does not transfer responsibility for the underlying decision. A “pass” means the row met the published rule in this test, on this date, with these inputs.

The errors and open items matter more than a polished average. In money and business workflows, one missed assumption, stale fact, false match, or overconfident definition can dominate many correct low-risk rows. That is why the artifact keeps row-level outcomes and why a human reviews exceptions rather than receiving only a percentage.

Reproducibility also has limits. A reader can repeat the steps and inspect the same answer key, but a probabilistic model may not return identical wording. A useful rerun should therefore compare atomic claims, calculations, citations, and escalation decisions—not superficial phrasing.

Why this topic needs its own boundary

Personal-finance data are unusually revealing: merchant strings, payment timing, balances, and repeated amounts can expose identity and behavior even when an obvious account number is removed. A privacy-first test asks whether the task can be answered with synthetic, aggregated, or minimized inputs.

The safe outcome is an organized draft for review, not a decision about a real household. Taxes, debt, dependants, currency exposure, and emergency needs are contextual facts that a small synthetic benchmark cannot know.

A safer operating workflow

  • Do not click or call using message-supplied details.
  • Verify through an independently found official channel.
  • Treat urgency, unusual payment methods, and secrecy as risk signals.
  • Use a second control when money or credentials are requested.
  • Report suspected fraud to the relevant authority.

How each control changes the decision

Control 1: Do not click or call using message-supplied details. For this test, that control answers the bounded question “Can AI reliably flag common scam patterns without blocking every urgent legitimate message?” without extending the result into an untested decision.

Control 2: Verify through an independently found official channel. For this test, that control answers the bounded question “Can AI reliably flag common scam patterns without blocking every urgent legitimate message?” without extending the result into an untested decision.

Control 3: Treat urgency, unusual payment methods, and secrecy as risk signals. For this test, that control answers the bounded question “Can AI reliably flag common scam patterns without blocking every urgent legitimate message?” without extending the result into an untested decision.

Control 4: Use a second control when money or credentials are requested. For this test, that control answers the bounded question “Can AI reliably flag common scam patterns without blocking every urgent legitimate message?” without extending the result into an untested decision.

Control 5: Report suspected fraud to the relevant authority. For this test, that control answers the bounded question “Can AI reliably flag common scam patterns without blocking every urgent legitimate message?” without extending the result into an untested decision.

Keep data collection, model preparation, deterministic validation, and approval as separate stages. Use the least sensitive input that can answer the question. If removing personal or confidential data makes the result ambiguous, route the case to an approved human process instead of restoring secrets to an unapproved tool.

Calculations need an independent formula; current facts need a current primary source; classifications need an “uncertain” route; and irreversible actions need explicit authorization outside the model. Logs should capture the version, prompt, source date, output, reviewer, correction, and final disposition without retaining unnecessary personal data.

Limitations and professional boundary

Scams adapt, language and cultural context matter, and a missed scam can cause serious harm. This is not a security product evaluation.

This publication provides general educational information. It does not know a reader’s finances, duties, jurisdiction, contracts, tax treatment, credit position, or risk tolerance. A qualified financial, accounting, tax, legal, lending, security, or other professional should review decisions with material consequences.

Primary sources

Verified 2026-07-26. Primary-source links can change; use the publication date and linked source to check for a newer version.

Bottom line

The model produced 18 true positives, 17 true negatives, two false positives, and three false negatives: 87.5% accuracy, 85.7% recall, and 90% precision. The practical lesson is to make AI produce inspectable work inside a controlled process—not to make fluency the final control.