Skip to content

Can AI Summarize an Annual Report Accurately?

A claim-by-claim audit of a public Form 10-K summary shows why citations need section-level verification.

Editorial illustration of summary cards connected by citation threads to exact sections of an unbranded annual report.

Answer in brief: Nine claims were supported, two omitted material context, and one turned a conditional management statement into an unsupported certainty.

Can an AI summary preserve the meaning of a Form 10-K without turning management language into fact? A useful answer has to be narrower than a product claim. This article tests a bounded workflow, publishes the scoring surface, and keeps consequential approval with a person. It does not turn a controlled result into personalized financial advice.

What we tested or analyzed

We used OpenAI GPT-5.6 in a Codex editorial session to summarize a public SEC filing structure, then checked 12 atomic claims against the filing sections and SEC guidance. The companion table preserves the disposition.

The original asset is a twelve-claim annual-report verification table. The complete machine-readable table is available as CSV. The evidence visual below summarizes the primary criterion; its values are also written in text and shown in the table, so the chart is not the only way to obtain the result.

9 of 12 items passed the primary criterion; 3 required review, failed, or remained open.
Twelve-claim annual-report verification table. Original LuckyToKnow evidence, 2026-07-26.
Twelve claim cards grouped as nine supported, two incomplete, and one unsupported across annual-report sections.
Controlled claim audit of a public Form 10-K. Outcomes reflect the published test; the diagram is not a company performance chart.

The measured result

Nine claims were supported, two omitted material context, and one turned a conditional management statement into an unsupported certainty.

The row-level outcome distribution was incomplete: 2, supported: 9, unsupported: 1. Those labels are deliberately more descriptive than one blended score. A partial, review, stale, exception, or unsupported row can carry a different operational risk from a plainly wrong row, so the CSV preserves the reason beside the disposition.

Complete scored asset. The same rows are available in the downloadable CSV.
ItemOutcomeEvidence or note
Claim 01supportedBusiness
Claim 02supportedRisk Factors
Claim 03supportedMD&A
Claim 04supportedFinancial Statements
Claim 05supportedCash flows
Claim 06supportedSegment note
Claim 07supportedDebt note
Claim 08supportedControls
Claim 09supportedAuditor report
Claim 10incompleteRisk ordering
Claim 11incompleteNon-GAAP context
Claim 12unsupportedForward growth certainty

Reading the evidence row by row

  • Claim 01 was recorded as supported. The evidence note is “Business”; the disposition remains visible so it cannot be averaged away.
  • Claim 02 was recorded as supported. The evidence note is “Risk Factors”; the disposition remains visible so it cannot be averaged away.
  • Claim 03 was recorded as supported. The evidence note is “MD&A”; the disposition remains visible so it cannot be averaged away.
  • Claim 04 was recorded as supported. The evidence note is “Financial Statements”; the disposition remains visible so it cannot be averaged away.
  • Claim 05 was recorded as supported. The evidence note is “Cash flows”; the disposition remains visible so it cannot be averaged away.
  • Claim 06 was recorded as supported. The evidence note is “Segment note”; the disposition remains visible so it cannot be averaged away.
  • Claim 07 was recorded as supported. The evidence note is “Debt note”; the disposition remains visible so it cannot be averaged away.
  • Claim 08 was recorded as supported. The evidence note is “Controls”; the disposition remains visible so it cannot be averaged away.
  • Claim 09 was recorded as supported. The evidence note is “Auditor report”; the disposition remains visible so it cannot be averaged away.
  • Claim 10 was recorded as incomplete. The evidence note is “Risk ordering”; the disposition remains visible so it cannot be averaged away.
  • Claim 11 was recorded as incomplete. The evidence note is “Non-GAAP context”; the disposition remains visible so it cannot be averaged away.
  • Claim 12 was recorded as unsupported. The evidence note is “Forward growth certainty”; the disposition remains visible so it cannot be averaged away.

The expected label or control was fixed before review. The visible note explains why the row received its disposition. The chart uses the published primary criterion, but the table is authoritative because it preserves exceptions that a single percentage would hide.

How to reproduce the check

  1. Download the CSV and read its labels, units, and synthetic/public-data notice before using it.
  2. Write the expected answers or decision rule before looking at a model response.
  3. Use the same bounded prompt and record the model or tool, access surface, and date.
  4. Preserve the raw response. Break prose into atomic claims rather than grading the tone of the whole answer.
  5. Recompute arithmetic with deterministic formulas and verify definitions against the linked primary sources.
  6. Record correct, partial, wrong, uncertain, and refused outcomes separately. Do not silently repair the model output before scoring it.
  7. Repeat material checks after a model, source, or workflow changes.

What the result means

The value of this result is diagnostic. It shows where a structured assistant can reduce search, formatting, or first-pass review work. It does not transfer responsibility for the underlying decision. A “pass” means the row met the published rule in this test, on this date, with these inputs.

The errors and open items matter more than a polished average. In money and business workflows, one missed assumption, stale fact, false match, or overconfident definition can dominate many correct low-risk rows. That is why the artifact keeps row-level outcomes and why a human reviews exceptions rather than receiving only a percentage.

Reproducibility also has limits. A reader can repeat the steps and inspect the same answer key, but a probabilistic model may not return identical wording. A useful rerun should therefore compare atomic claims, calculations, citations, and escalation decisions—not superficial phrasing.

Why this topic needs its own boundary

Market research becomes unreliable when dates disappear. A reported fact, a management forecast, a market price, and an analyst inference have different sources and time horizons. The workflow keeps those evidence types separate and rejects persuasive language as a substitute for a filing or time stamp.

None of the tests estimates intrinsic value or recommends a security. Their purpose is claim hygiene: finding where a statement came from, what period it describes, what could contradict it, and whether the statement is supportable at all.

A safer operating workflow

  • Use the filed document, not a copied summary.
  • Break prose into atomic claims.
  • Link every number to a statement or note.
  • Mark management opinion as opinion.
  • Read risk factors and footnotes before drawing conclusions.

How each control changes the decision

Control 1: Use the filed document, not a copied summary. For this test, that control answers the bounded question “Can an AI summary preserve the meaning of a Form 10-K without turning management language into fact?” without extending the result into an untested decision.

Control 2: Break prose into atomic claims. For this test, that control answers the bounded question “Can an AI summary preserve the meaning of a Form 10-K without turning management language into fact?” without extending the result into an untested decision.

Control 3: Link every number to a statement or note. For this test, that control answers the bounded question “Can an AI summary preserve the meaning of a Form 10-K without turning management language into fact?” without extending the result into an untested decision.

Control 4: Mark management opinion as opinion. For this test, that control answers the bounded question “Can an AI summary preserve the meaning of a Form 10-K without turning management language into fact?” without extending the result into an untested decision.

Control 5: Read risk factors and footnotes before drawing conclusions. For this test, that control answers the bounded question “Can an AI summary preserve the meaning of a Form 10-K without turning management language into fact?” without extending the result into an untested decision.

Keep data collection, model preparation, deterministic validation, and approval as separate stages. Use the least sensitive input that can answer the question. If removing personal or confidential data makes the result ambiguous, route the case to an approved human process instead of restoring secrets to an unapproved tool.

Calculations need an independent formula; current facts need a current primary source; classifications need an “uncertain” route; and irreversible actions need explicit authorization outside the model. Logs should capture the version, prompt, source date, output, reviewer, correction, and final disposition without retaining unnecessary personal data.

Limitations and professional boundary

This test is not a valuation, audit, or investment recommendation. Filing formats, amendments, footnotes, and company complexity vary.

This publication provides general educational information. It does not know a reader’s finances, duties, jurisdiction, contracts, tax treatment, credit position, or risk tolerance. A qualified financial, accounting, tax, legal, lending, security, or other professional should review decisions with material consequences.

Primary sources

Verified 2026-07-26. Primary-source links can change; use the publication date and linked source to check for a newer version.

Bottom line

Nine claims were supported, two omitted material context, and one turned a conditional management statement into an unsupported certainty. The practical lesson is to make AI produce inspectable work inside a controlled process—not to make fluency the final control.