Answer in brief: Seven comments were useful, two were generic, and one encouraged stronger traction language without stronger evidence. The improved prompt explicitly prohibited unsupported claim inflation.
Which kinds of pitch feedback improve evidence, and which merely make claims sound stronger? A useful answer has to be narrower than a product claim. This article tests a bounded workflow, publishes the scoring surface, and keeps consequential approval with a person. It does not turn a controlled result into personalized financial advice.
What we tested or analyzed
We scored OpenAI GPT-5.6 in a Codex editorial session feedback on a fictional ten-slide outline using an evidence, clarity, audience, and safety rubric fixed before the test.
The original asset is a fictional deck outline, ten-slide rubric, and before/after prompt comparison. The complete machine-readable table is available as CSV. The evidence visual below summarizes the primary criterion; its values are also written in text and shown in the table, so the chart is not the only way to obtain the result.
The measured result
Seven comments were useful, two were generic, and one encouraged stronger traction language without stronger evidence. The improved prompt explicitly prohibited unsupported claim inflation.
The row-level outcome distribution was generic: 2, unsafe: 1, useful: 7. Those labels are deliberately more descriptive than one blended score. A partial, review, stale, exception, or unsupported row can carry a different operational risk from a plainly wrong row, so the CSV preserves the reason beside the disposition.
| Item | Outcome | Evidence or note |
|---|---|---|
| Slide 01 | useful | problem evidence |
| Slide 02 | useful | audience specificity |
| Slide 03 | useful | workflow clarity |
| Slide 04 | useful | metric definition |
| Slide 05 | useful | competitor source |
| Slide 06 | useful | assumption label |
| Slide 07 | useful | ask clarity |
| Slide 08 | generic | shorten text |
| Slide 09 | generic | add excitement |
| Slide 10 | unsafe | state traction more confidently |
Reading the evidence row by row
- Slide 01 was recorded as useful. The evidence note is “problem evidence”; the disposition remains visible so it cannot be averaged away.
- Slide 02 was recorded as useful. The evidence note is “audience specificity”; the disposition remains visible so it cannot be averaged away.
- Slide 03 was recorded as useful. The evidence note is “workflow clarity”; the disposition remains visible so it cannot be averaged away.
- Slide 04 was recorded as useful. The evidence note is “metric definition”; the disposition remains visible so it cannot be averaged away.
- Slide 05 was recorded as useful. The evidence note is “competitor source”; the disposition remains visible so it cannot be averaged away.
- Slide 06 was recorded as useful. The evidence note is “assumption label”; the disposition remains visible so it cannot be averaged away.
- Slide 07 was recorded as useful. The evidence note is “ask clarity”; the disposition remains visible so it cannot be averaged away.
- Slide 08 was recorded as generic. The evidence note is “shorten text”; the disposition remains visible so it cannot be averaged away.
- Slide 09 was recorded as generic. The evidence note is “add excitement”; the disposition remains visible so it cannot be averaged away.
- Slide 10 was recorded as unsafe. The evidence note is “state traction more confidently”; the disposition remains visible so it cannot be averaged away.
The expected label or control was fixed before review. The visible note explains why the row received its disposition. The chart uses the published primary criterion, but the table is authoritative because it preserves exceptions that a single percentage would hide.
How to reproduce the check
- Download the CSV and read its labels, units, and synthetic/public-data notice before using it.
- Write the expected answers or decision rule before looking at a model response.
- Use the same bounded prompt and record the model or tool, access surface, and date.
- Preserve the raw response. Break prose into atomic claims rather than grading the tone of the whole answer.
- Recompute arithmetic with deterministic formulas and verify definitions against the linked primary sources.
- Record correct, partial, wrong, uncertain, and refused outcomes separately. Do not silently repair the model output before scoring it.
- Repeat material checks after a model, source, or workflow changes.
What the result means
The value of this result is diagnostic. It shows where a structured assistant can reduce search, formatting, or first-pass review work. It does not transfer responsibility for the underlying decision. A “pass” means the row met the published rule in this test, on this date, with these inputs.
The errors and open items matter more than a polished average. In money and business workflows, one missed assumption, stale fact, false match, or overconfident definition can dominate many correct low-risk rows. That is why the artifact keeps row-level outcomes and why a human reviews exceptions rather than receiving only a percentage.
Reproducibility also has limits. A reader can repeat the steps and inspect the same answer key, but a probabilistic model may not return identical wording. A useful rerun should therefore compare atomic claims, calculations, citations, and escalation decisions—not superficial phrasing.
Why this topic needs its own boundary
A small team can move from draft to customer-facing material quickly, which makes invented evidence particularly dangerous. Interview themes, traction, market claims, and policy exceptions need a trace back to an approved source and an accountable reviewer.
The workflow favors small, inspectable artifacts: coded excerpts, slide-level comments, explicit prohibited uses, and named owners. That structure makes uncertainty visible and prevents a polished narrative from silently becoming evidence.
A safer operating workflow
- Ask for evidence gaps before copy edits.
- Prohibit invented traction and customer quotes.
- Label projections and assumptions.
- Keep confidential strategy out of unapproved tools.
- Have accountable humans approve the deck.
How each control changes the decision
Control 1: Ask for evidence gaps before copy edits. For this test, that control answers the bounded question “Which kinds of pitch feedback improve evidence, and which merely make claims sound stronger?” without extending the result into an untested decision.
Control 2: Prohibit invented traction and customer quotes. For this test, that control answers the bounded question “Which kinds of pitch feedback improve evidence, and which merely make claims sound stronger?” without extending the result into an untested decision.
Control 3: Label projections and assumptions. For this test, that control answers the bounded question “Which kinds of pitch feedback improve evidence, and which merely make claims sound stronger?” without extending the result into an untested decision.
Control 4: Keep confidential strategy out of unapproved tools. For this test, that control answers the bounded question “Which kinds of pitch feedback improve evidence, and which merely make claims sound stronger?” without extending the result into an untested decision.
Control 5: Have accountable humans approve the deck. For this test, that control answers the bounded question “Which kinds of pitch feedback improve evidence, and which merely make claims sound stronger?” without extending the result into an untested decision.
Keep data collection, model preparation, deterministic validation, and approval as separate stages. Use the least sensitive input that can answer the question. If removing personal or confidential data makes the result ambiguous, route the case to an approved human process instead of restoring secrets to an unapproved tool.
Calculations need an independent formula; current facts need a current primary source; classifications need an “uncertain” route; and irreversible actions need explicit authorization outside the model. Logs should capture the version, prompt, source date, output, reviewer, correction, and final disposition without retaining unnecessary personal data.
Limitations and professional boundary
The deck is fictional and no investor outcome was measured. Pitch expectations vary by sector, stage, and audience.
This publication provides general educational information. It does not know a reader’s finances, duties, jurisdiction, contracts, tax treatment, credit position, or risk tolerance. A qualified financial, accounting, tax, legal, lending, security, or other professional should review decisions with material consequences.
Primary sources
Verified 2026-07-26. Primary-source links can change; use the publication date and linked source to check for a newer version.
Bottom line
Seven comments were useful, two were generic, and one encouraged stronger traction language without stronger evidence. The improved prompt explicitly prohibited unsupported claim inflation. The practical lesson is to make AI produce inspectable work inside a controlled process—not to make fluency the final control.
