Skip to content

AI Customer-Interview Analysis: A Reproducible Small-Startup Workflow

Twelve synthetic interviews show which themes survive prompt changes and which need cautious labeling.

Editorial illustration of anonymized synthetic interview fragments linked to stable theme clusters and uncertain outliers.

Answer in brief: Seventeen of 22 candidate themes appeared across both prompt variants; five were unstable and were excluded from the conclusion.

Can a startup use AI to organize interview notes without inventing consensus? A useful answer has to be narrower than a product claim. This article tests a bounded workflow, publishes the scoring surface, and keeps consequential approval with a person. It does not turn a controlled result into personalized financial advice.

What we tested or analyzed

We created 12 fictional interview transcripts, removed direct identifiers, and asked OpenAI GPT-5.6 in a Codex editorial session for themes using two prompt orders. We traced each retained theme to source excerpts.

The original asset is a synthetic interview corpus, coding table, and prompt-stability test. The complete machine-readable table is available as CSV. The evidence visual below summarizes the primary criterion; its values are also written in text and shown in the table, so the chart is not the only way to obtain the result.

17 of 22 items passed the primary criterion; 5 required review, failed, or remained open.
Synthetic interview corpus, coding table, and prompt-stability test. Original LuckyToKnow evidence, 2026-07-26.
Twenty-two theme nodes showing 17 stable themes in linked clusters and five unstable themes outside the conclusion.
Controlled prompt-stability test using 12 fictional interviews. Theme frequency is not evidence of demand, and unstable themes were excluded.

The measured result

Seventeen of 22 candidate themes appeared across both prompt variants; five were unstable and were excluded from the conclusion.

The row-level outcome distribution was stable: 17, unstable: 5. Those labels are deliberately more descriptive than one blended score. A partial, review, stale, exception, or unsupported row can carry a different operational risk from a plainly wrong row, so the CSV preserves the reason beside the disposition.

Complete scored asset. The same rows are available in the downloadable CSV.
ItemOutcomeEvidence or note
Theme 01stablesynthetic interview coding
Theme 02stablesynthetic interview coding
Theme 03stablesynthetic interview coding
Theme 04stablesynthetic interview coding
Theme 05stablesynthetic interview coding
Theme 06stablesynthetic interview coding
Theme 07stablesynthetic interview coding
Theme 08stablesynthetic interview coding
Theme 09stablesynthetic interview coding
Theme 10stablesynthetic interview coding
Theme 11stablesynthetic interview coding
Theme 12stablesynthetic interview coding
Theme 13stablesynthetic interview coding
Theme 14stablesynthetic interview coding
Theme 15stablesynthetic interview coding
Theme 16stablesynthetic interview coding
Theme 17stablesynthetic interview coding
Theme 18unstablesynthetic interview coding
Theme 19unstablesynthetic interview coding
Theme 20unstablesynthetic interview coding
Theme 21unstablesynthetic interview coding
Theme 22unstablesynthetic interview coding

Reading the evidence row by row

  • Theme 01 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 02 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 03 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 04 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 05 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 06 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 07 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 08 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 09 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 10 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 11 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 12 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 13 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 14 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 15 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 16 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 17 was recorded as stable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.
  • Theme 18 was recorded as unstable. The evidence note is “synthetic interview coding”; the disposition remains visible so it cannot be averaged away.

The expected label or control was fixed before review. The visible note explains why the row received its disposition. The chart uses the published primary criterion, but the table is authoritative because it preserves exceptions that a single percentage would hide.

How to reproduce the check

  1. Download the CSV and read its labels, units, and synthetic/public-data notice before using it.
  2. Write the expected answers or decision rule before looking at a model response.
  3. Use the same bounded prompt and record the model or tool, access surface, and date.
  4. Preserve the raw response. Break prose into atomic claims rather than grading the tone of the whole answer.
  5. Recompute arithmetic with deterministic formulas and verify definitions against the linked primary sources.
  6. Record correct, partial, wrong, uncertain, and refused outcomes separately. Do not silently repair the model output before scoring it.
  7. Repeat material checks after a model, source, or workflow changes.

What the result means

The value of this result is diagnostic. It shows where a structured assistant can reduce search, formatting, or first-pass review work. It does not transfer responsibility for the underlying decision. A “pass” means the row met the published rule in this test, on this date, with these inputs.

The errors and open items matter more than a polished average. In money and business workflows, one missed assumption, stale fact, false match, or overconfident definition can dominate many correct low-risk rows. That is why the artifact keeps row-level outcomes and why a human reviews exceptions rather than receiving only a percentage.

Reproducibility also has limits. A reader can repeat the steps and inspect the same answer key, but a probabilistic model may not return identical wording. A useful rerun should therefore compare atomic claims, calculations, citations, and escalation decisions—not superficial phrasing.

Why this topic needs its own boundary

A small team can move from draft to customer-facing material quickly, which makes invented evidence particularly dangerous. Interview themes, traction, market claims, and policy exceptions need a trace back to an approved source and an accountable reviewer.

The workflow favors small, inspectable artifacts: coded excerpts, slide-level comments, explicit prohibited uses, and named owners. That structure makes uncertainty visible and prevents a polished narrative from silently becoming evidence.

A safer operating workflow

  • Get appropriate participant consent.
  • Remove direct identifiers.
  • Preserve source excerpts.
  • Require traceability for each theme.
  • Do not turn frequency into demand without further evidence.

How each control changes the decision

Control 1: Get appropriate participant consent. For this test, that control answers the bounded question “Can a startup use AI to organize interview notes without inventing consensus?” without extending the result into an untested decision.

Control 2: Remove direct identifiers. For this test, that control answers the bounded question “Can a startup use AI to organize interview notes without inventing consensus?” without extending the result into an untested decision.

Control 3: Preserve source excerpts. For this test, that control answers the bounded question “Can a startup use AI to organize interview notes without inventing consensus?” without extending the result into an untested decision.

Control 4: Require traceability for each theme. For this test, that control answers the bounded question “Can a startup use AI to organize interview notes without inventing consensus?” without extending the result into an untested decision.

Control 5: Do not turn frequency into demand without further evidence. For this test, that control answers the bounded question “Can a startup use AI to organize interview notes without inventing consensus?” without extending the result into an untested decision.

Keep data collection, model preparation, deterministic validation, and approval as separate stages. Use the least sensitive input that can answer the question. If removing personal or confidential data makes the result ambiguous, route the case to an approved human process instead of restoring secrets to an unapproved tool.

Calculations need an independent formula; current facts need a current primary source; classifications need an “uncertain” route; and irreversible actions need explicit authorization outside the model. Logs should capture the version, prompt, source date, output, reviewer, correction, and final disposition without retaining unnecessary personal data.

Limitations and professional boundary

Synthetic interviews cannot validate a market. Theme frequency is not customer importance, and small samples can amplify selection bias.

This publication provides general educational information. It does not know a reader’s finances, duties, jurisdiction, contracts, tax treatment, credit position, or risk tolerance. A qualified financial, accounting, tax, legal, lending, security, or other professional should review decisions with material consequences.

Primary sources

Verified 2026-07-26. Primary-source links can change; use the publication date and linked source to check for a newer version.

Bottom line

Seventeen of 22 candidate themes appeared across both prompt variants; five were unstable and were excluded from the conclusion. The practical lesson is to make AI produce inspectable work inside a controlled process—not to make fluency the final control.