Skip to content

How We Test AI

LuckyToKnow tests are designed to answer a bounded question—not prove that an AI tool is universally good or bad.

Before a test

  1. Define the reader question, success criteria, and foreseeable harm.
  2. Record the exact tool, model or version, access tier, and test date.
  3. Create controlled public or synthetic inputs with no personal or confidential data.
  4. Choose an authoritative answer key or a reproducible calculation.
  5. Write the scoring rule before seeing the result.

During and after a test

  1. Preserve prompts, outputs, datasets, and scripts needed to reproduce the claim.
  2. Check every number independently and flag ambiguous cases.
  3. Report errors, refusals, inconsistency, and unsupported certainty.
  4. Separate the measured result from editorial interpretation.
  5. State limits: sample size, timing, model variability, jurisdiction, and data quality.

Two independent review passes

The evidence review checks sources, calculations, reproducibility, and unsupported inference. The adversarial editorial review looks for copied phrasing, thin value, unsafe advice, misleading headlines, vague disclosure, and missing limitations. A blocking finding returns the article to draft.