Skip to content

A Privacy-First Way to Categorize Expenses With AI

Forty synthetic transactions compare unredacted and minimized inputs, including the accuracy cost of removing identifiers.

Editorial illustration of fictional receipts passing through a privacy screen before being sorted into categories and review.

Answer in brief: The redacted workflow classified 35 of 40 correctly; the unredacted version classified 38. The three-point gain did not justify exposing names, full merchant references, or account identifiers.

How much categorization accuracy is lost when transaction data are minimized before AI use? A useful answer has to be narrower than a product claim. This article tests a bounded workflow, publishes the scoring surface, and keeps consequential approval with a person. It does not turn a controlled result into personalized financial advice.

What we tested or analyzed

We asked OpenAI GPT-5.6 in a Codex editorial session to label the same 40 synthetic transactions twice: once with descriptive merchant text and once with minimized tokens. Expected labels were written first.

The original asset is a forty synthetic transactions and a redaction/accuracy comparison. The complete machine-readable table is available as CSV. The evidence visual below summarizes the primary criterion; its values are also written in text and shown in the table, so the chart is not the only way to obtain the result.

35 of 40 items passed the primary criterion; 5 required review, failed, or remained open.
Forty synthetic transactions and a redaction/accuracy comparison. Original LuckyToKnow evidence, 2026-07-26.
Side-by-side comparison showing 35 of 40 correct with minimized input and 38 of 40 correct with unredacted input.
Controlled test using 40 synthetic transactions. The comparison reports test accuracy and data exposure, not tax treatment or real account behavior.

The measured result

The redacted workflow classified 35 of 40 correctly; the unredacted version classified 38. The three-point gain did not justify exposing names, full merchant references, or account identifiers.

The row-level outcome distribution was correct: 35, wrong: 5. Those labels are deliberately more descriptive than one blended score. A partial, review, stale, exception, or unsupported row can carry a different operational risk from a plainly wrong row, so the CSV preserves the reason beside the disposition.

Complete scored asset. The same rows are available in the downloadable CSV.
ItemOutcomeEvidence or note
T01correctcorrect
T02correctcorrect
T03correctcorrect
T04correctcorrect
T05correctcorrect
T06correctcorrect
T07correctcorrect
T08correctcorrect
T09correctcorrect
T10correctcorrect
T11correctcorrect
T12correctcorrect
T13correctcorrect
T14correctcorrect
T15correctcorrect
T16correctcorrect
T17correctcorrect
T18correctcorrect
T19correctcorrect
T20correctcorrect
T21correctcorrect
T22correctcorrect
T23correctcorrect
T24correctcorrect
T25correctcorrect
T26correctcorrect
T27correctcorrect
T28correctcorrect
T29correctcorrect
T30correctcorrect
T31correctcorrect
T32correctcorrect
T33correctcorrect
T34correctcorrect
T35correctcorrect
T36wrongcorrect
T37wrongcorrect
T38wrongcorrect
T39wrongwrong
T40wrongwrong

Reading the evidence row by row

  • T01 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T02 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T03 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T04 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T05 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T06 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T07 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T08 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T09 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T10 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T11 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T12 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T13 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T14 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T15 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T16 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T17 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.
  • T18 was recorded as correct. The evidence note is “correct”; the disposition remains visible so it cannot be averaged away.

The expected label or control was fixed before review. The visible note explains why the row received its disposition. The chart uses the published primary criterion, but the table is authoritative because it preserves exceptions that a single percentage would hide.

How to reproduce the check

  1. Download the CSV and read its labels, units, and synthetic/public-data notice before using it.
  2. Write the expected answers or decision rule before looking at a model response.
  3. Use the same bounded prompt and record the model or tool, access surface, and date.
  4. Preserve the raw response. Break prose into atomic claims rather than grading the tone of the whole answer.
  5. Recompute arithmetic with deterministic formulas and verify definitions against the linked primary sources.
  6. Record correct, partial, wrong, uncertain, and refused outcomes separately. Do not silently repair the model output before scoring it.
  7. Repeat material checks after a model, source, or workflow changes.

What the result means

The value of this result is diagnostic. It shows where a structured assistant can reduce search, formatting, or first-pass review work. It does not transfer responsibility for the underlying decision. A “pass” means the row met the published rule in this test, on this date, with these inputs.

The errors and open items matter more than a polished average. In money and business workflows, one missed assumption, stale fact, false match, or overconfident definition can dominate many correct low-risk rows. That is why the artifact keeps row-level outcomes and why a human reviews exceptions rather than receiving only a percentage.

Reproducibility also has limits. A reader can repeat the steps and inspect the same answer key, but a probabilistic model may not return identical wording. A useful rerun should therefore compare atomic claims, calculations, citations, and escalation decisions—not superficial phrasing.

Why this topic needs its own boundary

Personal-finance data are unusually revealing: merchant strings, payment timing, balances, and repeated amounts can expose identity and behavior even when an obvious account number is removed. A privacy-first test asks whether the task can be answered with synthetic, aggregated, or minimized inputs.

The safe outcome is an organized draft for review, not a decision about a real household. Taxes, debt, dependants, currency exposure, and emergency needs are contextual facts that a small synthetic benchmark cannot know.

A safer operating workflow

  • Remove account numbers and personal names.
  • Replace exact merchant text with a bounded description when possible.
  • Keep an “uncertain—review” label.
  • Never upload bank credentials or full statements to an unapproved tool.
  • Let an accountant decide tax treatment.

How each control changes the decision

Control 1: Remove account numbers and personal names. For this test, that control answers the bounded question “How much categorization accuracy is lost when transaction data are minimized before AI use?” without extending the result into an untested decision.

Control 2: Replace exact merchant text with a bounded description when possible. For this test, that control answers the bounded question “How much categorization accuracy is lost when transaction data are minimized before AI use?” without extending the result into an untested decision.

Control 3: Keep an “uncertain—review” label. For this test, that control answers the bounded question “How much categorization accuracy is lost when transaction data are minimized before AI use?” without extending the result into an untested decision.

Control 4: Never upload bank credentials or full statements to an unapproved tool. For this test, that control answers the bounded question “How much categorization accuracy is lost when transaction data are minimized before AI use?” without extending the result into an untested decision.

Control 5: Let an accountant decide tax treatment. For this test, that control answers the bounded question “How much categorization accuracy is lost when transaction data are minimized before AI use?” without extending the result into an untested decision.

Keep data collection, model preparation, deterministic validation, and approval as separate stages. Use the least sensitive input that can answer the question. If removing personal or confidential data makes the result ambiguous, route the case to an approved human process instead of restoring secrets to an unapproved tool.

Calculations need an independent formula; current facts need a current primary source; classifications need an “uncertain” route; and irreversible actions need explicit authorization outside the model. Logs should capture the version, prompt, source date, output, reviewer, correction, and final disposition without retaining unnecessary personal data.

Limitations and professional boundary

Synthetic merchant descriptions are cleaner than real bank feeds. Categories are educational and do not decide tax deductibility.

This publication provides general educational information. It does not know a reader’s finances, duties, jurisdiction, contracts, tax treatment, credit position, or risk tolerance. A qualified financial, accounting, tax, legal, lending, security, or other professional should review decisions with material consequences.

Primary sources

Verified 2026-07-26. Primary-source links can change; use the publication date and linked source to check for a newer version.

Bottom line

The redacted workflow classified 35 of 40 correctly; the unredacted version classified 38. The three-point gain did not justify exposing names, full merchant references, or account identifiers. The practical lesson is to make AI produce inspectable work inside a controlled process—not to make fluency the final control.