Skip to content

Can AI Explain Mortgage Terms Correctly? A 25-Term Accuracy Test

Twenty-five definitions checked against CFPB materials expose one material APR error and three incomplete answers.

Editorial illustration of a house blueprint assembled from glossary tiles with incomplete and incorrect pieces visible.

Answer in brief: Twenty-one definitions were correct, three were incomplete, and one incorrectly treated APR as the interest rate. That error is material in a comparison workflow.

Can a general-purpose model explain mortgage vocabulary without collapsing important distinctions? A useful answer has to be narrower than a product claim. This article tests a bounded workflow, publishes the scoring surface, and keeps consequential approval with a person. It does not turn a controlled result into personalized financial advice.

What we tested or analyzed

We prompted OpenAI GPT-5.6 in a Codex editorial session for plain-language definitions and scored them against CFPB explanations written into the answer key before review.

The original asset is a twenty-five-term answer key and scoring table. The complete machine-readable table is available as CSV. The evidence visual below summarizes the primary criterion; its values are also written in text and shown in the table, so the chart is not the only way to obtain the result.

21 of 25 items passed the primary criterion; 4 required review, failed, or remained open.
Twenty-five-term answer key and scoring table. Original LuckyToKnow evidence, 2026-07-26.
Twenty-five-term tile map showing 21 correct definitions, three partial definitions, and one wrong APR definition.
Controlled 25-term test checked against CFPB materials. The highlighted APR error is a published test result, not consumer advice.

The measured result

Twenty-one definitions were correct, three were incomplete, and one incorrectly treated APR as the interest rate. That error is material in a comparison workflow.

The row-level outcome distribution was correct: 21, partial: 3, wrong: 1. Those labels are deliberately more descriptive than one blended score. A partial, review, stale, exception, or unsupported row can carry a different operational risk from a plainly wrong row, so the CSV preserves the reason beside the disposition.

Complete scored asset. The same rows are available in the downloadable CSV.
ItemOutcomeEvidence or note
principalcorrectCFPB check
interest ratecorrectCFPB check
APRwrongCFPB check
pointscorrectCFPB check
escrowcorrectCFPB check
closing costscorrectCFPB check
cash to closecorrectCFPB check
rate lockcorrectCFPB check
fixed ratecorrectCFPB check
adjustable ratepartialCFPB check
balloon paymentcorrectCFPB check
prepayment penaltycorrectCFPB check
negative amortizationpartialCFPB check
mortgage insurancecorrectCFPB check
property taxcorrectCFPB check
homeowners insurancecorrectCFPB check
loan termcorrectCFPB check
amortizationcorrectCFPB check
down paymentcorrectCFPB check
Loan EstimatecorrectCFPB check
Closing DisclosurecorrectCFPB check
TIPpartialCFPB check
servicercorrectCFPB check
appraisalcorrectCFPB check
assumptioncorrectCFPB check

Reading the evidence row by row

  • principal was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • interest rate was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • APR was recorded as wrong. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • points was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • escrow was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • closing costs was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • cash to close was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • rate lock was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • fixed rate was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • adjustable rate was recorded as partial. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • balloon payment was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • prepayment penalty was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • negative amortization was recorded as partial. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • mortgage insurance was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • property tax was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • homeowners insurance was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • loan term was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.
  • amortization was recorded as correct. The evidence note is “CFPB check”; the disposition remains visible so it cannot be averaged away.

The expected label or control was fixed before review. The visible note explains why the row received its disposition. The chart uses the published primary criterion, but the table is authoritative because it preserves exceptions that a single percentage would hide.

How to reproduce the check

  1. Download the CSV and read its labels, units, and synthetic/public-data notice before using it.
  2. Write the expected answers or decision rule before looking at a model response.
  3. Use the same bounded prompt and record the model or tool, access surface, and date.
  4. Preserve the raw response. Break prose into atomic claims rather than grading the tone of the whole answer.
  5. Recompute arithmetic with deterministic formulas and verify definitions against the linked primary sources.
  6. Record correct, partial, wrong, uncertain, and refused outcomes separately. Do not silently repair the model output before scoring it.
  7. Repeat material checks after a model, source, or workflow changes.

What the result means

The value of this result is diagnostic. It shows where a structured assistant can reduce search, formatting, or first-pass review work. It does not transfer responsibility for the underlying decision. A “pass” means the row met the published rule in this test, on this date, with these inputs.

The errors and open items matter more than a polished average. In money and business workflows, one missed assumption, stale fact, false match, or overconfident definition can dominate many correct low-risk rows. That is why the artifact keeps row-level outcomes and why a human reviews exceptions rather than receiving only a percentage.

Reproducibility also has limits. A reader can repeat the steps and inspect the same answer key, but a probabilistic model may not return identical wording. A useful rerun should therefore compare atomic claims, calculations, citations, and escalation decisions—not superficial phrasing.

Why this topic needs its own boundary

Banking and lending decisions can affect access to money, housing, and essential services. A model score or explanation must therefore sit inside data-quality, validation, notice, review, appeal, security, and monitoring controls rather than becoming the decision by itself.

The examples organize public terms and simplified scenarios. They do not evaluate eligibility, affordability, identity, creditworthiness, or a real product, and they do not replace the official disclosure for the reader’s jurisdiction.

A safer operating workflow

  • Check the jurisdiction and official disclosure.
  • Keep APR distinct from the note rate.
  • Ask what can change over time.
  • Read special features and cash-to-close fields.
  • Use a qualified local professional for a real mortgage.

How each control changes the decision

Control 1: Check the jurisdiction and official disclosure. For this test, that control answers the bounded question “Can a general-purpose model explain mortgage vocabulary without collapsing important distinctions?” without extending the result into an untested decision.

Control 2: Keep APR distinct from the note rate. For this test, that control answers the bounded question “Can a general-purpose model explain mortgage vocabulary without collapsing important distinctions?” without extending the result into an untested decision.

Control 3: Ask what can change over time. For this test, that control answers the bounded question “Can a general-purpose model explain mortgage vocabulary without collapsing important distinctions?” without extending the result into an untested decision.

Control 4: Read special features and cash-to-close fields. For this test, that control answers the bounded question “Can a general-purpose model explain mortgage vocabulary without collapsing important distinctions?” without extending the result into an untested decision.

Control 5: Use a qualified local professional for a real mortgage. For this test, that control answers the bounded question “Can a general-purpose model explain mortgage vocabulary without collapsing important distinctions?” without extending the result into an untested decision.

Keep data collection, model preparation, deterministic validation, and approval as separate stages. Use the least sensitive input that can answer the question. If removing personal or confidential data makes the result ambiguous, route the case to an approved human process instead of restoring secrets to an unapproved tool.

Calculations need an independent formula; current facts need a current primary source; classifications need an “uncertain” route; and irreversible actions need explicit authorization outside the model. Logs should capture the version, prompt, source date, output, reviewer, correction, and final disposition without retaining unnecessary personal data.

Limitations and professional boundary

Mortgage terms and disclosures differ across jurisdictions. This U.S.-source test is educational and does not apply automatically in Tunisia or elsewhere.

This publication provides general educational information. It does not know a reader’s finances, duties, jurisdiction, contracts, tax treatment, credit position, or risk tolerance. A qualified financial, accounting, tax, legal, lending, security, or other professional should review decisions with material consequences.

Primary sources

Verified 2026-07-26. Primary-source links can change; use the publication date and linked source to check for a newer version.

Bottom line

Twenty-one definitions were correct, three were incomplete, and one incorrectly treated APR as the interest rate. That error is material in a comparison workflow. The practical lesson is to make AI produce inspectable work inside a controlled process—not to make fluency the final control.