If a client has ever pasted a pay stub or a bank statement into ChatGPT and asked “does this look okay for my mortgage,” a new study says the answer it gave back had roughly a one in four chance of being wrong. Researchers from Columbia University, working with AI workflow startup Tidalwave, tested general-purpose AI models against a purpose-built mortgage tool on 90 questions drawn from 10 synthetic borrower scenarios. The ai chatbot mortgage document errors real estate agents the study surfaced are not a minor footnote. They are exactly the kind of mistake that can sink a deal at the worst possible moment.

What the Columbia and Tidalwave Study Actually Tested
According to Yahoo Finance’s coverage, the test scenarios reflected real problems loan officers face during origination: payroll mismatches, undisclosed liabilities, and suspicious deposits. Anthropic’s Claude 4.5 scored 71% overall accuracy, dropping to just 42% on straightforward yes-or-no compliance checks. Tidalwave’s purpose-built tool, SOLO, scored 84% overall and 95% on those same compliance checks. General AI models struggled with tasks as basic as matching account IDs across documents, spotting a large unexplained deposit on a bank statement, or confirming an employer name matched between a pay stub and a direct deposit record.

The Bias Finding Agents Cannot Ignore
Beyond raw accuracy, the study flagged something more serious: researchers found AI models making judgments about whether a borrower was foreign-born based on how their name sounded. Columbia doctoral student Matthew Toles, who worked on the research, called that approach to assessing an applicant’s background exactly the kind of pattern that turns into fair-lending liability fast. A chatbot quietly inferring an applicant’s national origin from a name and letting that color its output is not a hypothetical risk. It is the kind of thing a regulator or a plaintiff’s attorney would flag immediately if it ever showed up in a real file.
Where AI Chatbot Mortgage Document Errors Real Estate Agents Should Watch For
Agents are not underwriters, but plenty of agents and their assistants use general AI chatbots informally to help a buyer sanity-check paperwork before it goes to a lender. Three practices are worth adopting after this study.
First, treat a general-purpose chatbot’s read on a financial document as a rough first pass, not a verification. If a buyer asks whether their bank statement looks fine before submission, a “looks okay” from ChatGPT is not the same as an actual compliance check.
Second, route anything document-specific to a tool built for the job, or to the lender directly, rather than a general assistant. The accuracy gap in this study between a general model and a purpose-built one was not small, it was more than 40 points on compliance-style questions.
Third, watch for AI-generated judgments that touch protected characteristics, even indirectly. A tool inferring anything about national origin, ethnicity, or immigration status from a name is a fair housing and fair lending red flag regardless of whether a human ever explicitly asked it to.

Not All AI Is Built for the Same Job
The gap between Claude’s 71% and SOLO’s 84% is the real lesson here. A general-purpose chatbot trained to be broadly helpful is not the same category of tool as one trained specifically on mortgage documents and compliance rules. That distinction applies well beyond mortgages, and it is the direct explanation for why ai chatbot mortgage document errors real estate agents showed up so consistently in the study. Agents evaluating any AI tool for a document-heavy, compliance-sensitive task should ask what the tool was actually built and tested for, not just whether it can produce a confident-sounding answer.

Related Reading
If you are evaluating which AI assistant your team should actually trust with client-facing tasks, our guide to AI virtual assistants for real estate breaks down which platforms are built for narrow, verifiable tasks versus general conversation. It pairs well with our coverage of AI hallucination risk in real estate listings, since both studies point to the same underlying issue: confident-sounding AI output is not the same as accurate AI output.
Final Thoughts
A one-in-four error rate on tasks as basic as matching an account ID is not a reason to avoid AI in real estate transactions. It is a reason to be precise about which AI is doing which job. The ai chatbot mortgage document errors real estate agents this study documented happened with some of the most capable general models available. The fix is not less AI, it is pointing the right AI at the right task, and keeping a human in the loop for anything that touches a client’s actual financing.
