Where AI gets taxes wrong
Six patterns account for most of it: stale numbers, invented citations, the missing follow-up question, marginal rate confusion, state law, and agreeing with you. They share one cause. The model has no access to your facts and no stake in the outcome, so it optimises for a good answer to the question you typed.
1. Stale numbers
Almost every figure in a return moves. Standard deductions, brackets, contribution limits, mileage rates, phase-out thresholds and credit amounts are adjusted annually, and the IRS publishes the new ones each autumn for the following year.
A model knows the figures that existed when it was trained. Ask it for the IRA contribution limit and it will answer with a number, correctly formatted and confidently stated, drawn from whatever year dominated its training. It will not say which year that was.
Statutory changes are worse, because they arrive irregularly. Public Law 119-21, enacted in 2025, changed the standard deduction. A model whose cutoff fell before it will describe the previous regime as current, in detail, without hesitation.
The tell is that there is no tell. Look every current-year figure up in Publication 17 or on IRS.gov, always.
2. Invented citations
Models produce references shaped like real ones. A section number, a publication title, a page. The format is right because the format is highly predictable, which is exactly why it is easy to generate and hard to spot.
A fabricated citation is more dangerous than a plain wrong answer, because it carries the appearance of having been checked. If an answer cites IRC 280A or Publication 587, open the source and read it. Roughly half the effort of using AI for research is verifying its output, which eliminates most of the time it saved.
3. The question you did not ask
Ask how to deduct a home office and you will get an explanation of the home office deduction.
Tax Topic 509 sets out who computes it: self-employed people on Schedule C, partners on Schedule E, farmers on Schedule F. If you are a W-2 employee working from a spare bedroom, the question that decides your answer is which of those categories you are in, and the model will rarely ask, because you did not raise it.
This generalises. A preparer's first ten minutes are spent finding out which facts you have. A model starts from the facts you volunteered, and it has no way to sense the shape of what you left out.
4. Marginal rate confusion
"Will this raise put me in a higher bracket and cost me money?" is the most common tax misunderstanding there is, and models reproduce it because their training data is full of it. Only the income above a threshold is taxed at the higher rate.
The related error runs the other way: applying a single marginal rate to a whole year's income to estimate tax. The answer is wrong by a wide margin, and it is arithmetically clean, so nothing looks amiss.
5. State law
Federal material dominates the training data. State rules are more fragmented, change on their own schedules, and do not automatically follow federal changes. Some states decouple from specific federal provisions deliberately.
A part-year resident who moved between states, or someone who worked remotely across a state line, is in the territory where AI output is least reliable and most confidently delivered.
6. It agrees with you
Push back and the model will usually revise toward what you appear to want. That is a good conversational instinct and a bad research method.
The IRS has already watched what happens when confident, agreeable, wrong tax advice spreads at scale. Misleading advice circulating on social media produced a wave of false claims for the Fuel Tax Credit and the Sick and Family Leave Credit, and the agency has assessed more than $162 million in penalties across more than 32,000 penalties in response. Nobody in that group believed they were doing something wrong. They had seen a confident explanation of why the credit applied to them.
A chatbot that tells you what you want to hear is the same mechanism with better grammar.
What they have in common
None of these are bugs to be fixed in the next release. They follow from what the tool is: a system that generates the most plausible text given your prompt, with no access to your records, no knowledge of what you withheld, and nothing at risk if it is wrong.
Antinozzi and Cooper scored ChatGPT on common tax questions across two filing seasons and found 39 to 47 percent of answers correct, with accuracy falling on the questions that were most common, most complex, most dependent on the taxpayer's own facts, or newest. The Taxpayer Advocate Service found that two leading tax companies' chatbots answered complex questions inaccurately or irrelevantly up to half the time. The IRS's 2026 Dirty Dozen tells taxpayers not to rely on AI-generated answers to complex questions and to verify any calculation an AI provides.
The cost of being wrong is not symmetrical with the cost of being slow. The accuracy-related penalty is 20 percent of the underpayment, on top of the tax and the interest, and you cannot point at the model.
Common questions
Why does AI give outdated tax numbers?
A model learns from text gathered up to a cutoff date. Contribution limits, brackets, standard deductions and mileage rates change every January, and Congress changes the rules underneath them at unpredictable times. A model with an earlier cutoff answers with the older figure in the same confident voice.
Does AI make up tax code sections?
It can. Models generate references that follow the pattern of real citations, so an invented section number looks exactly like a real one. Check every citation against IRS.gov before relying on it.
Why do I get different answers to the same tax question?
Generation is probabilistic, so rephrasing a prompt changes the output. Models also tend to move toward whatever the user seems to want, which means asking a third time for a deduction often produces it.
Can AI handle state taxes?
It is weaker there. States do not automatically adopt federal changes, and the rules differ in ways that are poorly represented in training data compared with federal material. Multi-state years are where the gap is widest.
What is the single most common AI tax error?
Answering the question that was asked instead of the question that mattered. The model has no way to know which of your facts it was not told.
Sources
- Taxpayer Advocate Service, Is AI-generated tax advice making the grade?
- IRS, Dirty Dozen tax scams for 2026 (IR-2026-30)
- IRS, Retirement topics: IRA contribution limits
- Public Law 119-21 (2025)
- IRS, Topic no. 509, Business use of home
- IRS, About Publication 17, Your Federal Income Tax
- IRS, Recordkeeping
- IRS, Accuracy-related penalty
- IRS, IRS assesses $162 million in penalties over false tax credit claims tied to social media (IR-2025-90)
- Antinozzi and Cooper, Is ChatGPT an Accurate Source of Information for Uninformed Taxpayers?, Journal of Emerging Technologies in Accounting 22(1) 23-43 (2025)
Written by Tax Shop in Lone Tree, Colorado, where returns are prepared and signed by an Enrolled Agent. Checked against the sources above on . Penalty amounts and inflation-adjusted figures change each January, and every return is different, so this is not advice about your own situation.
