Silvia Beats Every Major AI Product on Tax

Read about the latest tax benchmarks
We took seven of the most popular AI products and put them through ten of the hardest tax questions we could write. The kind with year-specific figures, provisions that interact in strange ways, and state rules that break from the federal treatment. Every answer was graded from 1 to 10 for factual accuracy by a judge who did not know which product produced it.
Silvia scored 8.73, the highest of any product tested. The next-best came in at 8.40, and the rest of the field trailed from there.
That result is the headline. The more interesting story is why it happened, and we ran a much larger study to answer that, which we will get to below.
Featured on CNBC. ProCap Financial CEO Anthony Pompliano discussed Silvia’s tax results on CNBC’s Squawk Box. Watch the segment.
Why Silvia comes out ahead
Most AI tools answer tax questions from memory. They absorbed a lot of text during training, and when you ask a question, they produce a fluent answer based on what they remember. That works fine for common questions and falls apart on the hard ones, where the answer turns on the exact wording of a statute or a figure that changed this year.
Silvia works differently. We built a curated library of primary tax law: the Internal Revenue Code, Treasury regulations, IRS publications and instructions, and the statutes and revenue-department guidance for individual states. When you ask Silvia a tax question, it retrieves from that library at the moment you ask, and it can run a live web search when it needs current information. Instead of hoping the AI remembers the law correctly, we give it the actual law to work from, and we make it cite the source so you can check the answer yourself.
That is the difference the chart is measuring. On hard tax questions, retrieving from the real law beats answering from memory.
We went deeper than ten questions
Ten questions make a clean headline, but they cannot tell you how much the tooling matters or when. So we ran a second, much larger study: 200 expert-level questions, more than 1,500 answers, and nearly 3,000 blind judgments, with the whole thing run twice.
This time we tested the same system four ways: with no tools at all, with only our curated tax library, with only web search, and with the full setup. That let us isolate exactly what the tooling adds. We also scored every answer on two separate things. Accuracy: is the answer correct? Grounding: can you verify it, meaning does it cite the real statute, and when you look that citation up, does it actually say what the answer claims? An answer can be perfectly correct and still useless to a tax professional if there is no way to check it, which is why we scored the two separately.
What the deeper study found
On the hardest questions, the tooling made a large difference. The full system beat the no-tools version by more than two points on accuracy and more than three points on grounding, both statistically solid.
On everyday client planning questions, something more interesting happened, and it is worth being honest about. The AI was already quite accurate on these without any tools. Adding tooling barely changed the accuracy score, because the model already handles common planning questions well.
What the tooling changed was verifiability. On those everyday questions, the curated library was the thing that let the answer cite the actual law so a professional could check it. The accuracy was already there. The ability to trust and verify it was not, until the tooling supplied it.
That distinction matters more than any single number. Web search and a curated law library do different jobs. Web search brings in facts the system does not already know. The library supplies the citation that lets someone verify the answer. A system built on only one of them will be weak somewhere. The full setup was the only one never meaningfully beaten on either test, and that is the property that counts when you cannot know in advance what kind of question someone will bring.
We published all of it
We put all 200 questions and their scores online under an open license, free for anyone to inspect. If we are going to say Silvia is the most accurate AI on tax, the right thing to do is let people check that claim against the actual data rather than take our word for it.
There is a broader reason too. Generic AI benchmarks tell you which model is smartest in general. They do not tell you whether an AI system actually works in a specific field like tax, where the answer has to match the real law, for the real year, with a citation you can follow. The only way to know that is to test it in the field itself and show your work.
Explore the full benchmark here: huggingface.co/datasets/cfosilvia/silvia-tax-bench
For the complete methodology, the confidence intervals, and every chart, the full research write-up is here: https://www.cfosilvia.com/ai-lab/tax-benchmark-analysis
Financial information notice
This content is for informational purposes only. It is not financial, investment, or legal advice. Past performance does not guarantee future results. Consult a qualified professional before making financial decisions.
Share this article
Your money deserves superintelligence.
Free forever. No card required. Give Silvia five minutes and see what an AI CFO trained on your money actually knows.


