Evaluate AI tax preparation software across four pillars, not a feature list. Accuracy: demand a source-to-field audit trail, confidence flags, and clear exception handling. Integrations: confirm tested support for your specific software and version, not a generic claim. Security: verify encryption, access controls, retention, WISP fit, and whether the vendor trains on your data. Review controls: require explicit human-review checkpoints and a clear answer to who signs. Treat security certifications as a question you ask the vendor to prove—never an assumption.
The short answer: evaluate against four pillars, not a demo
The market for AI tax preparation software is crowded, and nearly every vendor leads with the same two promises: pinpoint accuracy and dramatic time savings. Neither claim is verifiable from a scripted demo. To evaluate AI tax preparation software the way a firm owner should—and the way the rules governing your practice implicitly require—you need a structured framework that tests the product against your obligations, not the vendor's marketing.
Four pillars cover the ground that matters: accuracy you can audit (a source-to-field trail, confidence flags, and disciplined exception handling), integrations that fit your specific software and version (tested and supported, not a generic "works with everything" claim), security and data governance (encryption, access controls, retention, whether the vendor trains on your data, and how the tool fits your Written Information Security Plan), and human-review controls (explicit checkpoints and an unambiguous answer to who reviews and signs).
One rule runs through all four: a credentialed professional remains responsible for every return, so the software's job is to make that professional faster and better-informed, never to replace their judgment. And one caution runs through the security pillar specifically—do not assume any vendor, including any you already like, holds a security certification such as SOC 2. Certifications are something you make the vendor prove, in writing, with a current report. This guide walks each pillar in the order you should test it, and ends with a checklist you can take into a demo.
Why a framework beats a feature list
Feature lists reward whoever writes the longest one. A structured evaluation rewards the tool that actually reduces your risk and rework. The difference matters because the buyer of AI tax software is not buying a gadget—they are inserting an automated system into a workflow where a human being signs a legal document and is "primarily responsible for the overall substantive accuracy" of it. The evaluation has to be built around that responsibility.
Two professional standards make the framing concrete. The AICPA's revised Statements on Standards for Tax Services, effective January 1, 2024, added a standard on reliance on tools (SSTS No. 1, §1.4) stating that a member may reasonably rely on a tool but that "use of the tool does not absolve the member of their professional obligations," and that a practitioner should take reasonable steps to determine the tool is appropriate for its purpose. A parallel data-protection standard (§1.3) requires reasonable efforts to safeguard taxpayer data and references the FTC Safeguards Rule directly. In other words, evaluating the tool is itself part of your professional duty—the "reasonable steps" the standard asks for. A framework is how you document that you took them.
Treasury Department Circular 230 reinforces the point from the other direction: a practitioner must exercise due diligence, and reliance on another's work product is reasonable only when the professional uses reasonable care in engaging, supervising, and evaluating that work. An AI tool is work product you engage and supervise. Evaluate it accordingly.
Pillar 1: Accuracy you can audit, not accuracy you're told about
Every vendor claims high accuracy. The number is meaningless on its own—it depends on document quality, form mix, and how "accuracy" is measured—and no headline figure protects you when a single wrong entry understates a client's tax. Accuracy under IRC §6694 is enforced against the signing preparer, not the software vendor, so what you are really evaluating is not a percentage. It is whether the tool lets a reviewer verify its work quickly and completely.
The source-to-field audit trail
The single most important accuracy feature is a source-to-field audit trail: for every figure the AI enters into the return, the reviewer can see exactly which document and which field it came from, ideally with the source image displayed alongside the entered value. This turns review from blind trust into a fast visual reconciliation. Without it, a reviewer either re-keys the return to check it—destroying the time savings—or accepts figures on faith, which is precisely the "accepting output blindly" that professional standards warn against. Ask to see the trail in a live walkthrough, on a messy real-world document, not a clean sample.
Confidence flags and low-certainty extraction
A trustworthy tool knows what it is unsure about and says so. Look for confidence flags that surface low-certainty extractions—a smudged figure, an ambiguous box on a substitute form, a handwritten annotation—and route them to the reviewer's attention rather than entering them silently. Large language and vision models can produce confident-sounding output that is wrong, so a system that presents every field as equally certain is hiding exactly the information a reviewer needs. The goal is to concentrate professional attention on the fields most likely to be wrong.
Exception handling and missing-document detection
Accuracy is not only about the figures that are present; it is about the ones that are missing. Evaluate how the tool handles exceptions: does it compare against the prior year and flag a 1099 that appeared last year but not this year? Does it read a brokerage summary and notice a referenced form that was never uploaded? Does it flag a year-over-year swing that deserves a second look? Silent gaps are how an incomplete return reaches a client. A strong tool treats missing and inconsistent data as questions to raise, not blanks to leave. For deeper treatment, see our guides on missing-document detection and complex brokerage statements.
Pillar 2: Integrations tested for your software and your version
"Works with your tax software" is one of the most abused phrases in the category, and it is a real source of buyer's regret. There is a large gap between a vendor that has built and tested a supported connection to Drake, ProSeries, or Lacerte for the current tax year, and one that exports a generic file you re-import by hand or, worse, re-key. An unconfirmed integration does not save time—it relocates the data-entry work and adds a reconciliation step.
Ask about your specific program and version
Integrations are version-sensitive. Tax software changes every year, and a connection validated for last season is not automatically valid for this one. When you evaluate, name your exact program and the version or tax year you run, and ask the vendor to confirm that combination is tested and supported now—not "on the roadmap." If your firm runs more than one program across offices, confirm each one. Our program-specific walkthroughs cover the mechanics for Drake, ProSeries, and Lacerte.
How the drafted return is delivered
Confirm the delivery mechanism, because it determines whether the tool fits your existing review workflow or forces a new one. The best outcome is that the AI deposits a populated, review-ready return inside the tax program your team already prepares and reviews in, so the reviewer opens a familiar file with nothing to re-enter. A tool that lands the data in a separate portal, requiring your team to look in two places or copy figures across, adds friction that erodes the benefit. Ask to watch a return travel end to end, from document upload to a drafted return open in your software.
Scope of forms and entity types
Finally, match the tool's coverage to your book of business. A product tuned for simple 1040s with W-2 income may stumble on the K-1s, multi-account brokerage statements, rental schedules, or business returns your firm actually files. Bring representative documents from your hardest returns to the evaluation, not the vendor's easy demo file.
Pillar 3: Security and data governance you can verify
Security is not a soft "nice to have" in this category—it is a legal obligation, and the AI tool you adopt becomes part of the environment your compliance program must cover. Paid tax preparers are treated as "financial institutions" under the Gramm-Leach-Bliley Act, which places them under the FTC Safeguards Rule. That rule requires a written information security program, a designated Qualified Individual to oversee it, a written risk assessment, access controls, multi-factor authentication, and encryption of customer data in transit and at rest. The IRS reinforces the same expectations through Publication 4557, Safeguarding Taxpayer Data, and the Security Summit's WISP template in Publication 5708. There is no small-firm exception.
Encryption, access controls, and retention
Translate the Safeguards Rule into vendor questions. Is client data encrypted both in transit and at rest? Who at the vendor can access your clients' files, and under what controls—is access role-based and logged? Is multi-factor authentication enforced for your team and the vendor's staff? How long is data retained after a return is complete, and can you require deletion on a schedule that matches your own retention policy? These are not paranoid questions; they are the same controls your WISP has to describe. Whatever the vendor answers becomes a line item in your plan, so get it in writing.
Does the vendor train on your data? An IRC §7216 question
Ask directly whether the vendor uses your clients' tax data to train its models or for any purpose beyond preparing your clients' returns. This is not merely a privacy preference—it can be a legal line. IRC §7216 imposes criminal penalties on preparers who knowingly or recklessly use or disclose a taxpayer's return information for purposes other than preparing that return, unless an exception or the taxpayer's specific consent applies, and the Treasury regulations were written to cover electronic and software-based processing. If a tool feeds client return information into model training or shares it with third parties, you may need informed §7216 consent from the taxpayer before that data ever flows. A vendor that cannot give you a clear, documented answer on data use is a vendor you cannot evaluate on security.
Certifications: ask, do not assume
Security certifications such as SOC 2, ISO 27001, or a completed third-party penetration test are useful signals—but only if they are real, current, and scoped to the service you are buying. Treat certification as a question you make the vendor answer, not a claim you accept. Ask whether they hold a SOC 2 report, request to review it (or at least a summary and the report date) under NDA, and confirm what it actually covers. Never assume a tool is certified because it looks polished or because a competitor is; and be skeptical of any marketing that gestures at "bank-level security" without a document behind it. The honest posture, for you and for any vendor, is that a certification is proven with a report, not asserted in a sales deck. You can see how we present these questions on our own security page, and our security checklist for AI tax software goes deeper.
Pillar 4: Human-review controls and who signs
The fourth pillar is where many AI tools quietly fail, because it is the one most in tension with an "automation" sales pitch. A tool that races to "done" without forcing a human checkpoint is a liability, not a feature. Under federal law the paid preparer must review the return, sign it, and enter a valid PTIN; failure to sign or furnish a PTIN carries penalties under IRC §6695, and the IRS actively pursues "ghost preparers" who omit their identity. The software can never be the signer. So evaluate how deliberately the tool preserves the human review step.
Explicit review checkpoints
Look for a workflow with a defined, unskippable review checkpoint between "AI has drafted the return" and "return is approved." The tool should present the draft as ready for review, not ready to file, and it should make the reviewer's job efficient by surfacing exactly what needs attention: the confidence flags, the exceptions, the year-over-year anomalies, and the source-to-field trail from Pillar 1. The distinction to test is whether the product is designed to keep a professional in the loop by default, or whether human review is an optional setting a rushed firm can turn off.
Who reviews, who approves, who signs
Map the tool's workflow to your firm's roles. Can you route a return through a preparer and then a reviewer, with sign-off recorded at each stage? Does the system capture who approved what, and when—the kind of trail that supports the due-diligence documentation Circular 230 and the credit due-diligence rules expect? The answer to "who signs" is always a credentialed human, but a good tool makes the path to that signature auditable. This is also where AI naturally extends past preparation: a well-reviewed return is the starting point for year-round planning, and the same connected client context supports notice response later.
| Pillar | What to ask the vendor | What good looks like | Red flag |
|---|---|---|---|
| Accuracy | Can a reviewer trace every entered figure to its source document and field? | Source-to-field audit trail with the document image shown next to each value; confidence flags on low-certainty fields | A headline accuracy percentage and no way to verify individual entries |
| Integrations | Is my specific program and this tax year's version tested and supported now? | A named, supported connection that deposits a review-ready return inside your existing tax software | "Works with everything," generic export files, or a "coming soon" roadmap answer |
| Security | Is data encrypted in transit and at rest, and do you train on my clients' data? | Encryption both ways, role-based logged access, MFA, defined retention, and a clear "no training / consent-gated" answer | Vague data-use answers, or model training on client return information without §7216 consent |
| Security (certification) | Do you hold a current SOC 2 report, and may I review its scope and date? | A current third-party report you can review under NDA, scoped to the service | "Bank-level security" or an assumed certification with no report to show |
| Review controls | Is there an unskippable human-review checkpoint, and who signs? | A ready-for-review state, role-based sign-off recorded per stage, PTIN preparer always the signer | Human review is optional or off by default; the tool markets "no review needed" |
The evaluation checklist you can take into a demo
Use the table above as your scorecard and the questions below as your script. Run the same questions past every vendor so you are comparing like with like, and insist on live demonstration over verbal assurance wherever the item is demonstrable.
- Accuracy. Show me the source-to-field audit trail on a poor-quality real document. Where do confidence flags appear? How does the tool surface a missing form and a year-over-year anomaly?
- Integrations. Is my exact program and current tax-year version tested and supported today? Walk a return from upload to a drafted return open in that software. Can it handle my hardest form types—K-1s, multi-account brokerage, business returns?
- Security. Is data encrypted in transit and at rest? Who can access it and how is access logged? Is MFA enforced? What is the retention period and can I control deletion? Do you use my clients' data to train models, and if so, how do you handle §7216 consent?
- Certification. Do you hold a SOC 2 (or comparable) report? May I see its date and scope under NDA? What does it actually cover?
- Review controls. Where is the mandatory review checkpoint? Can it be disabled? How are preparer and reviewer sign-offs recorded, and who ends up as the signing, PTIN-holding professional?
Score each pillar, weight security and review controls heavily—they carry the legal exposure—and let a tool fail the evaluation on a single non-negotiable rather than averaging its way to a passing grade.
Running a defensible evaluation
Finally, treat the evaluation as a process you can point to later, because the AICPA reliance standard effectively asks you to. Bring your own documents, not the vendor's. Involve the person who will actually review returns in the tool, not only the person who signs the contract. Put the vendor's answers on data use, retention, and certification in writing—an email or a completed security questionnaire—so your Written Information Security Plan can reference something concrete.
Then run a small pilot before you commit the firm. Push a handful of real returns, across your typical form mix, through the tool during a quiet period. Measure the two things that matter: how much genuine time the reviewer saves after accounting for verification, and how often the tool's flags catch something a rushed manual process might have missed. A tool that clears the four pillars and survives a real pilot is one you can adopt with your name on the line—which, in this profession, is the only standard that counts. That is the honest promise of AI in tax preparation: not a replacement for the professional who signs, but a faster, better-documented path to the review only that professional can perform.
Automated Tax Prep you can actually evaluate
Tax Automate turns client documents into review-ready draft returns in Drake, ProSeries, and Lacerte—with a source-to-field audit trail, confidence flags, and explicit human-review checkpoints—so your evaluation has something real to inspect.
Explore Automated Tax Prep →Frequently asked questions
What is the most important thing to evaluate in AI tax preparation software?
A source-to-field audit trail. Because the signing preparer is responsible for accuracy under IRC §6694, the tool's most valuable feature is letting a reviewer verify every entered figure against its source document quickly. A headline accuracy percentage you cannot verify offers no protection; a visible, inspectable trail does.
How do I know if an AI tax tool really integrates with my software?
Name your exact program and the current tax-year version and ask the vendor to confirm that combination is tested and supported now, then watch a return travel from upload to a drafted return open in that software. Treat generic 'works with everything' claims, export files you re-import, or 'coming soon' answers as unconfirmed integrations that relocate the data-entry work rather than remove it.
Should I ask whether an AI tax vendor is SOC 2 certified?
Yes—as a question, not an assumption. Ask whether the vendor holds a current SOC 2 (or comparable) report, request to review its date and scope under NDA, and confirm what it covers. Never assume a tool is certified because it looks polished. A certification is proven with a report, not asserted in marketing.
Will an AI tax vendor train on my clients' data?
That depends entirely on the vendor, and you must ask directly. If a tool uses client return information beyond preparing that client's return—including model training or sharing with third parties—IRC §7216 may require specific, informed taxpayer consent, and the regulations cover electronic and software processing. A vendor that cannot give a clear, written answer on data use cannot be evaluated on security.
Does using AI tax software reduce my professional responsibility?
No. Under the AICPA's SSTS §1.4, a member may reasonably rely on a tool, but 'use of the tool does not absolve the member of their professional obligations.' A paid preparer must still review the return, sign it, and include a PTIN. Evaluate tools for how firmly they preserve that human-review checkpoint, not how completely they automate it away.
This article is based on published IRS guidance, the Internal Revenue Code preparer provisions, the FTC Safeguards Rule, IRS Publications 4557 and 5708, and AICPA professional standards, alongside Tax Automate product documentation. Any figures are illustrative and labeled as such—no statistical claims are made. Rules and requirements are current as of publication and should be verified for the applicable tax year.
- IRS — PTIN Requirements for Tax Return Preparers
- IRS — Tax Preparer Penalties (IRC §6694 and §6695)
- IRS — Circular 230, Regulations Governing Practice before the IRS
- IRS — Section 7216 Information Center
- IRS — Publication 4557, Safeguarding Taxpayer Data (PDF)
- IRS — Security Summit: Tax Pros Must Have a WISP (Publication 5708)
- FTC — Safeguards Rule: What Your Business Needs to Know
- AICPA — Statements on Standards for Tax Services (SSTS)