> ## Content Index
> Fetch the complete content index at: https://aiadoption.org/llms.txt
> Use this file to discover other available public pages before exploring further.

# AI in professional domains
- URL: https://aiadoption.org/ai-analysis/ai-in-professional-domains/
- Published: 2026-10-02T04:06:27.000Z
- Updated: 2026-10-02T04:06:27.000Z
- Description: AI now scores 60-90% on tax, corporate finance, legal reasoning, and mortgage benchmarks — with the top 15 models separated by 3 percentage points. The procurement framing that worked when 'capable enough for the domain' was a binary no longer matches the data.
- Author: Jassie
- Tags: AI Analysis, AI Procurement, AI Strategy, Large Language Models

AI now scores between 60% and 90% on tax, corporate finance, mortgage processing, and legal reasoning benchmarks. The top 15 models are separated by 3 percentage points on each. The procurement framing that worked when "capable enough for the domain" was a binary no longer matches the data.

*\[CHART fig\_258\_2026 — TaxEval v2: top model accuracy on US tax-return scenarios\]*

The capability picture, by domain:

TaxEval v2, covering US individual and corporate tax-return scenarios, shows top 15 models clustered in a narrow performance band. Performance differs by only a few percentage points across vendors. The benchmark covers tax-return preparation, entity classification, deductions, and credits.

CorpFin v2, covering corporate finance reasoning (financial statement analysis, valuation, and M&A scenarios), shows similar clustering. Kimi K2.5 leads at approximately 69%, with most top models within 3–5 points of the leader.

Chart: CorpFin v2, top model accuracy on corporate-finance scenarios. 

CaseLaw v2, covering US case-law reasoning, citation, and precedent analysis, shows GPT-5.1 leading at 73.4%, with GPT-4.1 at 69.9% and tight clustering across the rest of the top 15.

Chart: CaseLaw v2, top model accuracy on US legal reasoning. 

Finance Agent v1.1, covering agentic financial workflows (multi-step research and analysis tasks), shows more variation than the document-level benchmarks, with Claude leading. That suggests agentic capability in professional domains is differentiating faster than language-only capability.

The pattern across these benchmarks: frontier-tier capability is now competent across high-competency professional domains, but reliability, the dimension that actually matters for regulated workflow deployment, is not what these benchmarks measure.

A model that scores 73% on CaseLaw v2 is competent for case-law reasoning. It is not reliable for case-law reasoning at the level required to deploy in a regulated legal workflow without human review. Those are different procurement questions. The benchmark answers the first; the second requires evaluation methodology the public benchmarks do not provide.

## Three procurement implications follow.

**The first: "is this model capable enough for tax/finance/legal?" is now answered "yes, in the frontier tier."** The binary procurement question that filtered vendors based on capability threshold is no longer the right question. Every frontier vendor passes that threshold for these domains.

**The second: the operative procurement question has moved to "is this model reliable enough for our specific workflow at the required level of human review burden?"** That is a workflow-specific evaluation question, not a benchmark-comparison question. The vendor's TaxEval score does not answer it. Internal evaluation against the enterprise's specific tax workflow, with the enterprise's specific risk tolerance for errors, does.

**The third: failure-mode characterisation matters more than capability average.** A model that scores 87% on TaxEval is wrong on 13% of cases. The procurement question is not "which model scores highest on average": it is "what are the failure modes on the 13% it gets wrong, and are those failure modes compatible with our review and remediation workflow?" Some failure modes (clearly wrong outputs that a human reviewer easily catches) are compatible with light-touch human review. Other failure modes (plausibly wrong outputs that look right to a non-expert) are compatible only with expert review at every step, which negates the deployment economics.

The benchmark data does not directly characterise these failure modes. Vendor evaluation in regulated professional domains needs to include failure-mode analysis on the specific workflow, not just aggregate accuracy on the public benchmark.

There is a complication worth being explicit about. Each [Vals.ai](http://vals.ai/?ref=aiadoption.org) benchmark uses a methodology specific to the domain: tax-return scenarios for tax, M&A cases for finance, case-law precedents for legal. The methodologies are not directly comparable. A 73.4% on CaseLaw v2 is not equivalent to a 73.4% on TaxEval v2 in terms of what the score implies for downstream deployment. Procurement frameworks that treat scores across these benchmarks as comparable are introducing variance they cannot bound.

## Three concrete actions follow.

**Replace "vendor capability comparison" sections in procurement frameworks for regulated domains with "vendor reliability characterisation" sections.** The reliability section should include domain-specific accuracy, but should also include failure-mode taxonomy, error consistency under prompt variation, and behaviour under edge-case inputs. The benchmark scores feed this analysis; they do not substitute for it.

**Set explicit risk tolerance for false-confidence failures versus obvious-error failures.** False-confidence failures (where the model is wrong but presents the answer confidently and the answer looks right) are the procurement-relevant failure mode in regulated domains, because they survive the review process more often than obvious errors. Vendor evaluation should specifically test for this failure mode, with prompts designed to elicit it.

**Build vendor-portability into the workflow architecture.** Given that capability across frontier-tier vendors is now within procurement-comparable bands and the differentiation is on cost and reliability, the optimal vendor for a specific regulated workflow may shift more than annually. Workflow architectures that lock into a single vendor at the prompt or post-processing layer foreclose this optionality. Architectures that abstract vendor identity behind a uniform interface keep it.

> The prescription: shift procurement evaluation for regulated professional domains from "which vendor scores highest on the public benchmark?" to "which vendor's failure modes are most compatible with our review and remediation workflow at our target cost-per-task?" The first question is answerable from public data and produces flawed procurement decisions. The second is answerable only from workflow-specific evaluation and produces procurement decisions that align with downstream performance.

---

### Sources

- **Primary**: Stanford AI Index 2026, Chapter 2 (Technical Performance) 2.5 — [hai.stanford.edu/ai-index/2026](http://hai.stanford.edu/ai-index/2026?ref=aiadoption.org)
- **Domain-specific benchmark data**: [Vals.ai](http://vals.ai/?ref=aiadoption.org), 2026 — TaxEval v2, CorpFin v2, CaseLaw v2, Finance Agent v1.1
- **Reliability framing**: industry case material on regulated-domain AI deployment (tax-firm, BigLaw, mid-market finance), 2025–2026