Modus Labs, the research team at AI-native accounting firm Modus, has released Financial Audit Bench (FAB), an open-source benchmark that tests whether AI agents can complete the multi-step procedures of a financial statement audit, according to a post published October 1, 2026 by Lightspeed Venture Partners. The benchmark was run against 11 frontier AI models, with results, code, dataset and a live leaderboard made public.
What the benchmark measures
According to the Lightspeed post, FAB contains 45 workpapers across eight audit areas, modeled on the fieldwork a staff-level auditor does. To protect client records, Modus generated six synthetic company audits from aggregate statistics on past audits.
The post said Modus Labs paired PhDs and applied AI engineers with veterans of public accounting, including advisor Jim Burton, former Chief Auditor at Grant Thornton, and that auditors put more than 1,100 hours into developing and reviewing the benchmark.
Modus chief technology officer Pranav Pillai is quoted describing the design: "At firms today, senior auditors spot-check key details in a workpaper rather than re-performing audit procedures. We mirror this process in FAB."
What the results showed
Across the 11 frontier models tested, the models would pass more than 85% of the individual checks, but none fully passed more than 69% of the procedures it attempted, the post said. Because audit work has to hold up every time, each procedure was run eight times; when the top model was required to pass all eight runs, its success rate fell from 69.17% to 44.4%.
Performance varied sharply by area. Agents passed journal-entry completeness checks 94.5% of the time, while revenue testing came in at 8% and accounts receivable at 4%. The post said the weakest areas are the ones estimated to take auditors the most time.
Modus found that agents often stopped looking too early and trusted the wrong evidence. In one example cited, an agent tested receivables against the client's own cash-receipts report when it should have used the bank statement.
How Modus frames the findings
The post said the results only measure the tasks specifically scoped in the benchmark, and that whether an agent can run an audit alone is a separate question. "We therefore recommend deploying agents as co-pilots rather than autonomous preparers," Pillai said.
On why audit needs its own evaluation, Pillai said: "While time savings are straightforward to quantify, audit quality has historically been an area with less precise measures." The post added that it is estimated that 40% of audits contain significant errors, and that because auditors test samples, much of a company's activity never gets a direct look.
Model performance moved quickly during development: in the weeks Modus spent building the benchmark, the leading strict-pass rate rose from below 50% to 69.2%, the post said.
What comes next
Modus's auditors found that agents often wrote long descriptions of their process but didn't emphasize the key evidence or conclusions, according to the post, which argued future evaluations should also measure whether another auditor can review and follow the work.
Modus has released the paper, code, dataset and live leaderboard so anyone can rerun the benchmark and build on it. The write-up is published by Lightspeed Venture Partners, which notes the views expressed are those of the authors and do not necessarily represent the views of Lightspeed.
What to do
- Treat AI agents in audit work as co-pilots rather than autonomous preparers, as Modus recommends, and keep reviewer checks on their output.
- Expect weaker agent performance in revenue testing and accounts receivable, where Modus reported pass rates of 8% and 4%.
- Watch for agents stopping their search too early or relying on the wrong evidence source, such as using a client's cash-receipts report instead of a bank statement.
- Build systems that can test and adopt stronger models as they become available, as the post advises, given the leading strict-pass rate rose from below 50% to 69.2% during development.
- Use the publicly released paper, code, dataset and live leaderboard to rerun the benchmark on your own models.
Key facts and where they come from
- FAB is an open-source benchmark from Modus Labs testing AI agents on financial statement audit procedures.
These results come from Financial Audit Bench (FAB), an open-source benchmark released by Modus Labs, the research team at the AI-native accounting firm Modus.
- Eleven frontier models passed over 85% of individual checks but none fully passed more than 69% of procedures attempted.
the models would pass more than 85% of the individual checks. But none fully passed more than 69% of the procedures it attempted.
- Requiring the top model to pass all eight runs cut its success rate from 69.17% to 44.4%.
When the top model was required to pass all eight runs, its success rate fell from 69.17% to 44.4%.
- Accounts receivable pass rates averaged 4%, revenue testing 8%, and journal-entry completeness 94.5%.
Agents passed journal-entry completeness checks 94.5% of the time. However, revenue testing came in at 8%, and accounts receivable at 4%.
- The benchmark has 45 workpapers across eight audit areas, using six synthetic company audits.
The benchmark has 45 workpapers across eight audit areas, modeled on the fieldwork a staff-level auditor does.
- Auditors spent more than 1,100 hours developing and reviewing FAB.
Auditors put more than 1,100 hours into developing and reviewing the benchmark.
- Modus recommends using agents as co-pilots rather than autonomous preparers.
"We therefore recommend deploying agents as co-pilots rather than autonomous preparers," says Pranav.
