LLMs Outperform Human Invoice Reviewers
Challenge
Invoice reviews by humans are slow, inconsistent, and expensive, impacting billing accuracy and firm overhead.
Solution
A benchmark study compared LLMs (Large Language Models) with human reviewers—including early-career lawyers—on invoice accuracy and speed.
Results
•LLMs achieved up to 92% accuracy versus 72% for human reviewers.
•Line-item classification reached 81% F-score for AI vs. 43% for humans.
•LLMs processed invoices in 3.6 seconds, compared to 194–316 seconds for humans.
•Processing costs dropped by 99.97%, from $4.27 to mere cents per invoice.