research · 2026-09-10

What financial benchmarks actually measure

Four published sets, three different abilities, and one uncomfortable fact about contamination.

Primary sources, read and citedDeskworth · published · updated · 4 sources

The four sets that get quoted together do not measure the same thing, and treating them as interchangeable is how a leaderboard becomes meaningless.

SetWhat it isolates
FinQAMulti-step arithmetic, with the reasoning program itself scored 1
TAT-QAReading a table and its prose together 2
BizBenchKnowing which financial concept applies, separated from reading and from computing 3
FinanceBenchAnswering from a whole filing — and refusing when it does not answer 4

#Program accuracy is the metric that ages well

FinQA scores the answer and the reasoning separately. A system that reaches the right total through the wrong program is right this quarter and wrong the next, and only the second metric sees it coming.

#Contamination is the default assumption

Three of the four have been public since 2021 under permissive licences, on platforms built to be crawled. Assuming a modern model has never seen them is not conservative — it is unrealistic. Any score published without stating the model’s training cutoff next to the set’s publication date is not interpretable.

Still missingDeskworth has not run any of these benchmarks. There is no Deskworth score anywhere on this site, and the conditions under which one could appear are published in advance.

#Sources

  1. FinQA: A Dataset of Numerical Reasoning over Financial DataZhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, William Yang Wangpeer-reviewed paper · EMNLP 2021 · anthology 2021.emnlp-main.300 · published 2021-11-01 · read 2026-09-10 · CC BY 4.0https://aclanthology.org/2021.emnlp-main.300/
  2. TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in FinanceFengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, Tat-Seng Chuapeer-reviewed paper · ACL-IJCNLP 2021 · doi 10.18653/v1/2021.acl-long.254 · published 2021-08-01 · read 2026-09-10 · CC BY 4.0https://aclanthology.org/2021.acl-long.254/
  3. BizBench: A Quantitative Reasoning Benchmark for Business and FinanceRik Koncel-Kedziorski, Michael Krumdick, Viet Lai, Varshini Reddy, Charles Lovering, Chris Tannerpeer-reviewed paper · arXiv:2311.06602v2 · revised 2024-03-12 · published 2023-11-11 · read 2026-09-10 · CC BY-NC-SA 4.0https://arxiv.org/abs/2311.06602
  4. FinanceBench: A New Benchmark for Financial Question AnsweringPranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, Bertie Vidgenpeer-reviewed paper · arXiv:2311.11944v1 · published 2023-11-20 · read 2026-09-10 · CC BY-NC-ND 4.0https://arxiv.org/abs/2311.11944

Read next