How Deskworth defines a valid benchmark
Eleven conditions, written before we have a single score, so that no score can be published without them.
Publishing the conditions before the result is the only way to keep them honest. Once a number exists, every missing condition becomes negotiable.
#The line that gets skipped most often
“Which model?” is usually answered with a family name. A family name is not a model. Endpoints change under a stable label, so the exact identifier and the date of the run are two separate requirements, and both are mandatory.
#The line that hides the most variance
Retrieval. On an open-book financial benchmark, most of the spread between two reported results comes from how the passage was found, not from the model that read it. A result that does not describe its retrieval component is a result about an undisclosed system.
#Licences are part of the protocol
Two of the four sets we catalogue are non-commercial, and one forbids derivative works. That constrains what may be redistributed and what may be adapted, and it has to be checked before a run rather than after publication.
Still missingDeskworth has not run any benchmark. The catalogue publishes the protocol the next core pass will follow.
#Sources
- czyssrs/FinQA — official repositoryZhiyu Chen et al.https://github.com/czyssrs/FinQA
- NExTplusplus/TAT-QA — official repositoryNExT++, National University of Singaporehttps://github.com/NExTplusplus/TAT-QA
- PatronusAI/financebench — dataset cardPatronus AIhttps://huggingface.co/datasets/PatronusAI/financebench
- BizBench: A Quantitative Reasoning Benchmark for Business and FinanceRik Koncel-Kedziorski, Michael Krumdick, Viet Lai, Varshini Reddy, Charles Lovering, Chris Tannerhttps://arxiv.org/abs/2311.06602
Read next
- Evaluations and benchmarksFour published benchmarks for financial reasoning, documented from their primary sources — and a plain statement that Deskworth has not run any of them.
- A green test is not evidenceTested, packaged and observed are three different claims. Most arguments about what a system can do come from treating them as one.