Benchmark
General model leaderboards measure reasoning, code and trivia. None of them measure whether a model can write a defensible HEOR or market-access deliverable. We built that measurement for our own model selection, published it, and keep it current: the preprint is fixed at publication, this page tracks what is actually available.
Last updated 2026-09-16.
Models released after the publication, queued for the next round. We publish the result whichever way it goes.