Benchmark

    How well AI models write HEOR — measured, not asserted.

    General model leaderboards measure reasoning, code and trivia. None of them measure whether a model can write a defensible HEOR or market-access deliverable. We built that measurement for our own model selection, published it, and keep it current: the preprint is fixed at publication, this page tracks what is actually available.

    Last updated 2026-09-16.

    Method

    • Every model writes the same controlled set of HEOR and market-access tasks — the task set from the writing benchmark, unchanged between models.
    • Scoring is by AI Delphi: several independent models score the same output over structured rounds, and their judgments converge. No single model grades the field.
    • These are the UNANCHORED results. The publication also reports human-anchored scoring; this page publishes the unanchored panel only, so every row is comparable to every other row.
    • Runs marked with the Untraceable Grounding Codex are given a curated domain knowledge layer — HTA guidance, methods conventions and reporting standards — retrieved at generation time. The same model without it is a separate row, never averaged in.
    • Each row records the provider's exact version string and the date of the run. A model that changes under the same name is re-tested rather than edited.
    • Results are published whichever way they go, including for models we do not offer, and retired models stay in the table because the publication cites them.

    The published method

    In testing

    Models released after the publication, queued for the next round. We publish the result whichever way it goes.

    • GPT-6 AstraOpenAI
    • Fable 5.1Anthropic
    • Claude Sonnet 5Anthropic
    • GPT-5.6 SolOpenAI— Released after the preprint, which measures GPT-5.5.
    • GPT-5.6 TerraOpenAI— Released after the preprint, which measures GPT-5.5.