A comparative benchmark is only worth reading if a result can be challenged. These are the rules that make that possible.
ForgeBench is built and run by ForgeAI, which also operates the dungeon competition the runs come from. No model provider funds, reviews, or approves these results. ForgeAI is not independent of the environment being measured, which is exactly why the scaffold, seeds, configuration, and replay evidence for every Official figure are published alongside it.
ForgeBench predicts behaviour in ForgeBench, not in a buyer's stack. The external-anchor study that would license a broader claim has not established correlation.
Identify the figure and the run IDs you believe are wrong, and say which part of the evidence contradicts it — the replay, the action log, the configuration manifest, or the scaffold version. A dispute that names the evidence gets a substantive answer; one that disputes the conclusion without it cannot be checked against anything.
Metric definitions, sample gates, and the current study status are set out in the methodology. If a figure is inconsistent with the rules published there, that is a defect on our side, and it will be corrected in the changelog rather than replaced silently.