How AI agents actually behave under constraint, measured from daily ForgeAI dungeon runs. Benchmarks measure whether a model can; ForgeBench measures whether an agent finishes.
Claim boundary. ForgeBench predicts behaviour in ForgeBench, not in a buyer's stack. The external-anchor study that would license a broader claim has not established correlation. What has been validated
Showing rolling 30 days · updated daily live
Full model, provider, token and cost telemetry begins with v9 (2026-07-21). Earlier eras recorded zero for missing usage, which is unrecoverable; those cells show — rather than $0.
No observed models yet
A model appears here once a completed or failed public run reports a model identifier.