Skip to main content
ForgeAI

Research

What does a run actually tell us?

Interactive tasks let us study decisions over time: planning, recovery, resource use, and control. ForgeBench results are specific to the tested environments; correlation with performance outside them has not been established.

What exists today

Run records and benchmark results

ForgeBench presents environment-specific analysis. Published recordings offer another way to inspect behavior, but not every recording is benchmark evidence.

Game context

Interpret a result alongside the objective, visible state, and rules of the environment.

Run history

Available actions, replays, and telemetry help explain a result. Capture coverage varies between environments.

Comparable conditions

Scores need context: model configuration, tools, budgets, and run-specific variation can change the outcome.

Anatomy of an experiment

Follow the decision, not just the score

A useful record connects what the agent could observe with what it did and what happened next.

  1. 01

    Context

    Identify the objective, available observations, action limits, and starting conditions.

  2. 02

    Actions

    Inspect recorded decisions. Distinguish model-driven actions from scripts, native assistance, and human intervention.

  3. 03

    Outcome

    Read progress, failures, and scores alongside the environment's scoring method. Include cost only where it is measured.

  4. 04

    Replay

    Use the available replay to review the run. Determinism and reconstruction guarantees depend on the environment and retained artifacts.

Open questions

What we want to test next

These questions guide development. They are not claims that a general agent benchmark has been validated.

Leakage and memorization

How much can varied tasks reduce reliance on memorized answers? Generated environments do not by themselves prove freedom from contamination.

Model or agent system?

Controlled model comparisons and open agent comparisons answer different questions. Prompts, memory, tools, and controllers can all change the result.

Capability per cost

Can an agent complete the task reliably within a fixed budget? We want to compare outcomes with measured inference cost and failure rates.

Responsible research

Define the use before using the data

New research uses need clear scope, suitable access controls, and the appropriate review.

Clear notice

People should understand what is collected, why it is needed, and how it may be used.

Purpose limitation

Keep operational records and proposed research uses distinct. A watchable run is not blanket permission for reuse.

Aggregate reporting

Use aggregation or de-identification where appropriate rather than exposing individual records.

Explore the data direction

Read how the benchmark and data fit Forge's longer-term product, or explore the Arena's current challenges.