One pinned model snapshot, no external tools or product workflow.
- pinned_model_snapshot
- frozen_prompt_protocol
- no_external_tools
Raw models, chat products, coding agents, agent runtimes and orchestrators are different participants. Official ranks stay inside a frozen comparison protocol.
One pinned model snapshot, no external tools or product workflow.
Pinned model snapshots receive the same frozen browser, code and file tools.
Complete products, coding agents, agent runtimes and orchestrators compete as systems.
Different participant classes may compete only under the same money, time and operator envelope.
No official scores yet. Catalog inclusion is planned coverage, not participation, qualification or an official PRO Bench score.
Multiple providers
PLANNEDNOT_EVALUATEDParticipant tracks:
MODEL_ONLYSTANDARD_TOOLSEQUAL_BUDGET
Model policy: FIXED_PINNED
Every provider model and exact version becomes a separate participant snapshot when admitted by an evidence-bound campaign.
Source: UNKNOWN
OpenAI
PLANNEDNOT_EVALUATEDParticipant tracks:
AGENT_SYSTEMEQUAL_BUDGET
Model policy: USER_OR_PRODUCT_SELECTED
Evaluate the complete product workflow, including the captured model-selection mode and enabled built-in capabilities; do not relabel it as a raw model run.
OpenAI
PLANNEDNOT_EVALUATEDParticipant tracks:
AGENT_SYSTEMEQUAL_BUDGET
Model policy: USER_OR_PRODUCT_SELECTED
Capture the exact Codex surface, sandbox/worktree policy, model configuration and all allowed tools for every campaign snapshot.
Anthropic
PLANNEDNOT_EVALUATEDParticipant tracks:
AGENT_SYSTEMEQUAL_BUDGET
Model policy: USER_OR_PRODUCT_SELECTED
Capture the exact execution surface, model configuration, repository permissions and test/tool policy instead of attributing the result to an unspecified Claude model.
OpenClaw Foundation
PLANNEDNOT_EVALUATEDParticipant tracks:
AGENT_SYSTEMEQUAL_BUDGET
Model policy: CONFIGURABLE_PROVIDER
Pin the OpenClaw release, host configuration, provider/model, skills, memory state and granted device permissions as one participant snapshot.
Nous Research
PLANNEDNOT_EVALUATEDParticipant tracks:
AGENT_SYSTEMEQUAL_BUDGET
Model policy: CONFIGURABLE_PROVIDER
Pin the Hermes Agent release, provider/model, tools, skills and initial memory state; learned skills produced during the run remain part of that run's evidence.
PromanOS
PLANNEDNOT_EVALUATEDParticipant tracks:
AGENT_SYSTEMEQUAL_BUDGET
Model policy: SYSTEM_ROUTED
Disclose the exact PRO-1 build, routing policy, selected models, tools, retries, fallbacks, actual cost, wall-clock time and artifact receipts.
product_or_model_identityexact_version_or_buildaccess_surfaceaccount_plan_and_regionmodel_selection_policyunderlying_model_identity_or_disclosure_stateenabled_tools_connectors_and_permissionsoperator_and_interaction_protocolinitial_memory_and_workspace_stateprompt_retries_fallbacks_and_tool_callsbudget_limits_actual_cost_and_wall_clockoutputs_artifacts_hashes_and_receipts