EVIDENCE FIRST · HISTORICAL RESULTS ARE NOT CURRENT PRODUCTION PROOF
Participant contract / R5

Compare like with like.Keep the boundary visible.

Raw models, chat products, coding agents, agent runtimes and orchestrators are different participants. Official ranks stay inside a frozen comparison protocol.

Comparison partitions

Participant tracks

comparison trackModel onlyMODEL_ONLY

One pinned model snapshot, no external tools or product workflow.

  • pinned_model_snapshot
  • frozen_prompt_protocol
  • no_external_tools
comparison trackModel + standard toolsSTANDARD_TOOLS

Pinned model snapshots receive the same frozen browser, code and file tools.

  • pinned_model_snapshot
  • identical_tool_contract
  • frozen_prompt_protocol
comparison trackAgent / systemAGENT_SYSTEM

Complete products, coding agents, agent runtimes and orchestrators compete as systems.

  • product_or_system_snapshot
  • interaction_protocol
  • tool_and_model_disclosure
comparison trackBest result under equal budgetEQUAL_BUDGET

Different participant classes may compete only under the same money, time and operator envelope.

  • shared_money_cap
  • shared_wall_clock_cap
  • shared_operator_protocol
  • actual_cost_and_time_receipts
Coverage

Planned coverage

No official scores yet. Catalog inclusion is planned coverage, not participation, qualification or an official PRO Bench score.

RAW_MODELPinned API model snapshots

Multiple providers

PLANNEDNOT_EVALUATED

Participant tracks:
MODEL_ONLYSTANDARD_TOOLSEQUAL_BUDGET

Model policy: FIXED_PINNED

  • API

Every provider model and exact version becomes a separate participant snapshot when admitted by an evidence-bound campaign.

Source: UNKNOWN

CLOUD_CHAT_PRODUCTChatGPT

OpenAI

PLANNEDNOT_EVALUATED

Participant tracks:
AGENT_SYSTEMEQUAL_BUDGET

Model policy: USER_OR_PRODUCT_SELECTED

  • WEB_APP
  • DESKTOP_APP
  • MOBILE_APP

Evaluate the complete product workflow, including the captured model-selection mode and enabled built-in capabilities; do not relabel it as a raw model run.

Declared product source

CODING_AGENTCodex

OpenAI

PLANNEDNOT_EVALUATED

Participant tracks:
AGENT_SYSTEMEQUAL_BUDGET

Model policy: USER_OR_PRODUCT_SELECTED

  • DESKTOP_APP
  • CLI
  • IDE
  • CLOUD_SANDBOX

Capture the exact Codex surface, sandbox/worktree policy, model configuration and all allowed tools for every campaign snapshot.

Declared product source

CODING_AGENTClaude Code

Anthropic

PLANNEDNOT_EVALUATED

Participant tracks:
AGENT_SYSTEMEQUAL_BUDGET

Model policy: USER_OR_PRODUCT_SELECTED

  • CLI
  • WEB_APP
  • DESKTOP_APP

Capture the exact execution surface, model configuration, repository permissions and test/tool policy instead of attributing the result to an unspecified Claude model.

Declared product source

AGENT_RUNTIMEOpenClaw

OpenClaw Foundation

PLANNEDNOT_EVALUATED

Participant tracks:
AGENT_SYSTEMEQUAL_BUDGET

Model policy: CONFIGURABLE_PROVIDER

  • SELF_HOSTED
  • MESSAGING_GATEWAY
  • DEVICE_APPS

Pin the OpenClaw release, host configuration, provider/model, skills, memory state and granted device permissions as one participant snapshot.

Declared product source

AGENT_RUNTIMEHermes Agent

Nous Research

PLANNEDNOT_EVALUATED

Participant tracks:
AGENT_SYSTEMEQUAL_BUDGET

Model policy: CONFIGURABLE_PROVIDER

  • SELF_HOSTED
  • CLI
  • MESSAGING_GATEWAY

Pin the Hermes Agent release, provider/model, tools, skills and initial memory state; learned skills produced during the run remain part of that run's evidence.

Declared product source

ORCHESTRATORPRO-1

PromanOS

PLANNEDNOT_EVALUATED

Participant tracks:
AGENT_SYSTEMEQUAL_BUDGET

Model policy: SYSTEM_ROUTED

  • API
  • WEB_APP

Disclose the exact PRO-1 build, routing policy, selected models, tools, retries, fallbacks, actual cost, wall-clock time and artifact receipts.

Declared product source

Admission

Required participant snapshot

  • product_or_model_identity
  • exact_version_or_build
  • access_surface
  • account_plan_and_region
  • model_selection_policy
  • underlying_model_identity_or_disclosure_state
  • enabled_tools_connectors_and_permissions
  • operator_and_interaction_protocol
  • initial_memory_and_workspace_state
  • prompt_retries_fallbacks_and_tool_calls
  • budget_limits_actual_cost_and_wall_clock
  • outputs_artifacts_hashes_and_receipts