Roadmap

What Phase 2 is designed to validate

Four workstreams follow from the proof-of-concept's limitations. Their direction is defined; their final allocation is not.

  • 01

    Measurement validation

    Test a multi-classifier panel across model families, add a small naive and informed human baseline, and begin predictive-validity studies against deployment-style agentic tasks.

  • 02

    Statistics and design

    The current instrument-validation plan is approximately 1,370–1,450 model runs. The roughly 175 runs per model/objective target is an optional powered expansion, deferred to adaptive top-ups or a later phase.

  • 03

    Environment scaling and anti-gaming

    A structure-only MUD orientation area exists. The original publishable graph, scenario-ready content, cast, and objective machinery remain in progress, alongside a public-development / hidden-evaluation split.

  • 04

    Configuration robustness

    Use a representative three-model subset for prompt, temperature, and reasoning-budget sweeps, rather than multiplying those cells across the full roster.

Get involved

Fund it. Build it. Run it.

Phase 1 cost $99.59, billing-verified against provider exports. Phase 2 keeps the same discipline: provisional assumptions now, calibration data before commitment, and actuals published as spent.

01

Fund it

The $3,500 figure is a provisional planning envelope, not a settled budget. The downloadable model exposes base, and high assumptions for model mix, live-NPC calls, human recruitment, reasoning-heavy runs, etc.

Partner on funding →
Provisional Phase 2 cash budget under base, and high scenarios
Plan component Base High
Model API + judge panel current instrument-validation plan · 1,410 / 1,450 runs $638 $1,604
Human baseline 12 / 30 participants · compensation and platform fees $1,035 $2,520
Storage, hosting, and exports $100 $250
Contingency provider drift, retries, and calibration reruns $266 $875
Current plan total $2,039 $5,249
Optional powered expansion adds about 3,100 model runs to reach roughly 4,550 total +$1,336 +$3,199
Combined planning range $3,375 $8,448
$3,500 provisional envelope · the base combined scenario currently fits; the high scenario does not.
Engineering and analysis · contributed in kind and excluded from cash totals.

Planning model, July 2026. Final N, provider mix, human-baseline design, and budget are calibration- and pilot-gated. The adaptation rule will be preregistered before paid pilot batteries.

Download budget workbook
02

Build it

Repository public after release prep

The single-run harness is working; the next value is in the benchmark slice itself. We are looking for collaborators who can help finish the environment and objectives, harden the minimum paid-batch path, and calibrate the instrument without expanding the harness beyond what the experiment needs.

Partner on the build →
  • Original environment + content · publishable graph, original prose, scenario-ready rooms
  • Objective binding · scored cast, trust/suspicion loops, and the two Phase 1 objectives
  • Minimum batch harness · enable paid batteries only after readiness and preregistration gates
  • Measurement calibration · multi-judge drift, agreement, and reference-band checks
  • Statistics + preregistration · paired contrast power, adaptation rule, and final allocation
  • Eval integrity · hidden cells, retire-and-publish workflow, and artifact audits
03

Run it

We are recruiting post-calibration pilot partners—not offering a validated diagnostic today. Pilot partners can help define deployment-relevant tasks and, after calibration and preregistration, run a model or agent stack through the completed instrument.

Join the post-calibration pilot cohort →
  • Before the pilot · agree on a deployment-relevant question and data-handling boundary
  • After calibration · full per-turn transcripts, telemetry, and algorithmic failure-mode traces
  • What we will not sell yet · composite scores or leaderboard placement before validation
Scope

Current facts, future targets

Proof-of-concept (released)

  • 12 rooms · 4 NPCs · 14 items
  • 7 command types
  • 2 hidden social objectives
  • 50 turns per run
  • 650 scored runs · 13 models
  • $99.59 billed cost

Phase 2 (active build)

  • Structure-only 81-room / 57-NPC MUD orientation area seeded and validated
  • Single-run harness and keyed live-model smoke complete
  • Original publishable layout and scenario-ready content in progress
  • Scored cast and objective machinery not yet bound to the orientation area
  • Paid batch switch deliberately gated
  • Calibration, preregistration, and final allocation pending
Partner

Partner on Phase 2

Use one form for funding, technical collaboration, or the post-calibration pilot cohort. It opens your email app with the details filled in; this site does not store the submission.

Or write directly