The tool

Pull the levers the controller pulled

At each assimilation step you get the diagnostic panel the controller had, and the same three actions: CONTINUE at an inflation factor you choose, STOP, or ESCALATE. The budget identity is enforced live. Choose first; then see what the language model, the heuristic and the fixed three-step arm each did, and the reason each gave.

This page replays logged runs. It does not re-run the simulator.

The arithmetic of your choice is exact: the budget ledger and the refusal are the driver’s own, reproduced here. Its consequences are not available. One ES-MDA step is a hundred realisations of a reservoir simulator through ERT, and nothing in a browser runs it.

So: if you pick an α no logged run picked, there is no next panel — the next panel is a consequence of the α that was applied — and there is no score, because a score is a property of an ensemble that does not exist. Neither is interpolated, modelled, or nearest-neighboured into being. The page shows a named real arm’s panel, labelled as that arm’s, or it says nothing.

33logged runs available to replay, across four armspublic/data/runs.json
117CONTINUE decisions, against 5 STOP and 0 ESCALATErecomputed from every step in the bundle
13 of 20runs that finished under Σ 1/α = 1, which the identity permitsruns.json → summary.under_budget_runs
35,939active cells rendered in the 3-D view, property POROpublic/data/manifest.json

ES-MDA requires sum(1/alpha_k) = 1 over the assimilations. The driver enforces it as an inequality, refusing any step whose 1/alpha exceeds the budget remaining, so a controller that stops early finishes UNDER budget rather than violating the identity. ESCALATE appears in the action contract and in a separate escalation suite, but it fires nowhere in the sweep this page replays, and the page does not imply otherwise.

Step through a run, and choose before you look

Under common random numbers every arm at a seed starts from the same prior draw, so the posteriors differ only by what the controller did.

Panels are a consequence of the α that was applied, so they belong to a run. This picks whose panels you read; it does not change what the other arms did.

Your schedule so far

nothing committed

depth 0 · Σ 1/α = 0.000 · 1.000 unspent

Diagnostic panel · step 0

Language model · llm-seed1
χ² per observation (mismatch against the data)28,255

First step; no trajectory yet.

Spread retention · WWPR1.0000

Above the 0.25 collapse threshold (marked). Training-window CRPS 86.657.

Spread retention · WBHP1.0000

Above the 0.25 collapse threshold (marked). Training-window CRPS 51.963.

Inflation budget remaining before this step1.000

ES-MDA requires Σ 1/α = 1. The driver enforces it as an inequality, so a step whose 1/α exceeds this number is refused.

Off-target fraction0.497%

The deliverability signal ESCALATE exists for. It stays low throughout — 0.000% to 0.497% across all 122 panels in this bundle — and no controller in this sweep chose ESCALATE. There is no threshold that fires it: the action is a judgement the controller makes, so the absence is a count, not a trip point that was never reached.

Per-well spread0.4506

Median across wells; IQR 0.2269; worst well NA3D, which is the same well in all 122 panels in this bundle.

Ensemble members carried100

Members the simulator marked OK at this step. Realisations drop out; the panel reports what survived rather than the nominal 100.

Parameters at boundnot computed

params_at_bound is null in every committed panel. It is shown as not computed rather than as a number, because it is not one.

Read from results/diag/llm-seed1/diag-0.json

Your decision

the same three actions the controller had
Action at this step
Budget before this step1.000
This step spends 1/α0.067
Budget after0.933
Smallest α you can still afford1

ES-MDA requires sum(1/alpha_k) = 1 over the assimilations. The driver enforces it as an inequality, refusing any step whose 1/alpha exceeds the budget remaining, so a controller that stops early finishes UNDER budget rather than violating the identity.

Choose first, then commit. What the four controllers did at this step is hidden until you have, because reading their answer before choosing turns the exercise into agreement rather than judgement.

Would the comparator score you?

Commit at least one step and this fills in. The rule is the comparator’s own: a difference in assimilation depth arises before any metric is computed, so it refuses to compute one across a depth mismatch.

The scored outcome at seed 1

What the logged runs actually produced, from results/controllers/<run>.json by way of public/data/runs.json. Depth and spend are shown in the same table as the metrics, because they are the reason the columns cannot simply be compared.

Every value read from the results/controllers/ artifact named in the first row. Columns are in the bundle’s display order and are never sorted by a metric: two arms that ran to different depth are not comparable, and the arms that did match are separated by less than this experiment can resolve.
MetricLanguage modelLanguage model, replicationHeuristic (the null)Fixed three-step (matched effort)
Runllm-seed1llm-rep2-seed1heuristic-seed1fixed3-seed1
Assimilation depth3343
Σ 1/α spent0.4000.4000.7331.000
α schedule15, 7.5, 515, 7.5, 515, 7.5, 3.75, 3.757, 3.5, 1.75
Members carried1001009899
WWPR — n = 336 observations
CRPS (fair)118.68118.68117.48115.28
CRPS (reference)207.8207.8207.8207.8
CRPSS0.42890.42890.43460.4452
Coverage0.69940.69940.63100.6488
Spread-skill ratio0.76510.76510.67190.6765
RMSE225.52225.52225.91228.18
Rank histogramnot in this artifact101 bins99 binsnot in this artifact
WBHP — n = 336 observations
CRPS (fair)52.48252.48253.0655.805
CRPS (reference)63.26263.26263.26263.262
CRPSS0.17040.17040.16130.1179
Coverage0.45830.45830.40480.3720
Spread-skill ratio0.25230.25230.20210.1840
RMSE101101100.59102.53
Rank histogramnot in this artifact101 bins99 binsnot in this artifact

Coverage sits below its nominal level throughout. ES-MDA under-disperses as a property of the method, not of any controller, so a low coverage here is not a mark against the arm it appears under. A rank histogram is present in only some of the scored artifacts; where it is absent that is said rather than filled in.

The posterior, in three dimensions

A posterior is a mean and a spread. Switch between them: the mean is what the ensemble believes, and the spread is how much it does not know. The narrowing from prior to posterior is the result, and it is only visible because the colour domain is fixed rather than rescaled per field.

3 seed-1 posteriors are exported, together with the single prior they share. 4 arms ran at seed 1 and 3 fields exist, because one posterior no longer does. The language-model arm here is llm-rep2-seed1, the disclosed replication, because llm-seed1’s posterior was deleted from disk by a cleanup defect after it had been scored. Its scores, decision log and transcript stand and are in the repository; nobody can re-derive them from an ensemble.

All 14 scored metrics the two runs produced agree exactly — sample size, fair and reference CRPS, CRPSS, coverage, spread-skill and RMSE, in both families, to the last digit. They are not byte-identical artifacts: the replication additionally carries a rank histogram in WWPR and WBHP that llm-seed1 does not, which the scorecard above prints as its own row. That is why the field shown here is the closest thing to llm-seed1’s posterior that exists, and it is still not that posterior: this is a different ensemble that happened to score the same, and nobody can compare the two fields because one of them is gone. The arm is not deterministic. On seed 9 the replication ran to depth 4 where the original stopped at 3. It stands beside the original in the scorecard and is never pooled into k.

Reservoir view

Roughly 1.4 MB of geometry and fields. The download starts on its own when this scrolls into range; the button is here so it is reachable from the keyboard, and so a browser that never fires the observer is not a dead end.

What this view cannot show you