The tool
Pull the levers the controller pulled
At each assimilation step you get the diagnostic panel the controller had, and the same three actions: CONTINUE at an inflation factor you choose, STOP, or ESCALATE. The budget identity is enforced live. Choose first; then see what the language model, the heuristic and the fixed three-step arm each did, and the reason each gave.
This page replays logged runs. It does not re-run the simulator.
The arithmetic of your choice is exact: the budget ledger and the refusal are the driver’s own, reproduced here. Its consequences are not available. One ES-MDA step is a hundred realisations of a reservoir simulator through ERT, and nothing in a browser runs it.
So: if you pick an α no logged run picked, there is no next panel — the next panel is a consequence of the α that was applied — and there is no score, because a score is a property of an ensemble that does not exist. Neither is interpolated, modelled, or nearest-neighboured into being. The page shows a named real arm’s panel, labelled as that arm’s, or it says nothing.
ES-MDA requires sum(1/alpha_k) = 1 over the assimilations. The driver enforces it as an inequality, refusing any step whose 1/alpha exceeds the budget remaining, so a controller that stops early finishes UNDER budget rather than violating the identity. ESCALATE appears in the action contract and in a separate escalation suite, but it fires nowhere in the sweep this page replays, and the page does not imply otherwise.
Step through a run, and choose before you look
Under common random numbers every arm at a seed starts from the same prior draw, so the posteriors differ only by what the controller did.
Panels are a consequence of the α that was applied, so they belong to a run. This picks whose panels you read; it does not change what the other arms did.
nothing committed
depth 0 · Σ 1/α = 0.000 · 1.000 unspent
Diagnostic panel · step 0
Language model · llm-seed1First step; no trajectory yet.
Above the 0.25 collapse threshold (marked). Training-window CRPS 86.657.
Above the 0.25 collapse threshold (marked). Training-window CRPS 51.963.
ES-MDA requires Σ 1/α = 1. The driver enforces it as an inequality, so a step whose 1/α exceeds this number is refused.
The deliverability signal ESCALATE exists for. It stays low throughout — 0.000% to 0.497% across all 122 panels in this bundle — and no controller in this sweep chose ESCALATE. There is no threshold that fires it: the action is a judgement the controller makes, so the absence is a count, not a trip point that was never reached.
Median across wells; IQR 0.2269; worst well NA3D, which is the same well in all 122 panels in this bundle.
Members the simulator marked OK at this step. Realisations drop out; the panel reports what survived rather than the nominal 100.
params_at_bound is null in every committed panel. It is shown as not computed rather than as a number, because it is not one.
Read from results/diag/llm-seed1/diag-0.json
Your decision
the same three actions the controller hadES-MDA requires sum(1/alpha_k) = 1 over the assimilations. The driver enforces it as an inequality, refusing any step whose 1/alpha exceeds the budget remaining, so a controller that stops early finishes UNDER budget rather than violating the identity.
Choose first, then commit. What the four controllers did at this step is hidden until you have, because reading their answer before choosing turns the exercise into agreement rather than judgement.
Would the comparator score you?
Commit at least one step and this fills in. The rule is the comparator’s own: a difference in assimilation depth arises before any metric is computed, so it refuses to compute one across a depth mismatch.
The scored outcome at seed 1
What the logged runs actually produced, from results/controllers/<run>.json by way of public/data/runs.json. Depth and spend are shown in the same table as the metrics, because they are the reason the columns cannot simply be compared.
| Metric | Language model | Language model, replication | Heuristic (the null) | Fixed three-step (matched effort) |
|---|---|---|---|---|
| Run | llm-seed1 | llm-rep2-seed1 | heuristic-seed1 | fixed3-seed1 |
| Assimilation depth | 3 | 3 | 4 | 3 |
| Σ 1/α spent | 0.400 | 0.400 | 0.733 | 1.000 |
| α schedule | 15, 7.5, 5 | 15, 7.5, 5 | 15, 7.5, 3.75, 3.75 | 7, 3.5, 1.75 |
| Members carried | 100 | 100 | 98 | 99 |
| WWPR — n = 336 observations | ||||
| CRPS (fair) | 118.68 | 118.68 | 117.48 | 115.28 |
| CRPS (reference) | 207.8 | 207.8 | 207.8 | 207.8 |
| CRPSS | 0.4289 | 0.4289 | 0.4346 | 0.4452 |
| Coverage | 0.6994 | 0.6994 | 0.6310 | 0.6488 |
| Spread-skill ratio | 0.7651 | 0.7651 | 0.6719 | 0.6765 |
| RMSE | 225.52 | 225.52 | 225.91 | 228.18 |
| Rank histogram | not in this artifact | 101 bins | 99 bins | not in this artifact |
| WBHP — n = 336 observations | ||||
| CRPS (fair) | 52.482 | 52.482 | 53.06 | 55.805 |
| CRPS (reference) | 63.262 | 63.262 | 63.262 | 63.262 |
| CRPSS | 0.1704 | 0.1704 | 0.1613 | 0.1179 |
| Coverage | 0.4583 | 0.4583 | 0.4048 | 0.3720 |
| Spread-skill ratio | 0.2523 | 0.2523 | 0.2021 | 0.1840 |
| RMSE | 101 | 101 | 100.59 | 102.53 |
| Rank histogram | not in this artifact | 101 bins | 99 bins | not in this artifact |
Coverage sits below its nominal level throughout. ES-MDA under-disperses as a property of the method, not of any controller, so a low coverage here is not a mark against the arm it appears under. A rank histogram is present in only some of the scored artifacts; where it is absent that is said rather than filled in.
The posterior, in three dimensions
A posterior is a mean and a spread. Switch between them: the mean is what the ensemble believes, and the spread is how much it does not know. The narrowing from prior to posterior is the result, and it is only visible because the colour domain is fixed rather than rescaled per field.
3 seed-1 posteriors are exported, together with the single prior they share. 4 arms ran at seed 1 and 3 fields exist, because one posterior no longer does. The language-model arm here is llm-rep2-seed1, the disclosed replication, because llm-seed1’s posterior was deleted from disk by a cleanup defect after it had been scored. Its scores, decision log and transcript stand and are in the repository; nobody can re-derive them from an ensemble.
All 14 scored metrics the two runs produced agree exactly — sample size, fair and reference CRPS, CRPSS, coverage, spread-skill and RMSE, in both families, to the last digit. They are not byte-identical artifacts: the replication additionally carries a rank histogram in WWPR and WBHP that llm-seed1 does not, which the scorecard above prints as its own row. That is why the field shown here is the closest thing to llm-seed1’s posterior that exists, and it is still not that posterior: this is a different ensemble that happened to score the same, and nobody can compare the two fields because one of them is gone. The arm is not deterministic. On seed 9 the replication ran to depth 4 where the original stopped at 3. It stands beside the original in the scorecard and is never pooled into k.
Roughly 1.4 MB of geometry and fields. The download starts on its own when this scrolls into range; the button is here so it is reachable from the keyboard, and so a browser that never fires the observer is not a dead end.
What this view cannot show you
- One property, one seed, in the 3-D view. PORO only, seed 1 only. The replay console covers every seed; the exported fields do not.
- Boxes, not corner-point geometry. Cells are axis-aligned boxes at their centroids from the exported dx/dy/dz. Faults, pillar tilt and non-planar faces are not represented.
- Ensemble statistics, not realisations. A hundred members are reduced to a mean and a standard deviation. No individual realisation was exported.
- A depth rank, not the EGRID layer index. The k index is not in the shipped files. The slider counts down from the top of each active column and says so on its own label.
- Quantised fields. Both statistics are uint16 over each file’s own range — far below plotting resolution, but not float precision.
- No well locations. They are not in the shipped data, so you cannot see why any particular cell narrowed, and the spatial pattern should not be over-read.
- No final panel for a run that spent its schedule. A panel is the state a controller saw before deciding, so a run that never stopped has no panel after its last assimilation. Only the scored result exists for those.