Pre-registered · UNISIM-I-H · ES-MDA · not peer reviewed
A language model in control of an ensemble smoother
AgentMatch hands a language-model agent the controls of ES-MDA, the ensemble smoother used for reservoir history matching, and runs it on the public UNISIM-I-H benchmark. At each assimilation step the agent reads a diagnostic panel, then chooses an inflation factor, or stops, or escalates. What comes back is a posterior ensemble of reservoir models, not a single best-fit model.
33scored controller runs, decisions and per-step panels committedpublic/data/runs.json → summary.n_runs
13 of 20llm + heuristic runs that finished UNDER the ES-MDA budgetruns.json → summary.n_under_budget
1, 4 and 9seeds where the two arms ran to different assimilation depthruns.json → summary.depth_mismatch_seeds
35,939active cells on the 81×58×20 gridpublic/data/manifest.json → grid.active
Every figure on this page was read from a file in the repository when the page was served. None of them is written into the page source. The files are linked in the footer; the source of each number is printed under it.
What we found, and it is not what we set out to find
We pre-registered a comparison against the same architecture with the language model removed and thresholds on the same panel. It returned no verdict, and the absence of a verdict is the result.
On 3 of the 10 paired seeds the two arms ran a different number of assimilations, because the agent used the one capability under test and stopped early. The comparator refuses to compute a statistic across arms of different depth, and on those pairs it prints nothing. Beneath that refusal, the difference we did observe on water-production CRPS is -0.8661 against a minimum detectable effect of 7.179. This experiment could not have resolved a difference of that size in either direction.
We do not claim the agent beats the heuristic. We claim we could not have told, and we state how far short of telling we were.
Three claims
C1 — the architecture
A model controls a smoother, and a posterior comes back
A language-model agent configures, runs and diagnoses ES-MDA on a public benchmark, and the loop returns a posterior ensemble. 33 scored controller runs are in the repository, each with the decisions the controller made, the diagnostic panel it saw when it made them, and, for 13 of them, the transcript in the model’s own words. This claim stands on its own and does not depend on any comparison.
C2 — a null, with the effect it could have resolved
Ten paired seeds, and not enough power to call it
The language model against a threshold rule reading the identical panel, under common random numbers. Observed between-arm difference on fair CRPS: -0.8661. Minimum detectable effect at k = 10: 7.179. Resolving a difference that size would need 546 paired seeds. A null reported without that bound cannot be told apart from a lack of power, so we never report one without it.
C3 — two confounders the field does not control for
Depth, and the prior draw
Depth: a controller free to stop changes the number of assimilations, and that difference exists before any metric is computed. Prior draw: seed-to-seed variation in the headline metric (sd 29.52) exceeds the within-pair sd (7.219) by enough that an unpaired single-run comparison is measuring the draw rather than the controller. Pairing bought a factor of 5.83. Neither is currently controlled for in this literature.
The comparator refuses to adjudicate, and it is right to
The language-model arm chose to stop after three assimilations on 3 of its 10 seeds — seeds 1, 4 and 9. Every heuristic seed ran four (0 stopped early). The comparator carries a guard that refuses to compare arms of different assimilation depth, and on those pairs it fires and returns no number.
The guard is correct. A difference in depth arises before any metric is computed, so no downstream statistic can separate it from the controller effect it would otherwise be credited to. What makes this a result rather than an obstacle is that the design could not have avoided it without giving up the thing being tested. A controller that may stop cannot be compared at fixed depth to one that may not, and a controller forbidden to stop is not the controller we set out to evaluate. The problem is general: any adaptive controller with a termination rule has it.
Recomputed from public/data/runs.json when this page was served: distinct α-schedules, runs that stopped early, and the α values each arm ever chose. Nothing in this table is written into the page.| Arm | Runs | Distinct schedules | Stopped early | α values ever chosen |
|---|
Language model llm | 10 | 5 | 3 | 1.875, 3, 3.75, 5, 7.5, 15 |
|---|
Language model, replication llm-rep2 | 3 | 2 | 2 | 1.667, 5, 7.5, 15 |
|---|
Heuristic (the null) heuristic | 10 | 2 | 0 | 1.875, 3.75, 7.5, 15 |
|---|
Fixed three-step (matched effort) fixed3 | 10 | 1 | 0 | 1.75, 3.5, 7 |
|---|
The depth-matched subset is a disclosed post-hoc sensitivity check at k = 7 (seeds 2, 3, 5, 6, 7, 8, 10), where the observed difference is +1.26 against an MDE of 2.321. It does not change the conclusion.
A third confounder we proposed, and its own data withdrew
Partway through the sweep, once five of the ten pairs were visible, we registered a third mechanism: that a controller which under-spends the inflation budget buys apparent calibration, because prior spread that was never removed reads to every calibration metric exactly like calibration that was achieved.
The arithmetic is not in dispute. ES-MDA reaches its intended posterior only when the inflation budget is spent, and an ensemble whose spread was never removed will look calibrated to any metric that only inspects spread. The evidence we offered for it is another matter. At an interim look of eleven runs the correlation between spend and calibration was strong. Across the complete 20 runs, spanning spend 0.400 to 1.000, it is -0.107 for coverage and -0.350 for spread-skill. The discriminating test is within-arm, holding the controller fixed, and there the coverage relationship is absent in both arms. Accuracy does not move as predicted either: +0.134 for fair CRPS and +0.206 for RMSE.
We report that as a null and we leave it on the page. A pre-registered mechanism that its own data declined to support is the most direct evidence we have for C3: the confounders here are strong enough to produce, and then withdraw, a result that looked solid at eleven runs. The claim was falsifiable, more data falsified it, and it was caught before submission rather than by a referee.
Source: paper/figures/f4-stats.tex → \FFourRCoverage, \FFourRSsr, \FFourRCrps, \FFourRRmse, \FFourSpendMin, \FFourSpendMax, \FFourN
Play the experiment
Not a video of the demo, and not the simulator either. The decisions and the budget arithmetic are live; the consequences are a replay.
Seed 1 gives all three controllers the same prior draw, so the three posteriors differ only by what the controller did. Pull the levers the controller pulled: the inflation factor at each step, STOP, ESCALATE. Read the same diagnostic panel the agent read at that step, then read what it actually chose and the reason it gave, in its own words.
The physics is enforced, not decorated. ES-MDA requires the inflation budget to satisfy the sum of 1/α over the assimilations equalling one. The driver enforces that as an inequality: it refuses any step that would overspend, so a controller may finish under budget, and 13 of 20 of the controller runs did. The 3-D view carries the posterior mean of PORO and its ensemble spread over 35,939 active cells, because a posterior is a mean and a spread, and a page that showed only the mean would be making the mistake this paper is about.
The panels, the reasons and the scores are the logged ones. Choose an α no logged run chose and the page gives you no next panel and no score, because one ES-MDA step is a hundred reservoir simulations and nothing in a browser runs it. It says so rather than interpolating one.
What this page does not claim
- Not that the agent beats the heuristic. It was better on 2 of 10 pairs and the difference is below what this experiment could resolve.
- Not that the two controllers are the same. A bounded null is not an equivalence result.
- Not that the posterior is well calibrated. ES-MDA under-disperses, and that is a property of the method rather than of either controller.
- Not that the language-model arm is reproducible run to run. It is not deterministic, which is why the seeds are paired and why the replication is reported as new samples.
- Not that contamination is ruled out. A direct memorisation probe over 14 scored probes returned no evidence of recitation — median relative error 0.443 against 0.144 for a scale-null. A probe can raise an alarm without certifying safety.
- One case, UNISIM-I-H. The breadth here is across seeds, not across fields.
- Not peer reviewed. Nothing in this repository has been reviewed end to end by a human.
0.815%realisation dropout: 22 of 2,700 realisation-runs across 27 runsresults/dropout-summary.json
GATE PASSED8 metrics at 0% largest relative differenceresults/repro-gate-verdict.json
15 pagesthe paper, counted from the PDF itself at export timepublic/paper/agentmatch.pdf
The reproduction gate’s own verdict file records that The driver has been edited since this hash (guards added 2026-09-02/03). The analysis code paths it certifies are unchanged; a re-run of the gate against the current driver is listed in the WORKPLAN. It does not certify the current driver.