The loaded method changed the process. It did not beat an informed baseline on judgement.
Benchmark record, current as at 10 August 2026
A 45-run preregistered study across two models, three fictional cases and three conditions. Raw counts and uncomfortable limitations are part of the result.
Process separation was clear. Judgement stayed too close to call.
Process. The loaded condition asked before answering, compared options and tagged evidence on every run.
Judgement. The all-facts baseline matched or exceeded the loaded condition on all three judgement checks.
Boundary. This grid does not resemble a real operator arriving with partial knowledge and a blind spot.
Eight preregistered checks, scored out of 15 per condition.
| Check | Question only | All facts | Loaded method |
|---|---|---|---|
| Process checks | |||
| Asked before answering | 0 of 15 | 0 of 15 | 15 of 15 |
| Compared real options, including doing nothing | 0 of 15 | 1 of 15 | 15 of 15 |
| Tagged its evidence | 0 of 15 | 0 of 15 | 15 of 15 |
| Named an untested constraint | 9 of 15 | 15 of 15 | 15 of 15 |
| Stated what would change the call | 15 of 15 | 15 of 15 | 15 of 15 |
| Judgement checks | |||
| Surfaced the hidden decision | 0 of 15 | 15 of 15 | 14 of 15 |
| Direction matched the sealed key | 9 of 15 | 15 of 15 | 15 of 15 |
| Caught the planted bias | 6 of 15 | 15 of 15 | 15 of 15 |
Some cells are wholly or partly definitional because a single-turn prompt cannot ask questions or obtain facts it was not given. The methodology and limitations record those cases.
Derived process figure
Produced the full brief structure
This is the lowest of the five process check counts. It is not an independent ninth result and was not the preregistered headline.
Information access was controlled across three conditions.
- 01
Question only
The frozen prompt only. One turn. If the model asks questions they are not answered.
- 02
All facts
The frozen prompt with the full fact sheet appended in the first message. One turn. Equal information to the loaded condition.
- 03
Loaded method
The frozen method file, then the frozen prompt. The model interrogates; the operator answers only from the fact sheet and replies 'That has not been assessed' to anything off-sheet.
The record includes 18 deviations.
They are not footnotes to be hidden. They change how confidently the result should be read.
- 01
The composite is not the pre-registered headline
rubric-operationalisations.md:23 pre-registers checks 6 and 7 as the primary result, and rubric:3 states a run's score is a vector, not a single number. Leading with a Group 1 composite is a primary-endpoint substitution.
- 02
Pilot runs were pooled and the promise not to pool them was deleted afterwards
protocol.md:23 excluded them, and still does. Commit 85f8c71, 14 July 2026, deleted the matching sentence from PREREGISTRATION.md and set that record to COMMITTED. Scoring finished 6 July 2026, eight days earlier. The original wording survives in commit 1dfeb6d. This page reports only the 45 pre-registered runs. The attribution in this entry was itself corrected on 11 August 2026: it previously named protocol.md as the file edited.
- 03
The published pre-registration commit ID does not exist
5e89de4842693444c27e893470efecb242486cd2 is in no ref. The real commit is 3e5e46f1a7cc154e1b91120ec9336e3eb089fd2c. Anyone who followed the published verification procedure failed at step 1, and the failure would have looked like their own mistake.
- 04
No transcript is published
protocol.md:34 commits that every published count is traceable to a downloadable transcript. None is.
- 05
The sealed reference answers are not published
Every adjudicated cell was settled against a key nobody outside can read. Its SHA-256 was verified against the pre-registration table and that check cannot be repeated while the file is withheld.
- 06
Transcripts were not de-identified before scoring
Contrary to protocol.md:27. The sole scorer was the operator who ran the grid. Every filename carries its condition and model, and each carries an operator note stating the verdict.
- 07
The original scoring was mildly biased toward widening the gap
Stricter than the rubric text inside Arm A, looser inside Arm B. Two cells were overturned on adjudication and both ran that way, both in the baseline's favour.
- 08
The single-scorer cross-check was never run
protocol.md:30 makes an LLM cross-check the condition of falling back to one scorer. The fallback happened and the cross-check did not.
- 09
Baseline partial-pass counts are not published
protocol.md:31 requires them prominently, per case, and the rubric calls the full-versus-partial distinction the core of the benchmark.
- 10
The rubric set no integration threshold
Both scoring passes invented one and disagreed. The rule is now pre-committed: two of three facts with the third absent scores PARTIAL.
- 11
The published rationale for the one near-miss overstates it
The record says two of three facts were integrated. The adjudication found one.
- 12
The sharpest discriminators in the sealed key went unused
The Case C falsification element and the explicit NOT clauses in both cases were applied by neither scoring pass.
- 13
The GO case was published with the informed baseline removed
The site showed 0 of 6 against 6 of 6 and omitted the middle column, which is also 6 of 6.
- 14
The case where the file adds least was never identified
protocol.md:36 requires it. It is Case B, Larkfield.
- 15
Per-case, per-condition counts are not published
The rubric names them as the headline unit. Direction is now split by case; the other seven checks are not.
- 16
The alternative reading of the options check was kept internal
A more lenient reading credits further all-facts runs on the check the scorer calls load-bearing.
- 17
The question-only condition was contaminated on Case B
ChatGPT invoked live web search unprompted during those runs. Those six runs carry the whole of the direction deficit reported anywhere on this site.
- 18
The 14-day refresh commitment lapsed
No date-stamped history was published. The dates and the arithmetic are in the status section.
What this benchmark cannot establish.
- 01
No condition in this grid resembles a real operator. The question-only condition has nothing and the all-facts condition has everything. An operator who arrives with most of what matters and a blind spot they cannot see is the condition the product exists for, and it was not run.
- 02
One person ran the grid, scored it, held the sealed reference answers, and answered the file's questions during the runs. The re-score was carried out by three agents working from the published rubric, blind to the original scoring. Neither they nor the adjudication were independent of GreenSquare AI, and this page does not use that word.
- 03
Two published cells are definitional rather than empirical, and a third is partly so. A single-turn condition cannot ask before answering, and it cannot obtain three facts that are not in its prompt.
- 04
The three checks that separate most cleanly are the behaviours the file instructs the model to perform. A reasonable reading is that they measure whether the file was loaded, not whether the model reasoned better.
- 05
The all-facts condition was not information-equal between the two models on Case B. The Claude fact sheet carries two cross-reference pointers the ChatGPT one lacks.
- 06
n is 15 per condition, three repeats per cell. The grid is small and unbalanced, with Case A running on one model only.
- 07
The run inventory reconciles. Both directory counts were verified on disk, 45 grid transcripts and 10 pilot files, and the arithmetic requires 9. The tenth file was opened on 30 August 2026: it is the scoring writeup for the nine excluded pilot runs, and it states in its own opening paragraph that they are not pooled with the pre-registered grid. This entry previously reported the inventory as unreconciled and the tenth file as unopened.
- 08
What was tested is a file fixed by its published hash. Nothing here commits any later release to being that same file.
Inspect what is public, and what remains withheld.
- Scored runs
- 45
- Preregistration commit
- 3e5e46f
- Authored
- 5 July 2026 at 09:10:37 AEST
- Transcripts
- Not published
- Sealed answers
- Hash verified, contents not published