GreenSquare AI

The loaded method changed the process. It did not beat an informed baseline on judgement.

A 45-run preregistered study across two models, three fictional cases and three conditions. Raw counts and uncomfortable limitations are part of the result.

Process separation was clear. Judgement stayed too close to call.

Process. The loaded condition asked before answering, compared options and tagged evidence on every run.

Judgement. The all-facts baseline matched or exceeded the loaded condition on all three judgement checks.

Boundary. This grid does not resemble a real operator arriving with partial knowledge and a blind spot.

Eight preregistered checks, scored out of 15 per condition.

Benchmark checks by condition
CheckQuestion onlyAll factsLoaded method
Process checks
Asked before answering0 of 150 of 1515 of 15
Compared real options, including doing nothing0 of 151 of 1515 of 15
Tagged its evidence0 of 150 of 1515 of 15
Named an untested constraint9 of 1515 of 1515 of 15
Stated what would change the call15 of 1515 of 1515 of 15
Judgement checks
Surfaced the hidden decision0 of 1515 of 1514 of 15
Direction matched the sealed key9 of 1515 of 1515 of 15
Caught the planted bias6 of 1515 of 1515 of 15

Some cells are wholly or partly definitional because a single-turn prompt cannot ask questions or obtain facts it was not given. The methodology and limitations record those cases.

Derived process figure

Produced the full brief structure

15 of 15Loaded method

This is the lowest of the five process check counts. It is not an independent ninth result and was not the preregistered headline.

Information access was controlled across three conditions.

  1. 01

    Question only

    The frozen prompt only. One turn. If the model asks questions they are not answered.

  2. 02

    All facts

    The frozen prompt with the full fact sheet appended in the first message. One turn. Equal information to the loaded condition.

  3. 03

    Loaded method

    The frozen method file, then the frozen prompt. The model interrogates; the operator answers only from the fact sheet and replies 'That has not been assessed' to anything off-sheet.

Read the full methodology

The record includes 18 deviations.

They are not footnotes to be hidden. They change how confidently the result should be read.

  1. 01
    The composite is not the pre-registered headline

    rubric-operationalisations.md:23 pre-registers checks 6 and 7 as the primary result, and rubric:3 states a run's score is a vector, not a single number. Leading with a Group 1 composite is a primary-endpoint substitution.

  2. 02
    Pilot runs were pooled and the promise not to pool them was deleted afterwards

    protocol.md:23 excluded them, and still does. Commit 85f8c71, 14 July 2026, deleted the matching sentence from PREREGISTRATION.md and set that record to COMMITTED. Scoring finished 6 July 2026, eight days earlier. The original wording survives in commit 1dfeb6d. This page reports only the 45 pre-registered runs. The attribution in this entry was itself corrected on 11 August 2026: it previously named protocol.md as the file edited.

  3. 03
    The published pre-registration commit ID does not exist

    5e89de4842693444c27e893470efecb242486cd2 is in no ref. The real commit is 3e5e46f1a7cc154e1b91120ec9336e3eb089fd2c. Anyone who followed the published verification procedure failed at step 1, and the failure would have looked like their own mistake.

  4. 04
    No transcript is published

    protocol.md:34 commits that every published count is traceable to a downloadable transcript. None is.

  5. 05
    The sealed reference answers are not published

    Every adjudicated cell was settled against a key nobody outside can read. Its SHA-256 was verified against the pre-registration table and that check cannot be repeated while the file is withheld.

  6. 06
    Transcripts were not de-identified before scoring

    Contrary to protocol.md:27. The sole scorer was the operator who ran the grid. Every filename carries its condition and model, and each carries an operator note stating the verdict.

  7. 07
    The original scoring was mildly biased toward widening the gap

    Stricter than the rubric text inside Arm A, looser inside Arm B. Two cells were overturned on adjudication and both ran that way, both in the baseline's favour.

  8. 08
    The single-scorer cross-check was never run

    protocol.md:30 makes an LLM cross-check the condition of falling back to one scorer. The fallback happened and the cross-check did not.

  9. 09
    Baseline partial-pass counts are not published

    protocol.md:31 requires them prominently, per case, and the rubric calls the full-versus-partial distinction the core of the benchmark.

  10. 10
    The rubric set no integration threshold

    Both scoring passes invented one and disagreed. The rule is now pre-committed: two of three facts with the third absent scores PARTIAL.

  11. 11
    The published rationale for the one near-miss overstates it

    The record says two of three facts were integrated. The adjudication found one.

  12. 12
    The sharpest discriminators in the sealed key went unused

    The Case C falsification element and the explicit NOT clauses in both cases were applied by neither scoring pass.

  13. 13
    The GO case was published with the informed baseline removed

    The site showed 0 of 6 against 6 of 6 and omitted the middle column, which is also 6 of 6.

  14. 14
    The case where the file adds least was never identified

    protocol.md:36 requires it. It is Case B, Larkfield.

  15. 15
    Per-case, per-condition counts are not published

    The rubric names them as the headline unit. Direction is now split by case; the other seven checks are not.

  16. 16
    The alternative reading of the options check was kept internal

    A more lenient reading credits further all-facts runs on the check the scorer calls load-bearing.

  17. 17
    The question-only condition was contaminated on Case B

    ChatGPT invoked live web search unprompted during those runs. Those six runs carry the whole of the direction deficit reported anywhere on this site.

  18. 18
    The 14-day refresh commitment lapsed

    No date-stamped history was published. The dates and the arithmetic are in the status section.

What this benchmark cannot establish.

  1. 01

    No condition in this grid resembles a real operator. The question-only condition has nothing and the all-facts condition has everything. An operator who arrives with most of what matters and a blind spot they cannot see is the condition the product exists for, and it was not run.

  2. 02

    One person ran the grid, scored it, held the sealed reference answers, and answered the file's questions during the runs. The re-score was carried out by three agents working from the published rubric, blind to the original scoring. Neither they nor the adjudication were independent of GreenSquare AI, and this page does not use that word.

  3. 03

    Two published cells are definitional rather than empirical, and a third is partly so. A single-turn condition cannot ask before answering, and it cannot obtain three facts that are not in its prompt.

  4. 04

    The three checks that separate most cleanly are the behaviours the file instructs the model to perform. A reasonable reading is that they measure whether the file was loaded, not whether the model reasoned better.

  5. 05

    The all-facts condition was not information-equal between the two models on Case B. The Claude fact sheet carries two cross-reference pointers the ChatGPT one lacks.

  6. 06

    n is 15 per condition, three repeats per cell. The grid is small and unbalanced, with Case A running on one model only.

  7. 07

    The run inventory reconciles. Both directory counts were verified on disk, 45 grid transcripts and 10 pilot files, and the arithmetic requires 9. The tenth file was opened on 30 August 2026: it is the scoring writeup for the nine excluded pilot runs, and it states in its own opening paragraph that they are not pooled with the pre-registered grid. This entry previously reported the inventory as unreconciled and the tenth file as unopened.

  8. 08

    What was tested is a file fixed by its published hash. Nothing here commits any later release to being that same file.

Inspect what is public, and what remains withheld.

Scored runs
45
Preregistration commit
3e5e46f
Authored
5 July 2026 at 09:10:37 AEST
Transcripts
Not published
Sealed answers
Hash verified, contents not published