How the benchmark separated access to information from the discipline used on it.
Methodology. Preregistered 5 July 2026; record current as at 10 August 2026.
Three conditions on the same frozen cases, so that the informed baseline holds the same facts as the loaded condition and only the method differs.
The claim under test
The loaded method is designed to make an AI model follow a disciplined process for one consequential business decision. It should ask for missing facts, test the framing, compare credible options, mark the basis of material claims and state what would change the recommendation.
The benchmark asks two questions. Does that process hold across models and repeated runs? Does asking before advising help the model connect the facts that reveal the underlying decision? It does not test whether experienced operators prefer the resulting brief.
The frozen study artefact used earlier internal product names. This page calls that arm the loaded condition. The study is not a test of the current Frame Free plan or the future Frame Pro plan.
The completed grid
- 45 scored runs, with 15 per condition.
- Two tested models, Claude Opus 4.8 and ChatGPT GPT-5.5; three fictional business cases; three repeats in each completed cell.
- Three conditions: question only, all facts provided, and the loaded method.
The preregistered protocol targeted about 180 runs, including Gemini and more repeats. The completed grid was reduced to 54 runs, then re-cut to 45 by removing nine design-phase pilot runs that the protocol said would not be pooled. Gemini was not run.
The cases
The three cases are fictional and cannot be found through search. Each begins with an attractive surface question while placing three facts elsewhere in the case that change the real decision.
- Meridian. A managed-service bid with concentration, regulatory and stranded-investment risk.
- Larkfield. A build-or-acquire choice where the evidence supports the more aggressive acquisition.
- Corvan. A continue-or-exit decision adapted from a documented Australian post-mortem.
The three conditions
- Question only. The frozen prompt only. One turn. If the model asks questions they are not answered.
- All facts. The frozen prompt with the full fact sheet appended in the first message. One turn. Equal information to the loaded condition.
- Loaded method. The frozen method file, then the frozen prompt. The model interrogates; the operator answers only from the fact sheet and replies 'That has not been assessed' to anything off-sheet.
The all-facts condition is the informed baseline. It shows how well the model reasons when someone already knows which information matters. The loaded condition tests whether the method can draw that information out through questions and then produce the structured brief.
The scorecard
Every run is scored on eight checks fixed before the scored grid began. Process checks are scored from the transcript against published one-paragraph operationalisations. Judgement checks are scored against the sealed reference answer for the case, which the method file does not name.
Process checks
- Asked before answering.
- Real options incl. do-nothing, compared on criteria.
- Evidence tagged (given, derived, inferred, unknown).
- Named an untested constraint or assumption.
- Stated kill criteria.
Judgement checks
- Surfaced the hidden decision by integrating the facts.
- Recommendation direction matched the sealed key.
- Caught the planted inconsistency or bias.
Preregistration and scoring
The cases, fact sheets, protocol, rubric and answer-key hashes were committed publicly before the scored grid began, at commit 3e5e46f, authored 5 July 2026 at 09:10:37 AEST. The method file and sealed answer were recorded by hash so their exact versions could be fixed without publishing their contents during the test.
The current grid has one scorer. Five disputed cells were adjudicated against the sealed reference answers on 10 August 2026. A second independent score has not been completed. The protocol called for de-identified transcripts, two scorers and an LLM cross-check; none of the three happened, and each is recorded on the benchmark page as a deviation.
Limitations of the design
- Three repeats per cell provide an early reliability signal and leave substantial uncertainty.
- GreenSquare AI wrote two of the three cases.
- The current grid covers two models. A third preregistered model was not run.
- One person ran and scored the grid and held the sealed answers. The adjudication was also internal.
- The senior-reviewer study and independent re-score remain planned work.
- Model behaviour changes, so the evidence needs periodic replication.
The results support a claim about process consistency and fact gathering across the tested conditions. They do not support a claim that the method improves an experienced operator’s judgement.
Read the benchmark record