GreenSquare AI

How the benchmark separated access to information from the discipline used on it.

Three conditions on the same frozen cases, so that the informed baseline holds the same facts as the loaded condition and only the method differs.

Evidence status, methodologyReviewed 10 August 2026
Claim tested
That the design can tell a change in process from a change in judgement, and that the informed baseline is a fair comparison for the loaded method.
Result
The design does separate them: process checks are scored from the transcript; judgement checks against a sealed key. The comparison held on information but not on turns, so two process cells are definitional.
Read with
Caution. The protocol’s scoring safeguards (de-identification, two scorers, an LLM cross-check) were not followed, and each lapse is recorded as one of the 18 deviations.
Still open
Independent re-scoring, the senior-reviewer study and the third preregistered model.
Product relevance
The method file under test predates Frame. The design says nothing about Frame Free or Frame Pro.

The claim under test

The loaded method is designed to make an AI model follow a disciplined process for one consequential business decision. It should ask for missing facts, test the framing, compare credible options, mark the basis of material claims and state what would change the recommendation.

The benchmark asks two questions. Does that process hold across models and repeated runs? Does asking before advising help the model connect the facts that reveal the underlying decision? It does not test whether experienced operators prefer the resulting brief.

The frozen study artefact used earlier internal product names. This page calls that arm the loaded condition. The study is not a test of the current Frame Free plan or the future Frame Pro plan.

The completed grid

  • 45 scored runs, with 15 per condition.
  • Two tested models, Claude Opus 4.8 and ChatGPT GPT-5.5; three fictional business cases; three repeats in each completed cell.
  • Three conditions: question only, all facts provided, and the loaded method.

The preregistered protocol targeted about 180 runs, including Gemini and more repeats. The completed grid was reduced to 54 runs, then re-cut to 45 by removing nine design-phase pilot runs that the protocol said would not be pooled. Gemini was not run.

The cases

The three cases are fictional and cannot be found through search. Each begins with an attractive surface question while placing three facts elsewhere in the case that change the real decision.

  • Meridian. A managed-service bid with concentration, regulatory and stranded-investment risk.
  • Larkfield. A build-or-acquire choice where the evidence supports the more aggressive acquisition.
  • Corvan. A continue-or-exit decision adapted from a documented Australian post-mortem.

The three conditions

  • Question only. The frozen prompt only. One turn. If the model asks questions they are not answered.
  • All facts. The frozen prompt with the full fact sheet appended in the first message. One turn. Equal information to the loaded condition.
  • Loaded method. The frozen method file, then the frozen prompt. The model interrogates; the operator answers only from the fact sheet and replies 'That has not been assessed' to anything off-sheet.

The all-facts condition is the informed baseline. It shows how well the model reasons when someone already knows which information matters. The loaded condition tests whether the method can draw that information out through questions and then produce the structured brief.

The scorecard

Every run is scored on eight checks fixed before the scored grid began. Process checks are scored from the transcript against published one-paragraph operationalisations. Judgement checks are scored against the sealed reference answer for the case, which the method file does not name.

Process checks

  • Asked before answering.
  • Real options incl. do-nothing, compared on criteria.
  • Evidence tagged (given, derived, inferred, unknown).
  • Named an untested constraint or assumption.
  • Stated kill criteria.

Judgement checks

  • Surfaced the hidden decision by integrating the facts.
  • Recommendation direction matched the sealed key.
  • Caught the planted inconsistency or bias.

Preregistration and scoring

The cases, fact sheets, protocol, rubric and answer-key hashes were committed publicly before the scored grid began, at commit 3e5e46f, authored 5 July 2026 at 09:10:37 AEST. The method file and sealed answer were recorded by hash so their exact versions could be fixed without publishing their contents during the test.

The current grid has one scorer. Five disputed cells were adjudicated against the sealed reference answers on 10 August 2026. A second independent score has not been completed. The protocol called for de-identified transcripts, two scorers and an LLM cross-check; none of the three happened, and each is recorded on the benchmark page as a deviation.

Inspect the frozen preregistration

Limitations of the design

  • Three repeats per cell provide an early reliability signal and leave substantial uncertainty.
  • GreenSquare AI wrote two of the three cases.
  • The current grid covers two models. A third preregistered model was not run.
  • One person ran and scored the grid and held the sealed answers. The adjudication was also internal.
  • The senior-reviewer study and independent re-score remain planned work.
  • Model behaviour changes, so the evidence needs periodic replication.

The results support a claim about process consistency and fact gathering across the tested conditions. They do not support a claim that the method improves an experienced operator’s judgement.

Read the benchmark record