How the current benchmark was run.
The current result covers 54 scored runs across two models, three cases and three test conditions. This page records what was tested, how it was scored, where the protocol changed, and which claims remain open.
The claim under test
The Decision is designed to make an AI model follow a reliable process for one business decision. It should ask for missing facts, test the framing, compare genuine options, mark the basis of important claims, and state what would change the recommendation.
The benchmark tests two related questions. First, does the process hold across models and repeated runs? Second, does asking before advising help the model connect the facts that reveal the underlying decision?
The completed grid
The scored result currently includes:
- Claude Opus 4.8 and ChatGPT GPT-5.5.
- Three fictional business cases.
- Three test conditions for every case and model.
- Three repeats in each cell.
- 54 runs in total, with 18 runs per condition.
The pre-registered protocol targeted about 180 runs, including Gemini and more repeats. The completed grid was reduced to 54 runs. Gemini remains pending. The current website reports the completed 54 run result as an interim two model grid.
The cases
The three cases are fictional and cannot be found through search. Each presents an attractive surface question while placing three facts elsewhere in the case that change the real decision.
- Meridian. A large managed service bid with concentration, regulatory and stranded investment risk.
- Larkfield. A build or acquire choice where the evidence supports the more aggressive acquisition.
- Corvan. A continue or exit decision adapted from a documented Australian post mortem.
Larkfield matters because it tests direction. A useful method should support a bold move when the facts warrant it as readily as it supports caution when the downside is stronger.
The three conditions
- Question only. The model receives the short question a business user might type into a new chat. It answers once. Any questions it asks receive no reply.
- All facts provided. The model receives the same short question plus the full case fact sheet in the first message. It answers once.
- The Decision loaded. The model receives the product file and the short question. It asks its own questions. The operator answers only from the same fact sheet used in the all facts condition.
The all facts condition is the controlled comparison. It shows how well the model reasons when someone already knows which information matters. The loaded condition tests whether the file can draw that information out through questions and then produce the structured brief.
The scorecard
Every run is scored on eight checks fixed before the scored grid began.
Process checks
The product file directly requests these behaviours:
- Asked questions before giving the recommendation.
- Compared real options, including keeping the current position.
- Marked the basis of important claims.
- Named an untested constraint or assumption.
- Stated the conditions that would change the recommendation.
Judgment checks
These checks come from the sealed answer prepared for each case:
- Connected the three relevant facts and surfaced the underlying decision.
- Recommended the same direction as the sealed answer.
- Caught the planted inconsistency or bias.
Pre-registration and scoring
The cases, fact sheets, protocol, rubric and answer key hashes were committed publicly before the scored grid began. The paid product and sealed answer were recorded by hash so their exact versions could be fixed without publishing their contents during the test.
The current grid has one scorer. The scorer used the published rubric and sealed answer. Results are reported as raw counts with the number of runs shown. A second independent score has yet to be completed.
What the evidence supports today
- The loaded file completed all five process checks in 18 of 18 runs.
- It surfaced the underlying decision in 17 of 18 runs.
- The question only condition surfaced the underlying decision in 0 of 18 runs.
- The all facts condition surfaced it in 15 of 18 runs and matched the sealed direction in 18 of 18.
- The loaded condition matched the sealed direction in 18 of 18.
These results support a claim about process consistency and fact gathering across the two tested models. A blind senior reviewer study has yet to test which brief experienced operators would prefer to send to an executive committee.
Transparency still pending
The run transcripts are saved but are not yet published. GreenSquare plans to publish them when the Gemini grid is complete. Until then, readers can inspect the cases, fact sheets, protocol, rubric and pre-registration record, while the reported counts cannot be independently re-scored from the public repository.
Limitations
- Three repeats per cell provide an early reliability signal and leave substantial uncertainty.
- GreenSquare wrote two of the three cases.
- The current grid covers two models. Gemini remains pending.
- One person scored the current runs.
- The senior reviewer study and independent re-score remain planned work.
- AI models change frequently, so every result needs its model version and test date.
Current status
The completed 54 run result is published on the Benchmark page. It is current as at 14 July 2026. Claude Opus 4.8 and ChatGPT GPT-5.5 are complete. Gemini, the independent re-score, the transcript release and the senior reviewer study remain pending.