# Correction transfer pilot — selected evidence

3 October 2026. This is a real model experiment using a synthetic energy-meter problem. It is separate from the later development inquiry shown at [/research](/research).

Three conditions received the same task: 21 versioned language records and 27 relations as JSON, equivalent prose, or no extra language package. Each condition ran twice. Each of the six pipelines used four fresh participant calls with explicitly reconstructed context: A1, B1, A2 and B2. The corrected measurement was supplied to A2; only its admitted notification reached B2. There was no A3 return call. All conditions received useful common instructions.

The requested model was gpt-6-astra with xhigh effort. The resolved backend identity was not independently exposed. A separate Codex grader read anonymized outputs; scores were locked before the condition mapping was opened. This was not human expert or Fable review.

| Condition | Repetition 1 | Repetition 2 |
|---|---:|---:|
| JSON | 16/16 | 16/16 |
| Equivalent prose | 16/16 | 16/16 |
| No extra package | 15/16 | 15/16 |

All six pipelines transmitted the correction and revised the recipient's local decision. The grader found no listed critical errors. The one-point deductions concern not explicitly stating that additional unknown dependencies could remain outside coverage. Both baseline responses identified the known missing dependency correctly and did not claim complete coverage. A more permissive reading would give both 16/16. The original deductions and alternative readings remain on record.

**No behavioral superiority of JSON over prose was established.** The pilot also does not establish a reliable advantage of the package, durable learning, independent discovery, value alignment or general intelligence. One scenario with two repetitions and dependent stages is a limited comparison. Both roles' initial investigation threshold was the same; a difficult disagreement between local policies was not tested.

The original runner misclassified two CLI startup notices as tool calls. The six first A1 responses were retained, and the classification was amended before the remaining 18 calls. There were no participant quality retries. The changed chronology and scope of the exception are recorded in the amendment below.

## Inspect the evidence

- [Frozen protocol](./PROTOCOL.md), [task](./TASK.md), [rubric](./RUBRIC.md), [fixture](./fixture.json).
- [Condition results](./RESULTS.json), [blinded review](./masked-review/REVIEW.md), [detailed scores and alternatives](./masked-review/SCORES.json), [mapping opened after the score lock](./BLIND-MAPPING.json).
- [Shared reference](./masked-review/REFERENCE.md) and raw participant envelopes with masked labels: [01](./masked-review/case-01.json), [02](./masked-review/case-02.json), [03](./masked-review/case-03.json), [04](./masked-review/case-04.json), [05](./masked-review/case-05.json), [06](./masked-review/case-06.json).
- [Execution amendment](./AMENDMENT-01.md) and its [separate methods review](./AMENDMENT-01-REVIEW.md).
- [Hashes and publication scope](./provenance.json). These selected artifacts are byte-identical to the preserved local experiment archive. This reading package excludes raw CLI logs and provider metadata; it is not the complete execution archive or an independent replication.

The development inquiry uses this result to separate an engineering choice—versioned, inspectable records—from the research question of whether a representation improves decisions. Its later reviews and return messages are new development work, not additional participants or stages in this pilot.
