Abstract / Summary
A null result and the post-mortem behind a correction. This paper recomputes, from the released records, the inference-time probe in Curated Context, Not Weight Surgery (concept DOI 10.5281/zenodo.21247053). It is the internal re-analysis of 10 August 2026 that the corrected version of that paper (v0.3, 10.5281/zenodo.23124700) refers to, written up in full. The data are public at github.com/saluca-labs/cct. We pre-registered (in a dated plan written before the run, released with the data) a probe testing whether a curated reasoning “mindset”, served as retrieved context, protects a language model against a poisoned reference answer: 6 conditions × 24 tasks on qwen3:32b, 144 generations. The original reading was that a poisoned reference was echoed 45.8% of the time and that the mindset cut this to 33.3%. That finding is an artifact, and correcting it removes the effect. Three defects compound. (i) The “correct” and “echoed” labels are both substring tests over the same 120-character window, so 8 of the 11 flagged echoes on the poison arm are records the same grader also marked correct, in which the model explicitly rejects the poison. (ii) The 1,200-token generation cap truncated 20 of 24 poison-arm generations against 3 of 24 at baseline, so the grader read the tail of an unfinished reasoning trace instead of an answer. (iii) The report code hid a 7.3% false-positive echo rate on the control arms, including a task whose planted wrong answer (“1”) is a substring of its correct answer (“12”). Under four progressively stricter adjudications the protective effect is −12.5, +12.5, +4.2 and +4.2 points. Only the published tier has the protective sign, and nothing is significant at any tier (Fisher p = 0.56, 0.46, 1.00, 1.00; paired exact McNemar p = 0.45, 0.38, 1.00, 1.00). Because most poison-arm generations never produced an answer, the run cannot say whether the poison was adopted at all. The transferable lesson is narrow and concrete: a string grader that falls back to unstructured chain-of-thought measures reasoning about answers, and it fails most on exactly the arm where reasoning is longest. The paper closes with eight inexpensive harness recommendations and a revision note listing what changed from the August draft. Deposited with the recomputation scripts (analysis.py, adjudicate.py, reconcile.py), which take a clone of the data repository as their argument, and the per-record adjudication. Cite the original paper by its concept DOI, which resolves to the corrected latest version.