Abstract / Summary
Large language models (LLMs) are often evaluated from isolated responses or final answers, although high-consequence use unfolds across interaction trajectories. We formalized trajectory integrity as a non-scalar profile separating endpoint recovery, criterion recovery, factual-state recovery, and retrospective fidelity, and applied it to nine LLM systems completing a simulated allocation of one intensive care unit (ICU) bed between two patients. Nine systems were evaluated in Spanish, English, German, and Simplified Chinese under four conditions, yielding 16 trajectories per system and 144 nested trajectories. Each trajectory elicited an allocation, governing criterion, prospective reversal threshold, responses to contextual pressures, an explicit reset, and retrospective audits. The model was the primary cross-model unit. Endpoint recovery after explicit reset occurred in 144/144 trajectories, but every system had at least one adjudicated trajectory with non-integral criterion recovery. Full factual recovery occurred in 132/144 trajectories; 12 were condition-specific, partial, or ambiguous. No system produced a fully faithful structured retrospective checksum across all 16 trajectories. Forced retrospective audits were completed in 144/144 trajectories, reflecting format compliance rather than spontaneous contradiction. A proposed composite metric failed semantic and arithmetic comparability across systems, supporting use of a profile rather than a single score. These results show that endpoint agreement can coexist with differences in recovered criterion, factual state, and retrospective account. The analysis is descriptive: each experimental cell contains one generation, the scenario is simulated, and several interpretive variables lack complete independent double-coding.