Abstract / Summary
The Frozen Token: A Predictive Theory of a Single-Token Encoder as the Always-On Tier-0 Layer Across Sensing Modalities, Validated Out-of-Sample Randolph James Ferlic, M.D. and Kimberly Kate Ferlic — Fieldstone Analytics, LLC, Austin, TX, USA Preprint · Zenodo DOI: 10.5281/zenodo.23002239 · CC-BY 4.0 · Community: spiral-domain-encoder-campaign · a characterization of previously filed and published methods; no new algorithmic subject matter is disclosed. Abstract Over a program of pre-registered studies we mapped one object — a frozen, deterministic, class-discriminant single-token encoder — as the always-on, sub-milliwatt Tier-0 layer beneath the agentic edge, across radar micro-Doppler, always-on audio, wrist photoplethysmography, industrial vibration, electrocardiography, and a ten-dataset edge survey. This capstone consolidates those results into an explicit predictive theory: four falsifiable laws for where the frozen token wins, pays, and fails. L1 (competence-zone): near-parity on per-channel time/frequency signals, failure on spatial/relational signals; difficulty appears as a priced tax or a low ceiling. L2 (per-entity-necessity): a self-twin is needed iff the anomaly is entity-specific, not for a universal extreme — a dissociation confirmed on five modalities. L3 (composition/portability): compose across orthogonal channels (free), but re-commission a codebook per sensor. L4 (cost-not-accuracy): the advantage is cost, determinism, auditability, and privacy, at an honest bounded tax. To test whether the laws predict rather than describe, we pre-registered their implications for a modality the encoder was never built on — a chemical gas-sensor e-nose — before computing any result, and confirmed all five predictions: near-parity 6-gas classification (token 0.986 vs full 0.982), high absolute accuracy, a purpose-built drift benchmark on which an early-trained codebook collapses (0.98 → 0.37) and re-commissioning fully recovers (→ 0.98, exactly L3), a label-free gate (AUROC 0.974), and a deterministic int8-free cost profile. A frozen single-token encoder is therefore a predictable instrument whose behavior a partner can forecast from the structure of their task — which is what an always-on Tier-0 layer, and an acquirable technology, should be. This is a characterization of previously described, filed methods; it discloses no new algorithmic subject matter, and the per-deployment / per-sensor selection of configuration is retained as trade secret. Highlights · Four falsifiable laws for a frozen single-token encoder — competence-zone (L1), per-entity-necessity (L2), composition/portability (L3), cost-not-accuracy (L4) — each supported by cross-modality evidence from the estate. · Out-of-sample validation on a modality outside the program (chemical e-nose): 5/5 pre-registered predictions confirmed, including a purpose-built sensor-drift benchmark that behaves exactly as the re-commission law requires (0.98→0.37 under drift; re-commission → 0.98). · The laws predict, not merely describe — the strong test of a characterization: pre-registered before analysis, confirmed on a new modality. · A consolidated product / acquisition posture: cost, determinism, auditability, and privacy, with a one-token capability platform (one encode → N heads, energy ~1/M), filed mechanisms, and a retained trade-secret configuration layer. · Honest-map methodology — pre-registration, grouping-disjoint splits, permutation nulls, bootstrap intervals, and anchoring negatives — makes the laws credible. What this record contains · Manuscript_Paper53.pdf — the capstone manuscript (4 figures, 59 references); and Manuscript_Paper53.docx, the editable source. · PAPER_53_ZENODO_ARCHIVE.zip — the reproducibility archive (md5 in ARCHIVE_MD5.txt): the frozen pre-registration (predictions registered before analysis); the out-of-sample e-nose runner + Modal fetch; the frozen token encoder module; the e-nose result record (JSON); the 4 figures; the figure builder; the manuscript source; and a README. Public data is not redistributed (fetched at run time); all paths and handles are scrubbed (PATH_TO_DATA/PATH_TO_SCRATCH, any cloud handle → MODAL_USER) and leak-scanned. Cite as R. J. Ferlic and K. K. Ferlic, "The frozen token: a predictive theory of a single-token encoder as the always-on Tier-0 layer across sensing modalities, validated out-of-sample," Zenodo, 2026, doi: 10.5281/zenodo.23002239. License and patent notice Released under CC-BY 4.0. Consistent with that license, no patent or other intellectual-property right of the authors is licensed, waived, or conveyed by this deposit. This work characterizes previously described, filed, or published methods and discloses no new algorithmic subject matter; gradient boosting, k-means / vector and product quantization, Fisher discriminant analysis, the information bottleneck, temperature scaling, autoencoder / isolation-forest / one-class anomaly detection, decision-level sensor fusion, and nearest-centroid novelty detection are established prior art, used only as tools. The methods characterized — the class-discriminant single-token codebook encoder and its nearest-centroid monitor, the multi-token / token-ladder and soft-readout mechanisms, the inference-time co-channel fusion, the foundation codebook, and the per-entity self-twin — are the subject of filed and pending U.S. patent applications held by the authors, including U.S. Provisional Application No. 64/095,354 (the encoder), the personalization / on-device-adaptation applications (priority U.S. Application No. 19/467,303 and its continuations), the multi-token / token-ladder application (No. 64/119,487), and the inference-time fusion application (No. 64/137,805). The four predictive laws stated here are properties of the frozen filed method, not new subject matter. The per-deployment / per-sensor selection of configuration and commissioning procedure is retained as a trade secret and is not disclosed here. © 2026 Fieldstone Analytics, LLC and the authors. Inquiries: randolphf@fieldstoneanalyticsllc.com. Companion deposits (spiral-domain-encoder-campaign) · The mmWave Micro-Doppler Tier-0 Gate (radar micro-Doppler map; the universal-extreme fall that anchors L2): doi:10.5281/zenodo.22999908 · The Wrist-PPG Tier-0 Gate (wearable physiological map; the entity-specific pole that completes L2, and the difficulty/composition evidence): doi:10.5281/zenodo.23000837 · The Acoustic Tier-0 Gate (always-on-audio map; machine-health self-twin): doi:10.5281/zenodo.22968134 · Where the Cheapest Token Wins, Pays, and Fails (the ten-dataset edge map; the EEG spatial-relational failure of L1): doi:10.5281/zenodo.22945419 · The Configurable Bottleneck (the platform / composition dial-map): doi:10.5281/zenodo.22884023 · Paying Down the Price of the Bottleneck (label-efficiency / foundation codebook): doi:10.5281/zenodo.22866039 · The Price of the Bottleneck (deployment / drift characterization): doi:10.5281/zenodo.22838118 · The Cheapest Digital Twin (per-entity self-twin personalization; the ECG entity-specific evidence for L2): doi:10.5281/zenodo.22922678 · The Predictive Reach of a Decision Token: doi:10.5281/zenodo.22736921 · Non-invertible but not anonymous (privacy): doi:10.5281/zenodo.22819210 · Unlinkable but not anonymous (privacy): doi:10.5281/zenodo.22838120 · Label-free inference-time channel fusion (composition mechanism for L3): doi:10.5281/zenodo.22046713 · Deterministic multi-token token ladder: doi:10.5281/zenodo.22003179 · Class-discriminant single-token codebook construction (the base encoder): doi:10.5281/zenodo.20788187 · The token as a bounded, threshold-free generative-edge cache key: doi:10.5281/zenodo.22148612 Keywords edge AI; TinyML; always-on sensing; Tier-0; single-token encoder; vector quantization; k-means codebook; nearest-centroid classifier; prototypical networks; information bottleneck; self-twin; digital twin; per-entity personalization; label-free anomaly detection; novelty detection; sensor drift; concept drift; domain shift; re-commissioning; cross-modality transfer; sensor fusion; capability composition; multiplexing; energy efficiency; performance-per-watt; neural processing unit; hierarchical inference; cascade classifiers; determinism; bit-exact inference; auditability; privacy; re-identification; predictive theory; pre-registration; out-of-sample validation; falsifiability; honest negatives; reproducibility; radar micro-Doppler; photoplethysmography; electrocardiography; acoustic scene classification; machine-health monitoring; gas sensor array; electronic nose; chemical sensing; cost-accuracy tradeoff; Pareto frontier; reproducible research