Abstract / Summary
Part one (new in this version, pre-registered). Forty-nine complete filovirus genomes from nine species were compared against order-2 and order-5 Markov models fitted to each 16,384-base window, with three surrogate replicates per window. Across the 19 genomes frozen as the primary set, all four declared measures depart from the model in the same direction (median differences +0.0657, +0.0172, +0.0520, +0.1707; Holm-corrected p = 1.5e-05). An independent replication on 30 genomes acquired afterwards, with the success rule fixed in writing before acquisition, confirms all four (Holm-corrected p = 7.5e-09). Replication magnitudes are systematically smaller than the primary ones, so the primary estimates are reported as upper bounds. A direct noise measurement, from the spread of the surrogate replicates, ranks the four measures by reproducibility. Part two (pre-registered). Whether the G-quadruplex motifs of the Bundibugyo lineage remained identical between the 2007 and the 2026 outbreaks. Under a strict motif definition the 2007 reference genome contains two motifs; one is identical in all 21 targets, against 139 of 208 control segments identical in all targets. With two motifs the comparison has no power and the outcome is inconclusive. A gap in the design is declared: no minimum number of motifs had been fixed for the primary comparison, only for a secondary one. Withdrawals. The claim made in version 4, that genomic structure was conserved across three outbreaks and nineteen years, is withdrawn, together with the section listing candidate structural targets that rested on it. The underlying density result (GQ-like counts 35.0, 36.0 and 37.6 across the three corpora) stands and is retained; what does not follow is the step from a stable count to conserved structures, and from there to targets. A negative control added in this version also withdraws the MICRA label. Shuffled genomes, which retain length and base composition but have no structure, obtain the same label as the real genomes: nine shuffles out of ten reach the strongest category for the two large genomes, and no pathogen separates from its own shuffles (empirical p from 0.097 to 0.742). The label tracks the number of windows examined rather than structure. Analyses from versions 1 to 4 on CCHF IbAr10200, VZV, PRV and the Marburg clades are retained as descriptive measurements with their validation controls, and are marked as not redone under pre-registration. Limitations. Computational predictions without experimental validation; no clinical conclusions and no proposed therapeutic targets; the source code is not available for external review; the document has not undergone peer review. Chapter 10 lists every substantive error corrected from version 1 onwards, with the version that introduced it and the version that corrected it. Author: Luca Lupinacci, independent researcher, Cosenza, Italy. The author is not a biologist, a geneticist or a computer scientist by training; the software and this report were produced with the assistance of artificial-intelligence systems. Cite with the concept DOI 10.5281/zenodo.22817161, which always resolves to the latest version. ---------------------------------------------------------------------- IMPORTANT: this version supersedes all previous ones. Readers who downloaded any earlier version are encouraged to use this one. CHANGELOG Version 5 - 2 Oct 2026 New pre-registered part on 49 genomes with an independent replication; the MICRA label withdrawn after a negative control; the claim about conservation of structure withdrawn; the section on targets removed; laboratory constructs excluded; reliability measurement added; window overlap declared; errors chapter added; numbering and DOI made consistent. Version 4 - 24 Sep 2026 - 10.5281/zenodo.22936607 MICRA recomputed at 1000 permutations for all 5 pathogens; cross-validation with QGRS Mapper and G4Hunter on BDBV and Marburg; positive controls on c-MYC and human telomere; Hamming distance analysis on the 2007 corpus; complete GC/N/GQ table for the 8 Zaire 2025 genomes; M32 moved to an appendix; genomic files re-verified. Version 3.1 A document revision, not published as a Zenodo version: the "pan-filoviral conserved" region renamed "candidate"; target language revised after independent review; exploratory thresholds declared as pre-defined. Version 3 - 20 Sep 2026 - 10.5281/zenodo.22861938 Substantive corrections: 3 genomic files not matching their accessions replaced; FJ217162 removed from statistical comparisons; MICRA recomputed at 1000 permutations; GQ-like positive control added; M32 removed; "GQ" renamed "GQ-like"; 95% CI for the 2026 corpus corrected to bootstrap percentile. Version 2 - 19 Sep 2026 - 10.5281/zenodo.22846534 Corpus and analyses extended. Version 1.0 - 17 Sep 2026 - 10.5281/zenodo.22817162 First publication. NOTE ON NUMBERING: version 3.1 was a document revision and was never published as a separate Zenodo version; it has no DOI of its own. From version 5 there is a single numbering scheme, no sub-versions, and the document cites only the concept DOI. FILES: the report in English and Italian, the pre-registration with its dated amendments, the verified genome files with checksums, the per-genome measurements, the motif table, the quality-control tables, the negative-control output, and a manifest listing every file with its SHA-256. The source code is not included.