Abstract / Summary
Abstract Background. A variable associated with both the outcome and a candidate gene moves that candidate's estimate, most for candidates most strongly correlated with it, so a screen of many correlated candidates is exposed wherever one axis of biology moves the candidates and the outcome together. Tumor proliferation is the clearest such case. How often it is omitted, and what omitting it costs, are questions about analytic practice. Methods. Fifty recent prognostic gene expression papers in hepatocellular carcinoma were sampled, and the 41 with a reachable full text were coded for whether their model adjusted for proliferation. For every expressed gene I then fitted two Cox models for overall survival, one adjusted for age, sex and stage, one adding a standardized proliferation score. The shift is the ratio of the two hazard ratios; the landscape is its log regressed on the gene's correlation with the score. That measurement is exploratory. A cross-cancer test of two predictions, with its inclusion rule and refutation thresholds, was registered in advance. Results. None of those 41 adjusted for proliferation (0 percent, 95 percent CI 0.0 to 8.6), 23 reported proliferation activity in their high-risk group, and 28 did not state their covariates in the text. In hepatocellular carcinoma the landscape slope was -0.278 (R-squared 0.785), reproduced at -0.180, and the genes surviving false discovery control fell from 338 to 1. Ten cohorts met the registered rule; two were aggregates, so the eight sharing no patients are primary. The prediction that the slope would be negative in at least 90 percent of them was refuted, at four of eight, as was the prediction of gene-level agreement. Post hoc, a first-order omitted variable expression explained 0.84 of the variance of the per-gene shift, short of the 0.95 fixed in advance. Conclusions. Adjusting for a covariate associated with both the outcome and the candidates moves nearly every estimate in a high-dimensional screen, by an amount the candidate's correlation with it predicts. Whether that removes bias depends on a causal structure these data cannot establish. Its direction tracks that association in the cohort at hand, which my registered prediction assumed away.