Abstract / Summary
Much of what the public is told about diet and chronic disease now originates not from the cohorts that built and validated dietary instruments, but from secondary analyses of open datasets by researchers who did not. This matters because the most widely shared dietary exposure data, FFQ responses, carry measurement properties that determine which analyses are valid and which are not. When those properties are ignored, the result is not a marginal loss of precision. It is the systematic generation of findings that look like evidence but are partly or wholly artefacts of how diet was measured [ 1 , 2 ]. A central distinction is well established. FFQs are designed to rank individuals by habitual intake, and they perform acceptably for this purpose [ 1 , 3 , 4 ]. However, they have limited accuracy for estimating absolute intake, with validation studies showing only moderate correlations with reference methods and systematic misestimation [ 5 , 6 ]. Accordingly, FFQs provide relative, not precise quantitative, measures of diet. These limitations have important consequences. Biomarker-based studies show substantial measurement error in FFQ-derived estimates and attenuation or distortion of diet–disease associations [ 2 ]. Self-reported diet is also subject to systematic and differential misreporting, further biasing inference [ 7 ]. Pooled analyses across multiple validation studies using recovery biomarkers confirm that these errors are substantial and systematic across cohorts, rather than idiosyncratic to any single instrument or population [ 8 , 9 ]. Despite this, FFQ variables in open datasets are often analysed as unbiased continuous exposures. This treats ranking-based measures as quantitative ones, disregarding their validation framework and error structure and risks generating findings that reflect properties of the instrument as much as underlying biology.