Abstract / Summary
Generative data augmentation often triggers mode collapse: iterative training on synthetic data causes models to drift from the true distribution, degrading performance. We argue the root cause is not insufficient realism but the erosion of discriminative information across generation rounds. From a hierarchical decomposition perspective, samples comprise discriminative features, intra-class variations, and noise. We introduce “data intrusion” to characterize this erosion. Accordingly, we propose a decoupled generation framework following a “decompose-then-generate” strategy: discriminative features are identified and frozen via identity mapping, while diversity is expanded only in the surface feature space. This structurally ensures that Discriminative Feature Purity (DFP) is identically 1. To avoid the generative proof paradox, we design a two-stage validation paradigm with four metrics (DFP, SFD, NPG, DBS). On five structured datasets spanning medical, financial, cybersecurity, and industrial domains, our method maintains discriminative consistency generally above 0.85 (reaching 0.98 on ECG), significantly outperforming competitors (approx. 0.5–0.7). Downstream classification remains stable, with gains up to 16.4% on sample-scarce industrial data and 4.2% on ECG, where most competing methods suffer negative gains. A controlled-correlation simulation delineates an applicability boundary at ρ≈0.8, within which all real-world datasets fall. Centered on discriminative protection, this paper offers a verifiable, controllable, and cross-domain-robust alternative for generative data augmentation.