Abstract / Summary
Abstract Over the past decade, machine learning (ML) has established a transformative paradigm for investigating amyloidogenic polypeptides. Characterized by their propensity to form ordered, β-sheet-rich fibrillar aggregates, these polypeptides are centrally implicated in the pathogenesis of numerous neurodegenerative diseases, including Alzheimer’s disease (AD) and Parkinson’s disease (PD). This review systematically evaluates the application of ML across critical domains of amyloid investigation: sequence-based amyloidogenicity prediction, structural biology, inhibitor design, and amyloid-related clinical objectives. We further summarize the diverse array of ML architectures currently employed in amyloid research, spanning classical algorithms (e.g., support vector machines and random forests) to advanced deep learning frameworks, such as the transformer framework, protein language models (PLMs), etc. Additionally, we spotlight emerging frontiers, namely, the integration of ML with molecular dynamics (MD) simulations, the use of generative AI for de novo inhibitor design, etc., while addressing persistent bottlenecks in the field, including data standardization, model interpretability, and rigorous experimental validation for in silico predictions. Ultimately, this review serves as a foundational roadmap for emerging researchers and a definitive reference for established investigators operating at the intersection of artificial intelligence and amyloid biology.