Abstract / Summary
Abstract Federated learning (FL) allows hospitals to collaboratively train deep learning models without exchanging patient images, but reported FL studies in medical imaging often assume independent and identically distributed (IID) client data and full participation in every communication round both conditions seldom hold in real clinical networks. Furthermore, although the entire premise of FL rests on the iterative averaging of client weight updates, the optimisation dynamics of those weights are rarely audited at the round level, limiting clinicians' and regulators' ability to inspect how the global model evolves under stress. This study presents an auditable federated framework for binary brain-tumour classification from a combined multimodal dataset of 4618 axial computed-tomography (CT) and 5000 magnetic-resonance imaging (MRI) slices, totalling 9618 labelled images. A lightweight three-block convolutional neural network is trained across five simulated clients using FedProx with proximal coefficient µ = 0.01 and FedAvg aggregation, under two contrasting client partitions: an IID scenario in which each client receives a balanced sample of healthy and tumour cases, and a deliberately aggressive non-IID scenario in which each client is assigned an 80% majority class. Clinically realistic stress is added by dropping a randomly chosen client in approximately 30% of rounds. At every round, for every active client, the framework records three weight summaries initial (received from the server), updated (after local training), and aggregated (new global model) together with per-client accuracy, loss, precision, recall, F1-score, communication payload, wall-clock round time, and process memory. After ten communication rounds, the global model attains 96.05% accuracy and an F1-score of 0.9641 on the held-out test set under the IID partition, and 94.96% accuracy with F1 0.9547 under the 80% label-skew non-IID partition, corresponding to an IID-to-non-IID degradation of only 1.09 percentage points despite the severe class skew and 30% per-round client dropout. Inspection of the weight-audit trace reveals a consistent and measurable shift in the output-layer mean under non-IID training that is invisible in conventional accuracy curves, suggesting that round-level weight auditing can serve as a lightweight forensic signal of distributional drift in federated medical AI.