Abstract / Summary
Abstract Colorectal cancer is a leading cause of cancer death worldwide, even though colonoscopy is an effective screening tool. Automated analysis of endoscopic images can help doctors find more lesions and make faster decisions, but current systems usually handle just one task, are tested on clear images from a single dataset, and are rarely compared against working clinicians. We propose HierTAC-Net, a six-stage convolutional classifier that combines inverted residual blocks, triple attention at three resolution levels, and ConvNeXt-style refinement. We train the model once on the eight-class Kvasir-v2 dataset, and from a single softmax output, it generates seven clinical readouts without retraining, including full diagnosis, anatomical landmark checks, pathological classification, and polyp screening. In stratified five-fold cross-validation, HierTAC-Net achieves 94.24% accuracy with a 110.14 MB footprint and 0.32 ms GPU latency, outperforming the strongest compact attention baseline while remaining far lighter than large fine-tuned models. Our ablation study isolates the contribution of attention placement from the operator itself and shows that early insertion with triple attention provides the largest consistent gain. We further evaluate the model under compounded blur, uneven illumination, and specular highlights at three severity levels, and on an aligned eight-class subset of HyperKvasir without fine-tuning, where it retains 95.98% accuracy and 96.00% F1-score. In a supervised reading session, HierTAC-Net outperforms five gastroenterologists with 1 to 8 years of experience on a shared 200-image subset. These results indicate that the proposed architecture is suitable for real-time deployment and remains stable under common artifacts and across two clinical collections.