Controlled Comparison of Multimodal Hybrid Deep Learning for Benign-Malignant Mammogram Classification on CBIS-DDSM
Keywords:
breast cancer, mammography, transfer learning, clinical metadata fusion, attention mechanism, explainable AI, patient-wise splitAbstract
Breast cancer is the leading cause of cancer death among women, and mammography remains the first line of screening. Many deep learning studies on CBIS-DDSM report accuracy above 95%, yet they partition the archive at image level, so views of one patient land in both training and test sets and the figures cannot be trusted. This study builds a controlled comparison of two image representations, the whole mammogram and the lesion region of interest, under one identical training recipe. Six models are trained on each representation: EfficientNet-B3, ConvNeXt-Tiny, a gated hybrid fusion of the two with the Convolutional Block Attention Module, two variants that fuse seven clinical descriptors through a perceptron, and a Mammo-CLIP backbone pretrained on mammograms. The official CBIS-DDSM partition is enforced at patient level and yields 2,148 training, 261 validation and 448 test mammograms from 1,104, 122 and 234 disjoint patients, and the BI-RADS assessment is dropped to avoid circular prediction. On the test set the strongest model, Hybrid Fusion+CBAM+Clinical, reaches an AUC of 0.856 at an accuracy of 78.1% on regions of interest, and an AUC of 0.845 at an accuracy of 78.4% on whole images. Clinical metadata is the largest single source of gain, close to six AUC points on both representations. After its BatchNorm adaptation is corrected, Mammo-CLIP recovers to an AUC of 0.789 and becomes the best image-only whole-image backbone. Grad-CAM shows that attention settles on the lesion rather than on artifacts, and the reported range is what remains once leakage is removed.




