Reproducibility and Generalization Analysis of OCAT Net S for Oral Cancer Image Classification
Main Article Content
Abstract
Clinical photographs could support earlier referral of suspicious oral lesions, but high internal accuracy does not show whether a model will remain reliable when image source, disease prevalence, or lesion definition changes. This study conducted a secondary experimental analysis of the aggregate measurements reported for OCAT-Net-S by [1]. The analysis examined internal discrimination, component ablation, synthetic perturbation tolerance, computational efficiency, probability calibration, threshold behavior, fold-level dependence, leakage controls, and external transfer. No new patient images or model-training runs were introduced. All new results in this paper are derived from the published tables. OCAT-Net-S achieved an internal AUC-ROC of 0.956, sensitivity of 0.942, and oral-cancer F1-score of 0.942 on 1,238 clinical oral images. Its AUC advantage over TinyViT-5M, the strongest evaluated baseline, was 0.015. The external AUC fell to 0.830 on SMART-OM and 0.840 on the Oral Images Dataset, producing absolute generalization gaps of 0.126 and 0.116. These gaps were 8.40 and 7.73 times larger than the internal advantage over TinyViT-5M. Structural ablations caused a mean AUC loss of 0.046, compared with 0.0128 for training-related ablations. Under the most severe Gaussian-noise condition, OCAT-Net-S retained 96.3% of its clean AUC, compared with 93.8% for TinyViT-5M and 90.3% for ResNet-50. The model used 3.55 million parameters and achieved 0.269 AUC per million parameters, although MobileNetV3-Large remained faster and required fewer operations. Calibration was favorable internally, with a Brier score of 0.0419 and ECE-15 of 0.0174. At an assumed 3% screening prevalence, however, estimated positive predictive value ranged from 0.213 to 0.420 across reported operating points. The findings support the internal technical value of texture-frequency spatial gating and compact local-global feature learning. They also show that distribution shift and prevalence exerted a larger practical effect than the margin separating OCAT-Net-S from its strongest internal comparator. Patient-linked multicenter evaluation, fixed external thresholds, prospective workflow testing, and quantitative lesion localization remain necessary before clinical use.
Article Details

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
References
[1] M. H. Rijvi et al., “A compact texture frequency spatial gating network for oral cancer classification from clinical oral images,” Sci. Rep., 2026, doi: 10.1038/s41598-026-59480-0.
[2] A. L. D. Araújo, C. M. Pedroso, P. A. Vargas, M. A. Lopes, and A. R. Santos-Silva, “Advancing oral cancer diagnosis and risk assessment with artificial intelligence: a review,” Explor. Digit. Heal. Technol., vol. 3, p. 101147, 2025, doi: 10.37349/edht.2025.101147.
[3] M. F. Kabir, M. Y. Ahmad, R. Uddin, M. Cordero, and S. Kant, “Accurate and lightweight oral cancer detection using SE-MobileViT on clinically validated image dataset,” Discov. Artif. Intell., vol. 5, p. 173, 2025, doi: 10.1007/s44163-025-00442-2.
[4] D. P. Yadav, B. Sharma, A. Noonia, and A. Mehbodniya, “Explainable label guided lightweight network with axial transformer encoder for early detection of oral cancer,” Sci. Rep., vol. 15, no. 1, p. 6391, 2025, doi: 10.1038/s41598-025-87627-y.
[5] P. Liu and K. Bagi, “A tailored deep learning approach for early detection of oral cancer using a 19-layer CNN on clinical lip and tongue images,” Sci. Rep., vol. 15, no. 1, p. 23851, 2025, doi: 10.1038/s41598-025-07957-9.
[6] F. R. D. Cruze, J. Wasima, M. F. Hosen, M. B. A. Miah, Z. Muhammad, and M. F. Al Masud, “Oral Cancer Diagnosis Using Histopathology Images: An Explainable Hybrid Transformer Framework,” Technologies, vol. 14, no. 1, p. 39, 2026, doi: 10.3390/technologies14010039.
[7] A. Hossain et al., “Transformer-Based Ensemble Model for Binary and Multiclass Oral Cancer Segmentation,” 2025 International Conference on Electrical, Computer and Communication Engineering (ECCE). IEEE, 2025. doi: 10.1109/ECCE64574.2025.11012921.
[8] J. Štifanić, D. Štifanić, N. Anđelić, and Z. Car, “Explainable AI for Oral Cancer Diagnosis: Multiclass Classification of Histopathology Images and Grad-CAM Visualization,” Biology (Basel)., vol. 14, no. 8, p. 909, 2025, doi: 10.3390/biology14080909.
[9] P. Kaushik, V. Kukreja, E. Jain, M. Ati, and S. Hariharan, “Oral tumor detection and localization using detection transformer (DETR) with multi-scale feature extraction and self-supervised learning,” Array, vol. 30, p. 100825, 2026, doi: 10.1016/j.array.2026.100825.
[10] C. Karmakar et al., “Hierarchical lesion-aware transformer for oral cancer image classification,” Discov. Artif. Intell., 2026, doi: 10.1007/s44163-026-01776-1.
[11] T. R. Sikder et al., “XACT-TB an explainable hybrid CNN–swin transformer framework for tuberculosis screening from chest X-ray images,” Discov. Artif. Intell., vol. 6, p. 1030, 2026, doi: 10.1007/s44163-026-01509-4.
[12] M. I. Faruk et al., “Channel attention recalibrated fusion of CoAtNet and frozen DINOv2 features for binary chest X ray classification,” Discov. Comput., vol. 29, p. 551, 2026, doi: 10.1007/s10791-026-10446-w.
[13] P. D. Madan Kumar et al., “A Smartphone-based Comprehensive Dataset of Annotated Oral Cavity Images for Enhanced Oral Disease Diagnosis,” Sci. Data, vol. 13, p. 676, 2026, doi: 10.1038/s41597-026-06954-5.