Reproducibility and Generalization Analysis of OCAT Net S for Oral Cancer Image Classification

Main Article Content

Tajul Islam Rafi
Md Abu Kawsar Prodhan Hemal

Abstract

Clinical photographs could support earlier referral of suspicious oral lesions, but high internal accuracy does not show whether a model will remain reliable when image source, disease prevalence, or lesion definition changes. This study conducted a secondary experimental analysis of the aggregate measurements reported for OCAT-Net-S by [1]. The analysis examined internal discrimination, component ablation, synthetic perturbation tolerance, computational efficiency, probability calibration, threshold behavior, fold-level dependence, leakage controls, and external transfer. No new patient images or model-training runs were introduced. All new results in this paper are derived from the published tables. OCAT-Net-S achieved an internal AUC-ROC of 0.956, sensitivity of 0.942, and oral-cancer F1-score of 0.942 on 1,238 clinical oral images. Its AUC advantage over TinyViT-5M, the strongest evaluated baseline, was 0.015. The external AUC fell to 0.830 on SMART-OM and 0.840 on the Oral Images Dataset, producing absolute generalization gaps of 0.126 and 0.116. These gaps were 8.40 and 7.73 times larger than the internal advantage over TinyViT-5M. Structural ablations caused a mean AUC loss of 0.046, compared with 0.0128 for training-related ablations. Under the most severe Gaussian-noise condition, OCAT-Net-S retained 96.3% of its clean AUC, compared with 93.8% for TinyViT-5M and 90.3% for ResNet-50. The model used 3.55 million parameters and achieved 0.269 AUC per million parameters, although MobileNetV3-Large remained faster and required fewer operations. Calibration was favorable internally, with a Brier score of 0.0419 and ECE-15 of 0.0174. At an assumed 3% screening prevalence, however, estimated positive predictive value ranged from 0.213 to 0.420 across reported operating points. The findings support the internal technical value of texture-frequency spatial gating and compact local-global feature learning. They also show that distribution shift and prevalence exerted a larger practical effect than the margin separating OCAT-Net-S from its strongest internal comparator. Patient-linked multicenter evaluation, fixed external thresholds, prospective workflow testing, and quantitative lesion localization remain necessary before clinical use.

Article Details

How to Cite
Rafi, T. I., & Hemal, M. A. K. P. (2026). Reproducibility and Generalization Analysis of OCAT Net S for Oral Cancer Image Classification. The Eastasouth Journal of Information System and Computer Science, 4(01), 122–135. https://doi.org/10.58812/esiscs.v4i01.1235
Section
Articles

References

[1] M. H. Rijvi et al., “A compact texture frequency spatial gating network for oral cancer classification from clinical oral images,” Sci. Rep., 2026, doi: 10.1038/s41598-026-59480-0.

[2] A. L. D. Araújo, C. M. Pedroso, P. A. Vargas, M. A. Lopes, and A. R. Santos-Silva, “Advancing oral cancer diagnosis and risk assessment with artificial intelligence: a review,” Explor. Digit. Heal. Technol., vol. 3, p. 101147, 2025, doi: 10.37349/edht.2025.101147.

[3] M. F. Kabir, M. Y. Ahmad, R. Uddin, M. Cordero, and S. Kant, “Accurate and lightweight oral cancer detection using SE-MobileViT on clinically validated image dataset,” Discov. Artif. Intell., vol. 5, p. 173, 2025, doi: 10.1007/s44163-025-00442-2.

[4] D. P. Yadav, B. Sharma, A. Noonia, and A. Mehbodniya, “Explainable label guided lightweight network with axial transformer encoder for early detection of oral cancer,” Sci. Rep., vol. 15, no. 1, p. 6391, 2025, doi: 10.1038/s41598-025-87627-y.

[5] P. Liu and K. Bagi, “A tailored deep learning approach for early detection of oral cancer using a 19-layer CNN on clinical lip and tongue images,” Sci. Rep., vol. 15, no. 1, p. 23851, 2025, doi: 10.1038/s41598-025-07957-9.

[6] F. R. D. Cruze, J. Wasima, M. F. Hosen, M. B. A. Miah, Z. Muhammad, and M. F. Al Masud, “Oral Cancer Diagnosis Using Histopathology Images: An Explainable Hybrid Transformer Framework,” Technologies, vol. 14, no. 1, p. 39, 2026, doi: 10.3390/technologies14010039.

[7] A. Hossain et al., “Transformer-Based Ensemble Model for Binary and Multiclass Oral Cancer Segmentation,” 2025 International Conference on Electrical, Computer and Communication Engineering (ECCE). IEEE, 2025. doi: 10.1109/ECCE64574.2025.11012921.

[8] J. Štifanić, D. Štifanić, N. Anđelić, and Z. Car, “Explainable AI for Oral Cancer Diagnosis: Multiclass Classification of Histopathology Images and Grad-CAM Visualization,” Biology (Basel)., vol. 14, no. 8, p. 909, 2025, doi: 10.3390/biology14080909.

[9] P. Kaushik, V. Kukreja, E. Jain, M. Ati, and S. Hariharan, “Oral tumor detection and localization using detection transformer (DETR) with multi-scale feature extraction and self-supervised learning,” Array, vol. 30, p. 100825, 2026, doi: 10.1016/j.array.2026.100825.

[10] C. Karmakar et al., “Hierarchical lesion-aware transformer for oral cancer image classification,” Discov. Artif. Intell., 2026, doi: 10.1007/s44163-026-01776-1.

[11] T. R. Sikder et al., “XACT-TB an explainable hybrid CNN–swin transformer framework for tuberculosis screening from chest X-ray images,” Discov. Artif. Intell., vol. 6, p. 1030, 2026, doi: 10.1007/s44163-026-01509-4.

[12] M. I. Faruk et al., “Channel attention recalibrated fusion of CoAtNet and frozen DINOv2 features for binary chest X ray classification,” Discov. Comput., vol. 29, p. 551, 2026, doi: 10.1007/s10791-026-10446-w.

[13] P. D. Madan Kumar et al., “A Smartphone-based Comprehensive Dataset of Annotated Oral Cavity Images for Enhanced Oral Disease Diagnosis,” Sci. Data, vol. 13, p. 676, 2026, doi: 10.1038/s41597-026-06954-5.