A Vision Transformer-Based Approach for Human Organ Identification across Multiple Medical Imaging Modalities
DOI:
https://doi.org/10.47839/ijc.25.2.4658Keywords:
Vision Transformer, Medical Image Classification, Human Organ Identification, Brain Tumor MRI Dataset (BR35H), Mammographic Image Analysis Society Dataset (MIAS)Abstract
Modern healthcare relies extensively on medical imaging, which helps with disease monitoring, diagnosis, and therapy planning. In many medical conditions, diseases start in one organ and spread to multiple organs (e.g., metastasis). Organ identification ensures that each organ is independently analyzed for associated abnormalities, offering a comprehensive diagnostic solution. In medical research and diagnostic procedures, the categorization of medical images is of utmost importance, especially for organs like the brain, breast, lungs, and liver are frequently mentioned in cancer reports. Computed Tomography (CT), Magnetic Resonance Imaging (MRI), Breast Mammograms (MIAS), Brain (Br35H), Lung (IQ-OTHNCCD), and Liver (LiTs) are among the datasets used. ViT demonstrates self-attention, allowing each part of the image to weigh its importance relative to other parts and capture long-range dependency between input sequences. The model utilizes the Adam optimizer with a learning rate of 0.0001, achieving strong performance across multiple imaging types. In validation of the aforementioned datasets, the model’s accuracy for the brain, breast, lung, and liver is 96%, 94%, 93%, and 92%, respectively. The results indicate that ViT can minimize handcrafted feature extraction and achieve state-of-the-art performance by applying its attention mechanism to medical imaging applications.
References
F. Bray, M. Laversanne, H. Sung, et al., “Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries,” CA: A Cancer Journal for Clinicians, vol. 74, no. 3, pp. 229–263, 2024, https://doi.org/10.3322/caac.21834.
R. L. Siegel, A. N. Giaquinto, and A. Jemal, “Cancer statistics, 2024,” CA: A Cancer Journal for Clinicians, vol. 74, no. 1, pp. 12–49, 2024, https://doi.org/10.3322/caac.21820.
J. Rong and Y. Liu, “Advances in medical imaging techniques,” BMC Methods, vol. 1, p. 10, 2024, https://doi.org/10.1186/s44330-024-00010-7.
S. Aljahdali, G. Azim, W. Zabani, S. Bafaraj, J. Alyami, and A. Abduljabbar, “Effectiveness of radiology modalities in diagnosing and characterizing brain disorders,” Neurosciences (Riyadh), vol. 29, no. 1, pp. 37–43, 2024, https://doi.org/10.17712/nsj.2024.1.20230048.
R. Archana and P. S. E. Jeevaraj, “Deep learning models for digital image processing: A review,” Artificial Intelligence Review, vol. 57, p. 11, 2024, https://doi.org/10.1007/s10462-023-10631-z.
S. Takahashi, Y. Sakaguchi, N. Kouno, et al., “Comparison of vision transformers and convolutional neural networks in medical image analysis: A systematic review,” Journal of Medical Systems, vol. 48, p. 84, 2024, https://doi.org/10.1007/s10916-024-02105-8.
O. Elharrouss, R. Damseh, A. N. Belkacem, et al., “Transformer-based image and video inpainting: Current challenges and future directions,” Artificial Intelligence Review, vol. 58, p. 124, 2025, https://doi.org/10.1007/s10462-024-11075-9.
M. Khalil, A. Khalil, and A. Ngom, “A comprehensive study of vision transformers in image classification tasks,” arXiv preprint arXiv:2312.01232v1, 2023.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, et al., “An image is worth 16×16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020.
I. Aboussaleh, J. Riffi, K. El Fazazy, A. M. Mahraz, and H. Tairi, “3DUV-NetR+: A 3D hybrid semantic architecture using transformers for brain tumor segmentation with multimodal MR images,” Results in Engineering, p. 101892, 2024. https://doi.org/10.1016/j.rineng.2024.101892.
G. Ayana, E. Lee, and S.-W. Choe, “Vision transformers for breast cancer human epidermal growth factor receptor 2 expression staging without immunohistochemical staining,” The American Journal of Pathology, vol. 194, no. 3, pp. 402–414, 2024, https://doi.org/10.1016/j.ajpath.2023.11.015.
J. Ko, S. Park, and H.-G. Woo, “Optimization of vision transformer-based detection of lung diseases from chest X-ray images,” BMC Medical Informatics and Decision Making, vol. 24, p. 191, 2024, https://doi.org/10.1186/s12911-024-02591-3.
T. C. Phan, C. H. Ho, and A. C. Phan, “Detection and classification of liver lesions using vision transformer and active learning,” Future Data and Security Engineering (FDSE 2024), ser. Communications in Computer and Information Science, vol. 2309. Singapore: Springer, 2024, https://doi.org/10.1007/978-981-96-0434-0_15.
J. Martín and J. Sánchez, “Evaluation of vision transformers for multi-organ tumor classification using MRI and CT imaging,” Electronics, vol. 14, no. 15, p. 2976, 2025, https://doi.org/10.3390/electronics14152976.
X. Liu, L. Qu, Z. Xie, Y. Shi, and Z. Song, “Deep mutual learning among partially labeled datasets for multi-organ segmentation,” arXiv preprint arXiv:2407.12611, 2024.
X. Liu, L. Qu, Z. Xie, et al., “Towards more precise automatic analysis: A systematic review of deep learning-based multi-organ segmentation,” BioMedical Engineering OnLine, vol. 23, p. 52, 2024, https://doi.org/10.1186/s12938-024-01238-8.
A. Agarwal, A. Chauhan, I. Al-Rwae, P. Tamizharasan, Y. Seo, and D. Mitra, “Constraint-based multi-organ identification in CT images using unsupervised learning,” Proc. IEEE Nuclear Science Symp. and Medical Imaging Conf. (NSS/MIC), Italy, 2022, pp. 1–4, https://doi.org/10.1109/NSS/MIC44845.2022.10398918.
D. Santhosh Reddy, P. Rajalakshmi, and M. A. Mateen, “A deep learning-based approach for classification of abdominal organs using ultrasound images,” Biocybernetics and Biomedical Engineering, vol. 41, no. 2, pp. 779–791, 2021, https://doi.org/10.1016/j.bbe.2021.05.004.
H. Xu, Q. Xu, F. Cong, J. Kang, C. Han, Z. Liu, A. Madabhushi, and C. Lu, “Vision transformers for computational histopathology,” IEEE Reviews in Biomedical Engineering, 2023. https://doi.org/10.1109/RBME.2023.3297604.
A. Hatamizadeh, Y. Tang, V. Nath, D. Yang, H. Roth, and D. Xu, “UNETR: Transformers for 3D medical image segmentation,” Proc. IEEE/CVF Winter Conf. on Applications of Computer Vision (WACV), 2022. https://doi.org/10.1109/WACV51458.2022.00181.
Kaggle, “Brain tumor detection dataset,” [Online]. Available at: https://www.kaggle.com/datasets/ahmedhamada0/brain-tumor-detection.
University of Cambridge Repository, “Dataset resource,” [Online]. Available at: https://www.repository.cam.ac.uk/handle/1810/250394.
Kaggle, “IQ-OTH/NCCD lung cancer dataset,” [Online]. Available at: https://www.kaggle.com/datasets/hamdallak/the-iqothnccd-lung-cancer-dataset.
Kaggle, “Liver tumor segmentation dataset,” [Online]. Available at: https://www.kaggle.com/datasets/andrewmvd/liver-tumor-segmentation.
S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,” arXiv preprint arXiv:2101.01169, 2021.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers and distillation through attention,” Proc. Int. Conf. on Machine Learning (ICML), 2021.
G. Litjens, T. Kooi, B. E. Bejnordi, et al., “A survey on deep learning in medical image analysis,” Medical Image Analysis, vol. 42, pp. 60–88, 2017. https://doi.org/10.1016/j.media.2017.07.005.
M. P. McBee, O. A. Awan, A. T. Colucci, et al., “Deep learning in radiology,” Academic Radiology, vol. 25, no. 11, pp. 1472–1480, 2018. https://doi.org/10.1016/j.acra.2018.02.018.
D. Hendrycks and K. Gimpel, “A baseline for detecting misclassified and out-of-distribution examples in neural networks,” arXiv preprint arXiv:1610.02136, 2017.
A. Paul, A. Ganesh, and S. Maheshwari, “Fine-tuning vision transformers for small datasets,” Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2022.
Viso.ai, “Vision transformer (ViT): Explanation and guide,” [Online]. Available at: https://viso.ai/deep-learning/vision-transformer-vit/.
H. Ghabri, M. S. Alqahtani, S. Ben Othman, et al., “Transfer learning for accurate fetal organ classification from ultrasound images: A potential tool for maternal healthcare providers,” Scientific Reports, vol. 13, p. 17904, 2023, https://doi.org/10.1038/s41598-023-44689-0.
E. Somasundaram, Z. Taylor, V. V. Alves, et al., “Deep learning models for abdominal CT organ segmentation in children: Development and validation in internal and heterogeneous public datasets,” American Journal of Roentgenology, vol. 223, no. 1, p. e2430931, 2024, https://doi.org/10.2214/AJR.24.30931.
Downloads
Published
How to Cite
Issue
Section
License
International Journal of Computing is an open access journal. Authors who publish with this journal agree to the following terms:• Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
• Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
• Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.