TongueViT: A Vision Transformer-Based Framework for Automated Tongue Image Classification

dc.contributor.authorPandya Darshanaben D.
dc.contributor.authorTamhankar Ishaan
dc.contributor.authorKulkarni Jaimini
dc.contributor.authorTrivedi Vishal
dc.contributor.authorDegadwala Sheshang
dc.contributor.authorVyas Dhairya
dc.date.accessioned2026-06-27T04:20:04Z
dc.date.issued2026
dc.description.abstractTongue diagnosis is a serious non-invasive technique used in both traditional and modern medical diagnosis to examine the internal health status. Manual interpretation however is subjective and suffers inconsistencies. In this regard, paper introduce TongueViT, a new Vision Transformer (ViT)-based sys-tem of automated classification of tongue images. Compared to traditional CNN-based methods, TongueViT with the self-attention mechanism of transformers can exploit global contextual features to discriminate accurately between affected and normal tongue images. This model was trained and tested on a preprocessed dataset containing two categories of affected and healthy with stupendous classifi-cation accuracy of 99%. The results shown that TongueViT massively outcompetes the traditional models both in terms of performance and interpretability. Atten-tion maps also demonstrate that the model looks at clinically interesting areas of the tongue, making it more probable to be translated into real-world diagnostic assistance. These results confirm the usefulness of transformer-based models in biomedical image analysis and open the prospect of AI-aided tongue diagnostics in telemedicine and clinical screening.
dc.identifier.citationPandya, D.D., Tamhankar, I., Kulkarni, J.S., Trivedi, V.R., Degadwala, S., Vyas, D. (2026). TongueViT: A Vision Transformer-Based Framework for Automated Tongue Image Classification. In: Lanka, S., Cabezuelo, A.S., Tugui, A. (eds) Trends in Sustainable Computing and Machine Intelligence. ICTSM 2025. Lecture Notes in Networks and Systems, vol 1755. Springer, Cham.
dc.identifier.urihttps://doi.org/10.1007/978-3-032-13177-5_4
dc.identifier.urihttp://160.160.1.15:4000/handle/123456789/548
dc.language.isoen
dc.publisherTongueViT: A Vision Transformer-Based Framework for Automated Tongue Image Classification. In: Lanka, S., Cabezuelo, A.S., Tugui, A. (eds) Trends in Sustainable Computing and Machine Intelligence. ICTSM 2025. Lecture Notes in Networks and Systems, vol 1755. Springer, Cham.
dc.subjectTongue diagnosis cdot Vision Transformer cdot image classification cdot medical imaging cdot deep learning
dc.titleTongueViT: A Vision Transformer-Based Framework for Automated Tongue Image Classification
dc.typeBook chapter

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Jaimini mam-6.pdf
Size:
78.02 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.45 KB
Format:
Item-specific license agreed to upon submission
Description: