TongueViT: A Vision Transformer-Based Framework for Automated Tongue Image Classification
| dc.contributor.author | Pandya Darshanaben D. | |
| dc.contributor.author | Tamhankar Ishaan | |
| dc.contributor.author | Kulkarni Jaimini | |
| dc.contributor.author | Trivedi Vishal | |
| dc.contributor.author | Degadwala Sheshang | |
| dc.contributor.author | Vyas Dhairya | |
| dc.date.accessioned | 2026-06-27T04:20:04Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | Tongue diagnosis is a serious non-invasive technique used in both traditional and modern medical diagnosis to examine the internal health status. Manual interpretation however is subjective and suffers inconsistencies. In this regard, paper introduce TongueViT, a new Vision Transformer (ViT)-based sys-tem of automated classification of tongue images. Compared to traditional CNN-based methods, TongueViT with the self-attention mechanism of transformers can exploit global contextual features to discriminate accurately between affected and normal tongue images. This model was trained and tested on a preprocessed dataset containing two categories of affected and healthy with stupendous classifi-cation accuracy of 99%. The results shown that TongueViT massively outcompetes the traditional models both in terms of performance and interpretability. Atten-tion maps also demonstrate that the model looks at clinically interesting areas of the tongue, making it more probable to be translated into real-world diagnostic assistance. These results confirm the usefulness of transformer-based models in biomedical image analysis and open the prospect of AI-aided tongue diagnostics in telemedicine and clinical screening. | |
| dc.identifier.citation | Pandya, D.D., Tamhankar, I., Kulkarni, J.S., Trivedi, V.R., Degadwala, S., Vyas, D. (2026). TongueViT: A Vision Transformer-Based Framework for Automated Tongue Image Classification. In: Lanka, S., Cabezuelo, A.S., Tugui, A. (eds) Trends in Sustainable Computing and Machine Intelligence. ICTSM 2025. Lecture Notes in Networks and Systems, vol 1755. Springer, Cham. | |
| dc.identifier.uri | https://doi.org/10.1007/978-3-032-13177-5_4 | |
| dc.identifier.uri | http://160.160.1.15:4000/handle/123456789/548 | |
| dc.language.iso | en | |
| dc.publisher | TongueViT: A Vision Transformer-Based Framework for Automated Tongue Image Classification. In: Lanka, S., Cabezuelo, A.S., Tugui, A. (eds) Trends in Sustainable Computing and Machine Intelligence. ICTSM 2025. Lecture Notes in Networks and Systems, vol 1755. Springer, Cham. | |
| dc.subject | Tongue diagnosis cdot Vision Transformer cdot image classification cdot medical imaging cdot deep learning | |
| dc.title | TongueViT: A Vision Transformer-Based Framework for Automated Tongue Image Classification | |
| dc.type | Book chapter |
