Research Article | Open Access | Download PDF
Volume 74 | Issue 9 | Year 2026 | Article Id. IJETT-V74I9P139 | DOI : https://doi.org/10.14445/22315381/IJETT-V74I9P139Hybrid Contextual Embedding Framework for Low-Resource Gujarati Word Sense Disambiguation
Avani N Dave, Sanjay M Shah, Nakul R Dave
| Received | Revised | Accepted | Published |
|---|---|---|---|
| 19 May 2026 | 19 Aug 2026 | 26 Aug 2026 | 30 Sep 2026 |
Citation :
Avani N Dave, Sanjay M Shah, Nakul R Dave, "Hybrid Contextual Embedding Framework for Low-Resource Gujarati Word Sense Disambiguation," International Journal of Engineering Trends and Technology (IJETT), vol. 74, no. 9, pp. 562-579, 2026. Crossref, https://doi.org/10.14445/22315381/IJETT-V74I9P139
Abstract
Word sense disambiguation continues to constitute significant challenges in natural language processing, particularly for languages such as Gujarati that exhibit considerable morphological variation. Thus, natural language processing has not progressed well for word-sense disambiguation in Gujarati. After all, there are essentially two issues here: insufficient data and the languages' intrinsic complexity. Given the various forms of Gujarati and the widespread use of idioms, models do not perform well and yield inaccurate results. In order to address a gap in the existing literature, this study presents a new Gujarati word sense disambiguation dataset, which has been manually word-sense-tagged and is specifically tailored for ambiguous Gujarati lexemes. The corpus includes 150 ambiguous words, each of which has been carefully annotated to show its exact meaning in context. This dataset contains 426 distinct meanings across 4,400 sentences, making it a complete test set for sense tagging. Still, natural language processing has a long way to go before it can correctly interpret a written word and universally understand its meaning in highly morphological languages like Gujarati. This paper introduces a context-aware hybrid architecture designed to improve the identification of polysemous words in Gujarat texts. The proposed model utilizes various contextual characteristics and combines lexical, statistical, and contextual components to address the problems associated with the use of single-component models. The testing process involved the use of five-fold cross-validation considering Random Forest and BERT algorithms. The resulting accuracy was remarkably high at 97.54%, with the corresponding Macro-F1 value being 97.40%. A suite of models that exhibits tolerance for the correct sense of the word without overfitting can be found in the well-balanced macro-precision value of 98.14%. The system's stability and consistency have been confirmed by the consistently low standard deviation, which is ±0.53 percentage points. Together, these results show that the method is viable when applied to word sense disambiguation in the Gujarati language.
Keywords
Machine learning, Natural Language Processing, Word Sense Disambiguation, BERT, Random Forest, Sense Annotated Corpus, Gujarati Language, Polysemy Words.
References
[1] Eluri Suneetha,
Miriyala Kanthirekha, and Dhana Lakshmi Gorle, “Word Sense Disambiguation of
Regional Language using Deep Learning,” I-Manager’s Journal on Computer
Science, vol. 13, no. 1, 2025.
[CrossRef]
[Google Scholar] [Publisher Link]
[2] Gerard Escudero Bakx, “Machine
Learning Techniques for Word Sense Disambiguation,” Polytechnic University
of Catalonia, 2006.
[Google Scholar] [Publisher Link]
[3] Ajith Abraham et al.,
“Improvement of Translation Accuracy for the Word Sense Disambiguation System
using Novel Classifier Approach,” International Arab Journal of Information
Technology, vol. 21, no. 6, pp. 1128-1142, 2024.
[CrossRef] [Google Scholar]
[4] C.P. Chandrika, and
Jagadish S. Kallimani, “Word Sense Disambiguation for Indian Regional Language
using BERT Model,” Proceedings of Fifth International Conference on Smart
Computing and Informatics, Singapore: Springer Nature Singapore, vol. 2,
pp. 127-137, 2022.
[CrossRef]
[Google Scholar] [Publisher Link]
[5] Sandip S. Patil et al.,
“Bert and Indowordnet Collaborative Embedding for Enhanced Marathi Word Sense
Disambiguation,” ICTACT Journal on Soft Computing, vol. 13, no. 2, pp.
2842-2849, 2023.
[CrossRef]
[Google Scholar] [Publisher Link]
[6] Anusha, K. Manjula
Shenoy, and Smitha N. Pai, “Embedding Digital News Titles using Wordnet
Knowledge and WSD Over Bert Model,” 2025 IEEE International Conference on Distributed
Computing, Mangalore, India, pp. 1-6, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[7] Roberto Navigli, “Word
Sense Disambiguation: A Survey,” ACM Computing Surveys, vol. 41, no. 2,
pp. 1-69, 2009.
[CrossRef]
[Google Scholar] [Publisher Link]
[8] Tinghua Wang, Junyang
Rao, and Qi Hu, “Supervised Word Sense Disambiguation using Semantic Diffusion
Kernel,” Engineering Applications of Artificial Intelligence, vol. 27,
pp. 167-174, 2014.
[CrossRef]
[Google Scholar] [Publisher Link]
[9] Abdulgabbar Saif et
al., “Building Sense Tagged Corpus using Wikipedia for Supervised Word Sense
Disambiguation,” Procedia Computer Science, vol. 123, pp. 403-412, 2018.
[CrossRef]
[Google Scholar] [Publisher Link]
[10] Sunita Rawat,
“Supervised Word Sense Disambiguation using Decision Tree,” International
Journal of Recent Technology and Engineering, vol. 8, no. 2, pp. 4043-4047,
2019.
[CrossRef]
[Google Scholar] [Publisher Link]
[11] Kaveh Taghipour, and
Hwee Tou Ng, “Semi-Supervised Word Sense Disambiguation using Word Embeddings
in General and Specific Domains,” Proceedings of the 2015 Conference of the
North American Chapter of the Association for Computational Linguistics: Human
Language Technologies, Denver, Colorado, pp. 314-323, 2015.
[CrossRef] [Google Scholar] [Publisher Link]
[12] Vinto Gujjar et al., “A
Literature Survey on Word Sense Disambiguation for the Hindi Language,” Information,
vol. 14, no. 9, pp. 1-25, 2023.
[CrossRef]
[Google Scholar] [Publisher Link]
[13] Alok Ranjan Pal et al.,
“Hybrid Approach to Word Sense Disambiguation Combining Supervised and
Unsupervised Learning,” International Journal of Artificial Intelligence and
Applications, vol. 4, no. 4, pp. 89-101, 2013.
[CrossRef] [Google Scholar] [Publisher Link]
[14] Tabor Wegi Geleta, and
Jara Muda Haro, “Semisupervised Learning-based Word-Sense Disambiguation using
Word Embedding for Afaan Oromoo Language,” Applied Computational
Intelligence and Soft Computing, vol. 2024, no. 1, pp. 1-11, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[15] Tarjni Vyas, and Amit
Ganatra, “Gujarati Language Model: Word Sense Disambiguation using Supervised
Technique,” International Journal of Recent Technology and Engineering,
vol. 8, no. 2S11, pp. 3740-3744, 2019.
[CrossRef] [Google Scholar]
[16] Bhagyada Desai, Krinal
Naik, and Preeti Bhatt, “Word Sense Disambiguation in Gujarati Language,” International
Journal of Innovative Research in Computer Science and Technology, vol. 3,
no. 1, pp. 44-47, 2015.
[Google Scholar] [Publisher Link]
[17] Parth J. Vasoya, and
Tarjni Vyas, “A Survey on Word Sense Disambiguation Approaches,” International
journal of Trend in Research and Development, vol. 1, no. 1, pp. 1-3, 2014.
[Publisher Link]
[18] Zankhana B. Vaishnav,
“Gujarati Word Sense Disambiguation using Genetic Algorithm,” International
Journal on Recent and Innovation Trends in Computing and Communication,
vol. 5, no. 6, pp. 635-639, 2017.
[Google Scholar]
[19] Tarjni Vyas, and Amit
Ganatra, “Gujarati Language: Research Issues, Resources and Proposed Method on
Word Sense Disambiguation,” International Journal of Recent Technology and
Engineering, vol. 8, no. 2S11, pp. 3745-3749, 2019.
[CrossRef] [Google Scholar]
[20] Manish Sinha et al.,
“Hindi Word Sense Disambiguation,” Machine Translation, pp. 1-7, 2000.
[Google Scholar]
[21] Pranjal Protim Borah,
Gitimoni Talukdar, and Arup Baruah, “Assamese Word Sense Disambiguation using
Supervised Learning,” 2014 International Conference on Contemporary
Computing and Informatics, Mysore, India, pp. 946-950, 2014.
[CrossRef] [Google Scholar] [Publisher Link]
[22] Jumi Sarmah, and
Shikhar Kr. Sarma, “Decision Tree based Supervised Word Sense Disambiguation
for Assamese,” International Journal of Computer Applications, vol. 141,
no. 1, pp. 42-48, 2016.
[Google Scholar] [Publisher Link]
[23] Sailendra Kumar, and
Rakesh Kumar, “Word Sense Disambiguation in the Hindi Language: Neural Network
Approach,” International Journal of Technical Research and Science, pp.
72-76, 2021.
[CrossRef] [Google Scholar] [Publisher Link]
[24] Nazreena Rahman, and
Bhogeswar Borah, “An Unsupervised Method for Word Sense Disambiguation,” Journal
of King Saud University-Computer and Information Sciences, vol. 34, no. 9,
pp. 6643-6651, 2022.
[CrossRef]
[Google Scholar] [Publisher Link]
[25] Anh-Cuong Le et al.,
“Semi-Supervised Learning Integrated with Classifier Combination for Word Sense
Disambiguation,” Computer Speech and Language, vol. 22, no. 4, pp.
330-345, 2008.
[CrossRef]
[Google Scholar] [Publisher Link]
[26] Chandrakant D. Kokane,
and Sachin D. Babar, “Supervised Word Sense Disambiguation with Recurrent
Neural Network Model,” International Journal of Engineering and Advanced
Technology, vol. 9, no. 2, pp. 1447-1453, 2019.
[CrossRef] [Google Scholar] [Publisher Link]
[27] Huei-Ling Lai et al.,
“Supervised Word Sense Disambiguation on Polysemy with Neural Network Models: A
Case Study of BUN in Taiwan Hakka,” International Journal of Asian Language
Processing, vol. 30, no. 3, 2020.
[CrossRef]
[Google Scholar] [Publisher Link]
[28] Anup Kumar Barman et
al., “Word Sense Disambiguation Applied to Assamese-Hindi Bilingual Statistical
Machine Translation,” Engineering, Technology and Applied Science Research,
vol. 14, no. 1, pp. 12581-12586, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[29] Debapratim Das Dawn et
al., “Lexeme Connexion Measure of Cohesive Lexical Ambiguity Revealing Factor:
A Robust Approach for Word Sense Disambiguation of Bengali Text,” Multimedia
Tools and Applications, vol. 83, no. 5, pp. 12939-12983, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[30] Binod Kumar Mishra, and
Suresh Jain, “An Innovative Method for Hindi Word Sense Disambiguation,” SN
Computer Science, vol. 4, no. 6, pp. 1-17, 2023.
[CrossRef] [Google Scholar] [Publisher Link]
[31] Debapratim Das Dawn et
al., “A Dataset for Evaluating Bengali Word Sense Disambiguation Techniques,” Journal
of Ambient Intelligence and Humanized Computing, vol. 14, no. 4, pp.
4057-4086, 2023.
[CrossRef]
[Google Scholar] [Publisher Link]
[32] Binod Kumar Mishra, and
Suresh Jain, “Word Sense Disambiguation for Indic Language using Bi-LSTM,” Multimedia
Tools and Applications, vol. 84, no. 16, pp. 16631-16656, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[33] Chhaya S. Patil, and
Vaishali B. Patil, “A Multilingual Exploration of Word Sense Disambiguation
using Transformer Models: Dravidian and Devanagari Languages, 1st
ed., Recent Advances in Computing Sciences, CRC Press, pp. 158-160, 2025.
[Google Scholar] [Publisher Link]
[34] Sunjae Kwon et al.,
“Vision Meets Definitions: Unsupervised Visual Word Sense Disambiguation
Incorporating Gloss Information,” Proceedings of the 61st Annual
Meeting of the Association for Computational Linguistics, pp. 1583-1598,
2023.
[CrossRef]
[Google Scholar] [Publisher Link]
[35] Robbel Habtamu Yigzaw,
Beakal Gizachew Assefa, and Elefelious Getachew Belay, “A Hybrid Contextual
Embedding and Hierarchical Attention for Improving the Performance of Word
Sense Disambiguation,” IEEE Access, vol. 13, pp. 21744-21758, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[36] Hlaudi Daniel Masethe
et al., “Hybrid Transformer-based Large Language Models for Word Sense
Disambiguation in the Low-Resource Sesotho sa Leboa Language,” Applied
Sciences, vol. 15, no. 7, pp. 1-33, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[37] Tawseef Ahmad Mir, and
Aadil Ahmad Lawaye, “Word Sense Disambiguation Corpus for Kashmiri,” Natural
Language Processing, vol. 31, no. 2, pp. 631-654, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[38] Ali Osman Mohammed
Salih et al., “Enhancing a Context-based Supervised Machine Learning Model for
Accurate Disambiguation of English Polysemous Word Senses,” Scientific
Culture, vol. 11, no. 4, 1945.
[Google Scholar]
[39] Shailendra Kumar Patel,
Rakesh Kumar, and Anuj Kumar Sirohi, “BERT for Hindi Word Sense
Disambiguation, 1st ed., Intelligent Computing and Communication
Techniques, CRC Press, pp. 747-753, 2025.
[Google Scholar] [Publisher Link]
[40] Madhuri Karnik et al.,
“State of the Art Analysis of Word Sense Disambiguation,” Intelligent Computing
for Sustainable Development, Springer, Cham, pp. 55-70, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[41] Aditi Barai et al., “A
Comprehensive Analysis of Word Sense Disambiguation in a Regional Language,” 2025
International Conference on Inventive Computation Technologies,
Kirtipur, Nepal, pp. 1260-1265, 2025.
[CrossRef]
[Google Scholar] [Publisher Link]
[42] Ratul Das, Alok Ranjan
Pal, and Diganta Saha, “Unsupervised Approach for Word Sense Disambiguation in Bengali,”
Computational Technologies and Electronics, Springer, Cham, pp. 196-206,
2025.
[CrossRef]
[Google Scholar] [Publisher Link]
[43] Youddha Beer Singh et
al., A Handbook of Computational Linguistics: Artificial Intelligence in
Natural Language Processing, Bentham Science Publishers, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[44] Lianwang Hao, Tao
Zhang, and Huaixin Liang, “A Novel Rotational Causal Random Forest Approach for
Word Sense Disambiguation of English Modal Verbs,” SSRN, pp. 1-9, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[45] Aarti Purohit, and
Kuldeep Kumar Yogi, “Enhancing Word Sense Disambiguation for Hindi Agriculture
Domain: Feature Engineering and Machine Learning Approaches,” Results in
Engineering, vol. 28, pp. 1-12, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[46] Anirudh S Nair et al.,
“Improving Word Sense Disambiguation by Adopting Refined Algorithms,” 2024
Third International Conference on Trends in Electrical, Electronics, and
Computer Engineering, Bangalore, India, pp. 135-140, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[47] Vivek A. Manwar, and
A.B. Manwar, “mBERT: A Query Refinement Model for Marathi Word Sense
Disambiguation,” 2025 IEEE International Students' Conference on
Electrical, Electronics and Computer Science, Bhopal, India, pp. 1-6, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[48] Pawan Makhija, and
Sanjay Tanwani, “An Empirical Study on Various Word Sense Disambiguation
Techniques in the Biomedical Domain,” International Conference on Recent
Advancements and Modernisations in Sustainable Intelligent Technologies and
Applications, Indore, India, 2025.
[CrossRef] [Google Scholar] [Publisher Link]