International Journal of Engineering
Trends and Technology

Research Article | Open Access | Download PDF
Volume 74 | Issue 9 | Year 2026 | Article Id. IJETT-V74I9P139 | DOI : https://doi.org/10.14445/22315381/IJETT-V74I9P139

Hybrid Contextual Embedding Framework for Low-Resource Gujarati Word Sense Disambiguation


Avani N Dave, Sanjay M Shah, Nakul R Dave

Received Revised Accepted Published
19 May 2026 19 Aug 2026 26 Aug 2026 30 Sep 2026

Citation :

Avani N Dave, Sanjay M Shah, Nakul R Dave, "Hybrid Contextual Embedding Framework for Low-Resource Gujarati Word Sense Disambiguation," International Journal of Engineering Trends and Technology (IJETT), vol. 74, no. 9, pp. 562-579, 2026. Crossref, https://doi.org/10.14445/22315381/IJETT-V74I9P139

Abstract

Word sense disambiguation continues to constitute significant challenges in natural language processing, particularly for languages such as Gujarati that exhibit considerable morphological variation. Thus, natural language processing has not progressed well for word-sense disambiguation in Gujarati. After all, there are essentially two issues here: insufficient data and the languages' intrinsic complexity. Given the various forms of Gujarati and the widespread use of idioms, models do not perform well and yield inaccurate results. In order to address a gap in the existing literature, this study presents a new Gujarati word sense disambiguation dataset, which has been manually word-sense-tagged and is specifically tailored for ambiguous Gujarati lexemes. The corpus includes 150 ambiguous words, each of which has been carefully annotated to show its exact meaning in context. This dataset contains 426 distinct meanings across 4,400 sentences, making it a complete test set for sense tagging. Still, natural language processing has a long way to go before it can correctly interpret a written word and universally understand its meaning in highly morphological languages like Gujarati. This paper introduces a context-aware hybrid architecture designed to improve the identification of polysemous words in Gujarat texts. The proposed model utilizes various contextual characteristics and combines lexical, statistical, and contextual components to address the problems associated with the use of single-component models. The testing process involved the use of five-fold cross-validation considering Random Forest and BERT algorithms. The resulting accuracy was remarkably high at 97.54%, with the corresponding Macro-F1 value being 97.40%. A suite of models that exhibits tolerance for the correct sense of the word without overfitting can be found in the well-balanced macro-precision value of 98.14%. The system's stability and consistency have been confirmed by the consistently low standard deviation, which is ±0.53 percentage points. Together, these results show that the method is viable when applied to word sense disambiguation in the Gujarati language.

Keywords

Machine learning, Natural Language Processing, Word Sense Disambiguation, BERT, Random Forest, Sense Annotated Corpus, Gujarati Language, Polysemy Words.

References

[1] Eluri Suneetha, Miriyala Kanthirekha, and Dhana Lakshmi Gorle, “Word Sense Disambiguation of Regional Language using Deep Learning,” I-Manager’s Journal on Computer Science, vol. 13, no. 1, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[2] Gerard Escudero Bakx, “Machine Learning Techniques for Word Sense Disambiguation,” Polytechnic University of Catalonia, 2006.
[Google Scholar] [Publisher Link]

[3] Ajith Abraham et al., “Improvement of Translation Accuracy for the Word Sense Disambiguation System using Novel Classifier Approach,” International Arab Journal of Information Technology, vol. 21, no. 6, pp. 1128-1142, 2024.
[CrossRef] [Google Scholar]      

[4] C.P. Chandrika, and Jagadish S. Kallimani, “Word Sense Disambiguation for Indian Regional Language using BERT Model,” Proceedings of Fifth International Conference on Smart Computing and Informatics, Singapore: Springer Nature Singapore, vol. 2, pp. 127-137, 2022.
[CrossRef] [Google Scholar] [Publisher Link]

[5] Sandip S. Patil et al., “Bert and Indowordnet Collaborative Embedding for Enhanced Marathi Word Sense Disambiguation,” ICTACT Journal on Soft Computing, vol. 13, no. 2, pp. 2842-2849, 2023.
[CrossRef] [Google Scholar] [Publisher Link]

[6] Anusha, K. Manjula Shenoy, and Smitha N. Pai, “Embedding Digital News Titles using Wordnet Knowledge and WSD Over Bert Model,” 2025 IEEE International Conference on Distributed Computing, Mangalore, India, pp. 1-6, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[7] Roberto Navigli, “Word Sense Disambiguation: A Survey,” ACM Computing Surveys, vol. 41, no. 2, pp. 1-69, 2009.
[CrossRef] [Google Scholar] [Publisher Link]

[8] Tinghua Wang, Junyang Rao, and Qi Hu, “Supervised Word Sense Disambiguation using Semantic Diffusion Kernel,” Engineering Applications of Artificial Intelligence, vol. 27, pp. 167-174, 2014.
[CrossRef] [Google Scholar] [Publisher Link]

[9] Abdulgabbar Saif et al., “Building Sense Tagged Corpus using Wikipedia for Supervised Word Sense Disambiguation,” Procedia Computer Science, vol. 123, pp. 403-412, 2018.
[CrossRef] [Google Scholar] [Publisher Link]

[10] Sunita Rawat, “Supervised Word Sense Disambiguation using Decision Tree,” International Journal of Recent Technology and Engineering, vol. 8, no. 2, pp. 4043-4047, 2019.
[CrossRef] [Google Scholar] [Publisher Link]

[11] Kaveh Taghipour, and Hwee Tou Ng, “Semi-Supervised Word Sense Disambiguation using Word Embeddings in General and Specific Domains,” Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Denver, Colorado, pp. 314-323, 2015.
[CrossRef] [Google Scholar] [Publisher Link]

[12] Vinto Gujjar et al., “A Literature Survey on Word Sense Disambiguation for the Hindi Language,” Information, vol. 14, no. 9, pp. 1-25, 2023.
[CrossRef] [Google Scholar] [Publisher Link]

[13] Alok Ranjan Pal et al., “Hybrid Approach to Word Sense Disambiguation Combining Supervised and Unsupervised Learning,” International Journal of Artificial Intelligence and Applications, vol. 4, no. 4, pp. 89-101, 2013.
[CrossRef] [Google Scholar] [Publisher Link]

[14] Tabor Wegi Geleta, and Jara Muda Haro, “Semisupervised Learning-based Word-Sense Disambiguation using Word Embedding for Afaan Oromoo Language,” Applied Computational Intelligence and Soft Computing, vol. 2024, no. 1, pp. 1-11, 2024.
[CrossRef] [Google Scholar] [Publisher Link]

[15] Tarjni Vyas, and Amit Ganatra, “Gujarati Language Model: Word Sense Disambiguation using Supervised Technique,” International Journal of Recent Technology and Engineering, vol. 8, no. 2S11, pp. 3740-3744, 2019.
[CrossRef] [Google Scholar]

[16] Bhagyada Desai, Krinal Naik, and Preeti Bhatt, “Word Sense Disambiguation in Gujarati Language,” International Journal of Innovative Research in Computer Science and Technology, vol. 3, no. 1, pp. 44-47, 2015.
[
Google Scholar] [Publisher Link]

[17] Parth J. Vasoya, and Tarjni Vyas, “A Survey on Word Sense Disambiguation Approaches,” International journal of Trend in Research and Development, vol. 1, no. 1, pp. 1-3, 2014.
[Publisher Link]

[18] Zankhana B. Vaishnav, “Gujarati Word Sense Disambiguation using Genetic Algorithm,” International Journal on Recent and Innovation Trends in Computing and Communication, vol. 5, no. 6, pp. 635-639, 2017.
[Google Scholar]

[19] Tarjni Vyas, and Amit Ganatra, “Gujarati Language: Research Issues, Resources and Proposed Method on Word Sense Disambiguation,” International Journal of Recent Technology and Engineering, vol. 8, no. 2S11, pp. 3745-3749, 2019.
[CrossRef] [Google Scholar]

[20] Manish Sinha et al., “Hindi Word Sense Disambiguation,” Machine Translation, pp. 1-7, 2000.
[Google Scholar]

[21] Pranjal Protim Borah, Gitimoni Talukdar, and Arup Baruah, “Assamese Word Sense Disambiguation using Supervised Learning,” 2014 International Conference on Contemporary Computing and Informatics, Mysore, India, pp. 946-950, 2014.
[CrossRef] [Google Scholar] [Publisher Link]

[22] Jumi Sarmah, and Shikhar Kr. Sarma, “Decision Tree based Supervised Word Sense Disambiguation for Assamese,” International Journal of Computer Applications, vol. 141, no. 1, pp. 42-48, 2016.
[Google Scholar] [Publisher Link]

[23] Sailendra Kumar, and Rakesh Kumar, “Word Sense Disambiguation in the Hindi Language: Neural Network Approach,” International Journal of Technical Research and Science, pp. 72-76, 2021.
[CrossRef] [Google Scholar] [Publisher Link]

[24] Nazreena Rahman, and Bhogeswar Borah, “An Unsupervised Method for Word Sense Disambiguation,” Journal of King Saud University-Computer and Information Sciences, vol. 34, no. 9, pp. 6643-6651, 2022.
[CrossRef] [Google Scholar] [Publisher Link]

[25] Anh-Cuong Le et al., “Semi-Supervised Learning Integrated with Classifier Combination for Word Sense Disambiguation,” Computer Speech and Language, vol. 22, no. 4, pp. 330-345, 2008.
[CrossRef] [Google Scholar] [Publisher Link]

[26] Chandrakant D. Kokane, and Sachin D. Babar, “Supervised Word Sense Disambiguation with Recurrent Neural Network Model,” International Journal of Engineering and Advanced Technology, vol. 9, no. 2, pp. 1447-1453, 2019.
[CrossRef] [Google Scholar] [Publisher Link]

[27] Huei-Ling Lai et al., “Supervised Word Sense Disambiguation on Polysemy with Neural Network Models: A Case Study of BUN in Taiwan Hakka,” International Journal of Asian Language Processing, vol. 30, no. 3, 2020.
[CrossRef] [Google Scholar] [Publisher Link]

[28] Anup Kumar Barman et al., “Word Sense Disambiguation Applied to Assamese-Hindi Bilingual Statistical Machine Translation,” Engineering, Technology and Applied Science Research, vol. 14, no. 1, pp. 12581-12586, 2024.
[CrossRef] [Google Scholar] [Publisher Link]

[29] Debapratim Das Dawn et al., “Lexeme Connexion Measure of Cohesive Lexical Ambiguity Revealing Factor: A Robust Approach for Word Sense Disambiguation of Bengali Text,” Multimedia Tools and Applications, vol. 83, no. 5, pp. 12939-12983, 2024.
[CrossRef] [Google Scholar] [Publisher Link]

[30] Binod Kumar Mishra, and Suresh Jain, “An Innovative Method for Hindi Word Sense Disambiguation,” SN Computer Science, vol. 4, no. 6, pp. 1-17, 2023.
[CrossRef] [Google Scholar] [Publisher Link]

[31] Debapratim Das Dawn et al., “A Dataset for Evaluating Bengali Word Sense Disambiguation Techniques,” Journal of Ambient Intelligence and Humanized Computing, vol. 14, no. 4, pp. 4057-4086, 2023.
[CrossRef] [Google Scholar] [Publisher Link]

[32] Binod Kumar Mishra, and Suresh Jain, “Word Sense Disambiguation for Indic Language using Bi-LSTM,” Multimedia Tools and Applications, vol. 84, no. 16, pp. 16631-16656, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[33] Chhaya S. Patil, and Vaishali B. Patil, “A Multilingual Exploration of Word Sense Disambiguation using Transformer Models: Dravidian and Devanagari Languages, 1st ed., Recent Advances in Computing Sciences, CRC Press, pp. 158-160, 2025.
[Google Scholar] [Publisher Link]

[34] Sunjae Kwon et al., “Vision Meets Definitions: Unsupervised Visual Word Sense Disambiguation Incorporating Gloss Information,” Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, pp. 1583-1598, 2023.
[CrossRef] [Google Scholar] [Publisher Link]

[35] Robbel Habtamu Yigzaw, Beakal Gizachew Assefa, and Elefelious Getachew Belay, “A Hybrid Contextual Embedding and Hierarchical Attention for Improving the Performance of Word Sense Disambiguation,” IEEE Access, vol. 13, pp. 21744-21758, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[36] Hlaudi Daniel Masethe et al., “Hybrid Transformer-based Large Language Models for Word Sense Disambiguation in the Low-Resource Sesotho sa Leboa Language,” Applied Sciences, vol. 15, no. 7, pp. 1-33, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[37] Tawseef Ahmad Mir, and Aadil Ahmad Lawaye, “Word Sense Disambiguation Corpus for Kashmiri,” Natural Language Processing, vol. 31, no. 2, pp. 631-654, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[38] Ali Osman Mohammed Salih et al., “Enhancing a Context-based Supervised Machine Learning Model for Accurate Disambiguation of English Polysemous Word Senses,” Scientific Culture, vol. 11, no. 4, 1945.
[Google Scholar]

[39] Shailendra Kumar Patel, Rakesh Kumar, and Anuj Kumar Sirohi, “BERT for Hindi Word Sense Disambiguation, 1st ed., Intelligent Computing and Communication Techniques, CRC Press, pp. 747-753, 2025.
[Google Scholar] [Publisher Link]

[40] Madhuri Karnik et al., “State of the Art Analysis of Word Sense Disambiguation,” Intelligent Computing for Sustainable Development, Springer, Cham, pp. 55-70, 2024.
[CrossRef] [Google Scholar] [Publisher Link]

[41] Aditi Barai et al., “A Comprehensive Analysis of Word Sense Disambiguation in a Regional Language,” 2025 International Conference on Inventive Computation Technologies, Kirtipur, Nepal, pp. 1260-1265, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[42] Ratul Das, Alok Ranjan Pal, and Diganta Saha, “Unsupervised Approach for Word Sense Disambiguation in Bengali,” Computational Technologies and Electronics, Springer, Cham, pp. 196-206, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[43] Youddha Beer Singh et al., A Handbook of Computational Linguistics: Artificial Intelligence in Natural Language Processing, Bentham Science Publishers, 2024.
[CrossRef] [Google Scholar] [Publisher Link]

[44] Lianwang Hao, Tao Zhang, and Huaixin Liang, “A Novel Rotational Causal Random Forest Approach for Word Sense Disambiguation of English Modal Verbs,” SSRN, pp. 1-9, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[45] Aarti Purohit, and Kuldeep Kumar Yogi, “Enhancing Word Sense Disambiguation for Hindi Agriculture Domain: Feature Engineering and Machine Learning Approaches,” Results in Engineering, vol. 28, pp. 1-12, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[46] Anirudh S Nair et al., “Improving Word Sense Disambiguation by Adopting Refined Algorithms,” 2024 Third International Conference on Trends in Electrical, Electronics, and Computer Engineering, Bangalore, India, pp. 135-140, 2024.
[CrossRef] [Google Scholar] [Publisher Link]

[47] Vivek A. Manwar, and A.B. Manwar, “mBERT: A Query Refinement Model for Marathi Word Sense Disambiguation,” 2025 IEEE International Students' Conference on Electrical, Electronics and Computer Science, Bhopal, India, pp. 1-6, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[48] Pawan Makhija, and Sanjay Tanwani, “An Empirical Study on Various Word Sense Disambiguation Techniques in the Biomedical Domain,” International Conference on Recent Advancements and Modernisations in Sustainable Intelligent Technologies and Applications, Indore, India, 2025.
[CrossRef] [Google Scholar] [Publisher Link]