Narula, Rachna, Gupta, Vedika
ORCID: https://orcid.org/0000-0002-8109-498X, Khan, Jawad, Rahim, Tariq and Kumar, Vijay
(2026)
Detecting Hate Speech in Hindi Digital Discourse Using Transformer–Long Short-Term Memory Models.
Journal of Communication Language and Culture, 6 (2).
pp. 303-325.
Multimedia University Press
.
ISSN 2805-444X
Available at: https://doi.org/10.33093/jclc.2026.6.2.16
Detecting Hate Speech in Hindi Digital Discourse.pdf - Published Version
Available under License Creative Commons Attribution Non-commercial No Derivatives.
Download (1MB) | Preview
Abstract
Hate speech on digital communication platforms has become a major obstacle to healthy online discourse, especially in multilingual societies such as India, where Hindi is a dominant language in social media interactions. Hostile, offensive, and defamatory speech is linguistically and socio-culturally complex because of colloquial idioms, regional variations, code-mixing, and culturally embedded references. However, effective detection is crucial for creating safer and more inclusive digital communication environments. This study evaluates advanced language models for analysing hate speech in Hindi social media content. A dataset of 21000 Hindi posts from Twitter and public repositories, categorized into five categories, was analysed: hate, offensive, fake, defamation, and non-hostile. General-purpose models were tested against Hindi-specific language models, including Hindi-BERT (based on Bidirectional Encoder Representations from Transformers, or BERT) and MuRIL, to investigate whether performance can be further enhanced by integrating Long Short-Term Memory (LSTM) layers. In tests, the best-performing model, Hindi-BERT (with sequential learning added), correctly identified 94 out of every 100 posts—a considerable improvement over simpler models. For social media platforms, it has real-world consequences: it can automatically identify harmful content to be reviewed by a human, minimize nuisance alerts that waste human moderators' time, and identify defamatory or bogus posts before they gain a lot of traction. The results offer a systematic approach to researching how hostility develops in the Hindi-speaking online community, how linguistic creativity (e.g., slang, sarcasm, code-mixing) can conceal or manifest hostility, and how decisions about content moderation influence public discourse for communication scholars. Overall, this paper illustrates that language-specific computational tools can be used for both platform governance and communication research, provided that the cultural context is considered. Finally, technical methods are combined with communication scholarship to explain and curb harmful speech in the online public sphere.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | India | Electoral reform | Political challenges | Federalism | Regional parties | Governance |
| Subjects: | Social Sciences and humanities > Arts and Humanities > Language and Linguistics Physical, Life and Health Sciences > Computer Science Social Sciences and humanities > Social Sciences > Behavioral Studies |
| Vol/Issue no. published date: | July 2026 |
| Depositing User: | Mr. Syed Anas Ali |
| Date Deposited: | 11 Aug 2026 07:02 |
| Last Modified: | 11 Aug 2026 07:03 |
| Official URL: | https://doi.org/10.33093/jclc.2026.6.2.16 |
| URI: | https://pure.jgu.edu.in/id/eprint/12181 |
Downloads
Downloads per month over past year
Dimensions
Dimensions