Detecting Hate Speech in Hindi Digital Discourse Using Transformer–Long Short-Term Memory Models

Narula, Rachna, Gupta, Vedika ORCID: https://orcid.org/0000-0002-8109-498X, Khan, Jawad, Rahim, Tariq and Kumar, Vijay (2026) Detecting Hate Speech in Hindi Digital Discourse Using Transformer–Long Short-Term Memory Models. Journal of Communication Language and Culture, 6 (2). pp. 303-325. Multimedia University Press . ISSN 2805-444X Available at: https://doi.org/10.33093/jclc.2026.6.2.16

[thumbnail of Detecting Hate Speech in Hindi Digital Discourse.pdf]
Preview
Text
Detecting Hate Speech in Hindi Digital Discourse.pdf - Published Version
Available under License Creative Commons Attribution Non-commercial No Derivatives.

Download (1MB) | Preview

Abstract

Hate speech on digital communication platforms has become a major obstacle to healthy online discourse, especially in multilingual societies such as India, where Hindi is a dominant language in social media interactions. Hostile, offensive, and defamatory speech is linguistically and socio-culturally complex because of colloquial idioms, regional variations, code-mixing, and culturally embedded references. However, effective detection is crucial for creating safer and more inclusive digital communication environments. This study evaluates advanced language models for analysing hate speech in Hindi social media content. A dataset of 21000 Hindi posts from Twitter and public repositories, categorized into five categories, was analysed: hate, offensive, fake, defamation, and non-hostile. General-purpose models were tested against Hindi-specific language models, including Hindi-BERT (based on Bidirectional Encoder Representations from Transformers, or BERT) and MuRIL, to investigate whether performance can be further enhanced by integrating Long Short-Term Memory (LSTM) layers. In tests, the best-performing model, Hindi-BERT (with sequential learning added), correctly identified 94 out of every 100 posts—a considerable improvement over simpler models. For social media platforms, it has real-world consequences: it can automatically identify harmful content to be reviewed by a human, minimize nuisance alerts that waste human moderators' time, and identify defamatory or bogus posts before they gain a lot of traction. The results offer a systematic approach to researching how hostility develops in the Hindi-speaking online community, how linguistic creativity (e.g., slang, sarcasm, code-mixing) can conceal or manifest hostility, and how decisions about content moderation influence public discourse for communication scholars. Overall, this paper illustrates that language-specific computational tools can be used for both platform governance and communication research, provided that the cultural context is considered. Finally, technical methods are combined with communication scholarship to explain and curb harmful speech in the online public sphere.

Item Type: Article
Uncontrolled Keywords: India | Electoral reform | Political challenges | Federalism | Regional parties | Governance
Subjects: Social Sciences and humanities > Arts and Humanities > Language and Linguistics
Physical, Life and Health Sciences > Computer Science
Social Sciences and humanities > Social Sciences > Behavioral Studies
Vol/Issue no. published date: July 2026
Depositing User: Mr. Syed Anas Ali
Date Deposited: 11 Aug 2026 07:02
Last Modified: 11 Aug 2026 07:03
Official URL: https://doi.org/10.33093/jclc.2026.6.2.16
URI: https://pure.jgu.edu.in/id/eprint/12181

Downloads

Downloads per month over past year

Actions (login required)

View Item
View Item