Augmenting Hate Speech Detection with Hybrid CNN-RNN Models and Dataset Distribution Techniques
Identifying hate speech is an important element in controlling toxicity and allowing people to function in digital spaces. In this work, we scale hybrid CNN-RNN architectures using augmentation and carefully balancing datasets to achieve better classification accuracy of hate speech. Unlike previous works that grappled with the one-sided persistent fight against imbalanced datasets, our work builds balanced datasets through synthetic oversampling like SMOTE, under sampling, and modern augmentation strategies that create stronger distributions for training. Moreover, we integrate other state-of-the-art machine and deep learning systems, including pre-trained language models and explainable AI. Explainable AI techniques helped enable transparency by providing insights into the flagged content which builds trust. This work achieves better accuracy (up to 0.908) and F1 scores (up to 0.914) while maintaining computational efficiency for real-time deployment. Bias mitigation as well as ethical frameworks for adaptation within torture communities strengthens ethical concerns. The societal interest motivates development into large-scale, multi-ligneous, big-impact hate-speech detection frameworks.