TamHateX
Description
This repository contains TamHateX, a hierarchically annotated Tamil hate-speech dataset collected from public YouTube comments between January and April 2026. Annotation was completed on 20 July 2026. The dataset contains 5,319 comments, divided into 4,255 training, 532 validation, and 532 test records. It includes Tamil script, Romanised Tamil, English, and Tamil–English code-mixed comments, preserving informal spelling, slang, and social-media language. Each comment is annotated at five levels: Level 1: Hate or Non-Hate Level 2: Target of the expression Level 3: Explicit or Implicit hate Level 4: Fine-grained abuse category Level 5: Social-bias category The dataset supports research on Tamil hate-speech detection, target identification, implicit and explicit hate recognition, fine-grained abuse classification, social-bias analysis, hierarchical classification, and knowledge distillation. All comments were de-identified and are provided only for academic research.