HATEBENCH
- cases
- 50
- models
- 34
- providers
- 17
- predictions
- 1,700
Measuring frontier language models on Hindi and Hinglish hate-speech classification
Leaderboard
HateBench score
Run cost (log)
most efficient ↖
Top 10 of 34
Pass@1 · 95% CI
All models receive the identical system prompt and JSON output contract via one runner. Read why.
English moderation benchmarks are saturated, and they say little about how a model behaves on the text Indian platforms actually receive. HateBench is a Hindi and Hinglish hate-speech benchmark built to separate frontier models on that text. It brings four things to the table:
- Real-world text: Cases are romanized Hinglish and Devanagari social-media comments with slang, emoji, and coded abuse, not sanitized English templates.
- One runner, one contract: Every model sees the identical system prompt and must return the same strict JSON. Scores reflect the model, not the scaffolding.
- Full accounting: Each run records cost, input and output tokens, and latency per model, so accuracy is always shown next to what it took to get there.
- Honest intervals: Each score ships with a 95% confidence interval, so close models read as close instead of being separated by noise.
The result is a leaderboard that reflects how models actually perform on Hindi moderation work, with the cost of every score attached.
Write-up
Read the full blog
The methodology, the runner, the metrics engine, and what 67 models reveal about hate-speech detection in Hindi — written up in full on the author's site.
The dataset and per-model predictions are not distributed with this site. If you need access to the data for research, please email me.