Bangla OCR Triple Threat 🔥 Benchmark

A same-test leaderboard for full-page Bangla handwriting recognition across 6,669 Bongabdo robustness renderings, 6,669 BN-HTRd test renderings, and 6,669 novel Chaos compositions.

WER · word error CER · character error Full-page HTRWriter-separated evaluation

Each dataset score is 50% WER + 50% CER; Overall 🔥 is the equal weighted aggregate: 20% Bongabdo + 20% BN-HTRd + 60% Chaos. Lower is better. No semantic-importance metric is used in this first release. Avg decode (ms) is mean latency per page; API timings include provider/network latency, while local batched timings are amortized per page. 🔒 marks closed paid APIs.

Frozen public benchmark: dataset and provenance manifests · Sources: Bongabdo · BN-HTRd Splitted

🔥 Best Overall 🥇 🔒 google/gemini-3.1-flash-lite · 22.3520% + 20% + 60% · WER/CER only
Frozen benchmark3 × 6,669 pages Bongabdo · BN-HTRd · Chaos

Sortable leaderboard — ranked by Overall

Sortable leaderboard — ranked by Overall
1
22.35
17.78
10.71
12.03
5.48
36.75
22.42
3732.4
Gemini
complete
OpenRouter; temperature 0; 20,007 predictions complete; Overall = 20% Bongabdo + 20% BN-HTRd + 60% private-before-eval Chaos.

Submit a model

Run the model on all three 6,669-row splits in the frozen public dataset. Submit UTF-8 JSONL with IDs bongabdo:0bongabdo:6668 and bn_htrd:0bn_htrd:6668. Keep raw model output; Bangla Unicode and whitespace normalization is applied centrally and identically to every submission.

Open a Community discussion with predictions, exact model/API revision, decoding configuration, runtime hardware, and whether inference was local or provider-hosted.