Megadose AI progress, ranked and analyzed.

BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models

· ArXiv · AI/CL/LG ·
The benchmark finds visual mis-recognition, not Bengali conjunct handling, is the main failure mode.

BanglaWild contains 2,535 in-the-wild Bengali scene text images with gold transcriptions and diagnostic labels. The paper evaluates 15 vision-language models and three OCR systems, plus LoRA fine-tuning on six open models. Its error taxonomy says visual mis-recognition makes up about 60% of errors in the strongest systems, while conjunct-related errors stay under 2%. Larger models within the same family did not consistently beat smaller ones. ArXiv · AI/CL/LG's note

score 4

Categories: Research