BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension
Abstract
Document comprehension is a challenging yet impactful taskfor Multimodal Large Language Models, especially as these systems seegrowing adoption in real-world, human-centric applications. However,this adoption is limited for low-resource languages such as Bangla dueto the scarcity of high-quality annotated data. To address this gap, weintroduce BaFCo, a benchmark dataset for Bangla form comprehensionwith a focus on Document Layout Analysis (DLA) and Key InformationExtraction (KIE). BaFCo curates 200 multi-page complex Bangladeshigovernment forms, sourced from across diverse sectors including agricul-ture, education, banking, and land management. To accurately capturethe structural and contextual complexity of these forms, we define a fine-grained annotation schema comprising 26 types of form entities, alongwith a separate coarse form entity set consisting of 5 types. We evaluatethe latest MLLMs from the ChatGPT, Gemini, Claude, Qwen, and Kimiseries using zero-shot and chain-of-thought prompts under both low andhigh reasoning setups. Our results reveal limitations in current MLLMs’ability in comprehending Bangla forms, particularly in accurately local-izing highly granular form entities. Our dataset and code is available at:https://huggingface.co/datasets/Mausul/bafco