paper-with-me

홈 › Papers

IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing

2025-12-23 · Oikantik Nath, Sahithi Kukkala, Mitesh Khapra, Ravi Kiran Sarvadevabhatla arxiv

Document layout analysis is essential for downstream tasks such as information retrieval, extraction, OCR, and digitization. However, existing large-scale datasets like PubLayNet and DocBank lack fine-grained region labels and multilingual diversity, making them insufficient for representing complex document layouts. In contrast, human-annotated datasets such as M6Doc and D4LA offer richer labels and greater domain diversity, but are too small to train robust models and lack adequate multilingual coverage. This gap is especially pronounced for Indic documents, which encompass diverse scripts yet remain underrepresented in current datasets, further limiting progress in this space. To address these shortcomings, we introduce IndicDLP, a large-scale foundational document layout dataset spanning 11 representative Indic languages alongside English and 12 common document domains. Additionally, we curate UED-mini, a dataset derived from DocLayNet and M6Doc, to enhance pretraining and provide a solid foundation for Indic layout models. Our experiments demonstrate that fine-tuning existing English models on IndicDLP significantly boosts performance, validating its effectiveness. Moreover, models trained on IndicDLP generalize well beyond Indic layouts, making it a valuable resource for document digitization. This work bridges gaps in scale, diversity, and annotation granularity, driving inclusive and efficient document understanding.

📄 PDF Abstract BibTeX arXiv:2512.20236

Code (0)

등록된 구현이 없습니다.

Tasks

Document Layout AnalysisInformation Retrieval

Similar Papers 제목 키워드 기반

Monolingual or Multilingual Instruction Tuning: Which Makes a Better Alpaca

2023-09-16 · Pinzhen Chen, Shaoxiong Ji, Nikolay Bogoychev, Andrey Kutuzov 외

Foundational large language models (LLMs) can be instruction-tuned to perform open-domain question answering, facilitating applications like chat assistants. While such efforts are often carried out in a single language,…

Instruction FollowingLarge Language ModelMultilingual NLPOpen-Domain Question Answering+2

MultiLegalSBD: A Multilingual Legal Sentence Boundary Detection Dataset

2023-05-02 · Tobias Brugger, Matthias Stürmer, Joel Niklaus

Sentence Boundary Detection (SBD) is one of the foundational building blocks of Natural Language Processing (NLP), with incorrectly split sentences heavily influencing the output quality of downstream tasks. It is a chal…

Boundary DetectionSentence

Text2Cypher Across Languages: Evaluating Foundational Models Beyond English

2025-06-26 · Makbule Gulcin Ozsoy, William Tai

Recent advances in large language models have enabled natural language interfaces that translate user questions into database queries, such as Text2SQL, Text2SPARQL, and Text2Cypher. While these interfaces enhance databa…

AttributeText2Sparql

The Multilingual Mind : A Survey of Multilingual Reasoning in Language Models

2025-02-13 · Akash Ghosh, Debayan Datta, Sriparna Saha, Chirag Agarwal

While reasoning and multilingual capabilities in Language Models (LMs) have achieved remarkable progress in recent years, their integration into a unified paradigm, multilingual reasoning, is at a nascent stage. Multilin…

Logical ReasoningSurvey

MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder

2024-09-21 · Khai Le-Duc, Phuc Phan, Tan-Hanh Pham, Bach Phan Tat 외

Multilingual automatic speech recognition (ASR) in the medical domain serves as a foundational task for various downstream applications such as speech translation, spoken language understanding, and voice-activated assis…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderDiversity+3