paper-with-me

홈 › Papers

DocXChain: A Powerful Open-Source Toolchain for Document Parsing and Beyond

2023-10-19 · Cong Yao

In this report, we introduce DocXChain, a powerful open-source toolchain for document parsing, which is designed and developed to automatically convert the rich information embodied in unstructured documents, such as text, tables and charts, into structured representations that are readable and manipulable by machines. Specifically, basic capabilities, including text detection, text recognition, table structure recognition and layout analysis, are provided. Upon these basic capabilities, we also build a set of fully functional pipelines for document parsing, i.e., general text reading, table parsing, and document structurization, to drive various applications related to documents in real-world scenarios. Moreover, DocXChain is concise, modularized and flexible, such that it can be readily integrated with existing tools, libraries or models (such as LangChain and ChatGPT), to construct more powerful systems that can accomplish more complicated and challenging tasks. The code of DocXChain is publicly available at:~\url{https://github.com/AlibabaResearch/AdvancedLiterateMachinery/tree/main/Applications/DocXChain}

📄 PDF Abstract BibTeX arXiv:2310.12430

Code (1)

alibabaresearch/advancedliteratemachinery 공식 구현 pytorch

Tasks

Document AIDocument Layout Analysisdocument understandingOptical Character Recognition (OCR)Scene Text DetectionScene Text RecognitionText Detection

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Transforming Sensitive Documents into Quantitative Data: An AI-Based Preprocessing Toolchain for Structured and Privacy-Conscious Analysis

2025-07-11 · Anders Ledberg, Anna Thalén arxiv

Unstructured text from legal, medical, and administrative sources offers a rich but underutilized resource for research in public health and the social sciences. However, large-scale analysis is hampered by two key chall…

SLGPT: Using Transfer Learning to Directly Generate Simulink Model Files and Find Bugs in the Simulink Toolchain

2021-05-16 · Sohil Lal Shrestha, Christoph Csallner

Finding bugs in a commercial cyber-physical system (CPS) development tool such as Simulink is hard as its codebase contains millions of lines of code and complete formal language specifications are not available. While d…

Deep LearningTransfer Learning

EDA Corpus: A Large Language Model Dataset for Enhanced Interaction with OpenROAD

2024-05-04 · Bing-Yue Wu, Utsav Sharma, Sai Rahul Dhanvi Kankipati, Ajay Yadav 외

Large language models (LLMs) serve as powerful tools for design, providing capabilities for both task automation and design assistance. Recent advancements have shown tremendous potential for facilitating LLM integration…

Language ModelingLanguage ModellingLarge Language Model

OpenFly: A Comprehensive Platform for Aerial Vision-Language Navigation

2025-02-25 · Yunpeng Gao, Chenhui Li, Zhongrui You, Junli Liu 외

Vision-Language Navigation (VLN) aims to guide agents by leveraging language instructions and visual cues, playing a pivotal role in embodied AI. Indoor VLN has been extensively studied, whereas outdoor aerial VLN remain…

BenchmarkingSemantic SegmentationVision-Language Navigation

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

2026-07-15 · Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu 외 hf

Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable t…

Action Segmentation