DOCMASTER: A Unified Platform for Annotation, Training, & Inference in Document Question-Answering
The application of natural language processing models to PDF documents is pivotal for various business applications yet the challenge of training models for this purpose persists in businesses due to specific hurdles. These include the complexity of working with PDF formats that necessitate parsing text and layout information for curating training data and the lack of privacy-preserving annotation tools. This paper introduces DOCMASTER, a unified platform designed for annotating PDF documents, model training, and inference, tailored to document question-answering. The annotation interface enables users to input questions and highlight text spans within the PDF file as answers, saving layout information and text spans accordingly. Furthermore, DOCMASTER supports both state-of-the-art layout-aware and text models for comprehensive training purposes. Importantly, as annotations, training, and inference occur on-device, it also safeguards privacy. The platform has been instrumental in driving several research prototypes concerning document analysis such as the AI assistant utilized by University of California San Diego's (UCSD) International Services and Engagement Office (ISEO) for processing a substantial volume of PDF documents.
Code (0)
등록된 구현이 없습니다.
Tasks
Privacy PreservingQuestion AnsweringSimilar Papers 제목 키워드 기반
DocMaster: A Hierarchical Structure-Aware System for Document Analysis
Leveraging large language models (LLMs) to analyze complex documents -- such as academic papers, technical manuals, and financial reports -- has emerged as a mainstream and critical task in both research and industry. In…
Question AnsweringThresh: A Unified, Customizable and Deployable Platform for Fine-Grained Text Evaluation
Fine-grained, span-level human evaluation has emerged as a reliable and robust method for evaluating text generation tasks such as summarization, simplification, machine translation and news generation, and the derived a…
Machine TranslationMulti-Task LearningNews GenerationText GenerationAwakeForest: An Interactive Geospatial Platform for Large-Scale Forest Imagery
Forest imagery analysis often involves multiple tightly coupled vision tasks, which must be performed under substantial variation in geographic regions, sensors, and acquisition conditions. However, practitioners often l…
DocSpiral: A Platform for Integrated Assistive Document Annotation through Human-in-the-Spiral
Acquiring structured data from domain-specific, image-based documents such as scanned reports is crucial for many downstream tasks but remains challenging due to document variability. Many of these documents exist as ima…
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Automating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities. We introduce Aguvis, a unified vision-based framework for autonomo…
Multimodal ReasoningNatural Language Visual GroundingNavigate