paper-with-me

Papers

How Robust are the Tabular QA Models for Scientific Tables? A Study using Customized Dataset

2024-03-30 · Akash Ghosh, B Venkata Sahith, Niloy Ganguly, Pawan Goyal, Mayank Singh

Question-answering (QA) on hybrid scientific tabular and textual data deals with scientific information, and relies on complex numerical reasoning. In recent years, while tabular QA has seen rapid progress, understanding their robustness on scientific information is lacking due to absence of any benchmark dataset. To investigate the robustness of the existing state-of-the-art QA models on scientific hybrid tabular data, we propose a new dataset, "SciTabQA", consisting of 822 question-answer pairs from scientific tables and their descriptions. With the help of this dataset, we assess the state-of-the-art Tabular QA models based on their ability (i) to use heterogeneous information requiring both structured data (table) and unstructured data (text) and (ii) to perform complex scientific reasoning tasks. In essence, we check the capability of the models to interpret scientific tables and text. Our experiments show that "SciTabQA" is an innovative dataset to study question-answering over scientific heterogeneous data. We benchmark three state-of-the-art Tabular QA models, and find that the best F1 score is only 0.462.

📄 PDF Abstract BibTeX arXiv:2404.00401

Code (1)

akash-ghosh-123/scitabqa 공식 구현

Tasks

Question Answering

Similar Papers 제목 키워드 기반

SemEval-2021 Task 9: Fact Verification and Evidence Finding for Tabular Data in Scientific Documents (SEM-TAB-FACTS)

2021-05-28 · SEMEVAL 2021 · Nancy X. R. Wang, Diwakar Mahajan, Marina Danilevsk. Sara Rosenthal

Understanding tables is an important and relevant task that involves understanding table structure as well as being able to compare and contrast information within cells. In this paper, we address this challenge by prese…

Fact Verification

TabLeX: A Benchmark Dataset for Structure and Content Information Extraction from Scientific Tables

2021-05-12 · Harsh Desai, Pratik Kayal, Mayank Singh

Information Extraction (IE) from the tables present in scientific articles is challenging due to complicated tabular representations and complex embedded text. This paper presents TabLeX, a large-scale benchmark dataset …

ArticlesTable Extraction

Language Model Representations for Efficient Few-Shot Tabular Classification

2026-01-21 · Inwon Kang, Parikshit Ram, Yi Zhou, Horst Samulowitz 외 arxiv

The Web is a rich source of structured data in the form of tables, from product catalogs and knowledge bases to scientific datasets. However, the heterogeneity of the structure and semantics of these tables makes it chal…

Tables to LaTeX: structure and content extraction from scientific tables

2022-10-31 · Pratik Kayal, Mrinal Anand, Harsh Desai, Mayank Singh

Scientific documents contain tables that list important information in a concise fashion. Structure and content extraction from tables embedded within PDF research documents is a very challenging task due to the existenc…

Language ModelingLanguage Modelling

iTBLS: A Dataset of Interactive Conversations Over Tabular Information

2024-04-19 · Anirudh Sundar, Christopher Richardson, William Gay, Larry Heck

This paper introduces Interactive Tables (iTBLS), a dataset of interactive conversations situated in tables from scientific articles. This dataset is designed to facilitate human-AI collaborative problem-solving through …

ArticlesMathematical Reasoningparameter-efficient fine-tuning