paper-with-me

홈 › Papers

AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation

2025-09-21 · Tiancheng Huang, Ruisheng Cao, Yuxin Zhang, Zhangyi Kang, Zijian Wang, Chenrun Wang, Yijie Luo, Hang Zheng, Lirong Qian, Lu Chen, Kai Yu arxiv

The growing volume of academic papers has made it increasingly difficult for researchers to efficiently extract key information. While large language models (LLMs) based agents are capable of automating question answering (QA) workflows for scientific papers, there still lacks a comprehensive and realistic benchmark to evaluate their capabilities. Moreover, training an interactive agent for this specific task is hindered by the shortage of high-quality interaction trajectories. In this work, we propose AirQA, a human-annotated comprehensive paper QA dataset in the field of artificial intelligence (AI), with 13,956 papers and 1,246 questions, that encompasses multi-task, multi-modal and instance-level evaluation. Furthermore, we propose ExTrActor, an automated framework for instruction data synthesis. With three LLM-based agents, ExTrActor can perform example generation and trajectory collection without human intervention. Evaluations of multiple open-source and proprietary models show that most models underperform on AirQA, demonstrating the quality of our dataset. Extensive experiments confirm that ExTrActor consistently improves the multi-turn tool-use capability of small models, enabling them to achieve performance comparable to larger ones.

📄 PDF Abstract BibTeX arXiv:2509.16952

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

NeuSym-RAG: Hybrid Neural Symbolic Retrieval with Multiview Structuring for PDF Question Answering

2025-05-26 · Ruisheng Cao, Hanchong Zhang, Tiancheng Huang, Zhangyi Kang 외

The increasing number of academic papers poses significant challenges for researchers to efficiently acquire key details. While retrieval augmented generation (RAG) shows great promise in large language model (LLM) based…

ChunkingLarge Language ModelQuestion AnsweringRAG+2

PWISeg: Point-based Weakly-supervised Instance Segmentation for Surgical Instruments

2023-11-16 · Zhen Sun, Huan Xu, Jinlin Wu, Zhen Chen 외

In surgical procedures, correct instrument counting is essential. Instance segmentation is a location method that locates not only an object's bounding box but also each pixel's specific details. However, obtaining mask-…

Instance SegmentationSegmentationSemantic SegmentationSurgical tool detection+2

Pose2Seg: Detection Free Human Instance Segmentation

2018-03-28 · CVPR 2019 6 · Song-Hai Zhang, Rui-Long Li, Xin Dong, Paul L. Rosin 외

The standard approach to image instance segmentation is to perform the object detection first, and then segment the object from the detection bounding-box. More recently, deep learning methods like Mask R-CNN perform the…

2D Human Pose EstimationHuman Instance SegmentationInstance SegmentationKeypoint Detection+6

SIT-FER: Integration of Semantic-, Instance-, Text-level Information for Semi-supervised Facial Expression Recognition

2025-03-24 · Sixian Ding, Xu Jiang, Zhongjing Du, Jiaqi Cui 외

Semi-supervised deep facial expression recognition (SS-DFER) has gained increasingly research interest due to the difficulty in accessing sufficient labeled data in practical settings. However, existing SS-DFER methods m…

Facial Expression Recognition

AppealCase: A Dataset and Benchmark for Civil Case Appeal Scenarios

2025-05-22 · YuTing Huang, Meitong Guo, Yiquan Wu, Ang Li 외

Recent advances in LegalAI have primarily focused on individual case judgment analysis, often overlooking the critical appellate process within the judicial system. Appeals serve as a core mechanism for error correction …

Decision MakingMulti-class ClassificationText ClassificationText Generation