paper-with-me

홈 › Papers

Query of CC: Unearthing Large Scale Domain-Specific Knowledge from Public Corpora

2024-01-26 · Zhaoye Fei, Yunfan Shao, Linyang Li, Zhiyuan Zeng, Conghui He, Hang Yan, Dahua Lin, Xipeng Qiu

Large language models have demonstrated remarkable potential in various tasks, however, there remains a significant scarcity of open-source models and data for specific domains. Previous works have primarily focused on manually specifying resources and collecting high-quality data on specific domains, which significantly consume time and effort. To address this limitation, we propose an efficient data collection method $\textit{Query of CC}$ based on large language models. This method bootstraps seed information through a large language model and retrieves related data from public corpora. It not only collects knowledge-related data for specific domains but unearths the data with potential reasoning procedures. Through the application of this method, we have curated a high-quality dataset called KNOWLEDGE PILE, encompassing four major domains, including stem and humanities sciences, among others. Experimental results demonstrate that KNOWLEDGE PILE significantly improves the performance of large language models in mathematical and knowledge-related reasoning ability tests. To facilitate academic sharing, we open-source our dataset and code, providing valuable support to the academic community.

📄 PDF Abstract BibTeX arXiv:2401.14624

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

IEPile: Unearthing Large-Scale Schema-Based Information Extraction Corpus

2024-02-22 · Honghao Gui, Lin Yuan, Hongbin Ye, Ningyu Zhang 외

Large Language Models (LLMs) demonstrate remarkable potential across various domains; however, they exhibit a significant performance gap in Information Extraction (IE). Note that high-quality instruction data is the vit…

Zero-shot Generalization

Unearthing Common Inconsistency for Generalisable Deepfake Detection

2023-11-20 · Beilin Chu, Xuan Xu, Weike You, Linna Zhou

Deepfake has emerged for several years, yet efficient detection techniques could generalize over different manipulation methods require further research. While current image-level detection method fails to generalize to …

Contrastive LearningDeepFake DetectionFace SwappingInductive Bias

Visual Analytics for Efficient Image Exploration and User-Guided Image Captioning

2023-11-02 · Yiran Li, Junpeng Wang, Prince Aboagye, Michael Yeh 외

Recent advancements in pre-trained large-scale language-image models have ushered in a new era of visual comprehension, offering a significant leap forward. These breakthroughs have proven particularly instrumental in ad…

Caption GenerationEfficient ExplorationImage Captioning

OASum: Large-Scale Open Domain Aspect-based Summarization

2022-12-19 · Xianjun Yang, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 외

Aspect or query-based summarization has recently caught more attention, as it can generate differentiated summaries based on users' interests. However, the current dataset for aspect or query-based summarization either f…

QdaVPR: A novel query-based domain-agnostic model for visual place recognition

2026-03-08 · Shanshan Wan, Lai Kang, Yingmei Wei, Tianrui Shen 외 arxiv

Visual place recognition (VPR) aiming at predicting the location of an image based solely on its visual features is a fundamental task in robotics and autonomous systems. Domain variation remains one of the main challeng…

Visual Place RecognitionStyle Transfer