paper-with-me

홈 › Papers

If the Sources Could Talk: Evaluating Large Language Models for Research Assistance in History

2023-10-16 · Giselle Gonzalez Garcia, Christian Weilbach

The recent advent of powerful Large-Language Models (LLM) provides a new conversational form of inquiry into historical memory (or, training data, in this case). We show that by augmenting such LLMs with vector embeddings from highly specialized academic sources, a conversational methodology can be made accessible to historians and other researchers in the Humanities. Concretely, we evaluate and demonstrate how LLMs have the ability of assisting researchers while they examine a customized corpora of different types of documents, including, but not exclusive to: (1). primary sources, (2). secondary sources written by experts, and (3). the combination of these two. Compared to established search interfaces for digital catalogues, such as metadata and full-text search, we evaluate the richer conversational style of LLMs on the performance of two main types of tasks: (1). question-answering, and (2). extraction and organization of data. We demonstrate that LLMs semantic retrieval and reasoning abilities on problem-specific tasks can be applied to large textual archives that have not been part of the its training data. Therefore, LLMs can be augmented with sources relevant to specific research projects, and can be queried privately by researchers.

📄 PDF Abstract BibTeX arXiv:2310.10808

Code (1)

gissygonzalez/kleiogpt 공식 구현 pytorch

Tasks

Question AnsweringRetrievalSemantic Retrieval

Similar Papers 제목 키워드 기반

Objectively Evaluating the Reliability of Cell Type Annotation Using LLM-Based Strategies

2024-09-24 · Wenjin Ye, Yuanchen Ma, Junkai Xiang, Hongjie Liang 외

Reliability in cell type annotation is challenging in single-cell RNA-sequencing data analysis because both expert-driven and automated methods can be biased or constrained by their training data, especially for novel or…

Language ModelingLanguage ModellingLarge Language Model

SMAC-Talk: A Natural Language Extension of the StarCraft Multi-Agent Challenge for Large Language Models

2026-06-02 · Joel Sol, Homayoun Najjaran arxiv

As LLMs become more widely deployed, they are increasingly expected to work alongside other AI agents rather than operating in isolation. Effective coordination in these settings requires agents to communicate, share inf…

Decision Making

A Knowledge Plug-and-Play Test Bed for Open-domain Dialogue Generation

2024-03-06 · Xiangci Li, Linfeng Song, Lifeng Jin, Haitao Mi 외

Knowledge-based, open-domain dialogue generation aims to build chit-chat systems that talk to humans using mined support knowledge. Many types and sources of knowledge have previously been shown to be useful as support k…

Dialogue GenerationResponse Generation

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

2026-07-25 · Sahil Deepak Gawande, Mayank Singh arxiv

Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between Engl…

Dialogue Generation

Can Language Models Make Fun? A Case Study in Chinese Comical Crosstalk

2022-07-02 · Benyou Wang, Xiangbo Wu, Xiaokang Liu, Jianquan Li 외

Language is the principal tool for human communication, in which humor is one of the most attractive parts. Producing natural language like humans using computers, a.k.a, Natural Language Generation (NLG), has been widel…

BenchmarkingMachine TranslationText Generation