paper-with-me

answerability prediction

1개 벤치마크 · 논문 9편 · 이 태스크의 논문 보기 →

Benchmarks

PeerQA

결과 6개

Most implemented

Language Models are Few-Shot Learners

2020-05-28 · 구현 67개

GPT-4 Technical Report

2023-03-15 · 구현 11개

Mistral 7B

2023-10-10 · 구현 6개

The Llama 3 Herd of Models

2024-07-31 · 구현 5개

Papers

PeerQA: A Scientific Question Answering Dataset from Peer Reviews

2025-02-19 · Tim Baumgärtner, Ted Briscoe, Iryna Gurevych

We present PeerQA, a real-world, scientific, document-level Question Answering (QA) dataset. PeerQA questions have been sourced from peer reviews, which contain questions that reviewers raised while thoroughly examining …

answerability predictionAnswer GenerationArticlesPassage Retrieval+5

The Llama 3 Herd of Models

2024-07-31 · Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey 외

Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, cod…

answerability predictionLanguage ModelingLanguage ModellingMulti-task Language Understanding+3

Towards Reliable and Factual Response Generation: Detecting Unanswerable Questions in Information-Seeking Conversations

2024-01-21 · Weronika Łajewska, Krisztian Balog

Generative AI models face the challenge of hallucinations that can undermine users' trust in such systems. We approach the problem of conversational information seeking as a two-step process, where relevant passages in a…

answerability predictionResponse Generation

Mistral 7B

2023-10-10 · Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford 외

We introduce Mistral 7B v0.1, a 7-billion-parameter language model engineered for superior performance and efficiency. Mistral 7B outperforms Llama 2 13B across all evaluated benchmarks, and Llama 1 34B in reasoning, mat…

answerability predictionArithmetic ReasoningChatbotCode Generation+11

GPT-4 Technical Report

2023-03-15 · Preprint 2023 3 · OpenAI, :, Josh Achiam, Steven Adler 외

We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs. While less capable than humans in many real-world scenarios, GPT-4 exhibits human-level…

answerability predictionArithmetic ReasoningBug fixingCode Generation+18

Towards Confident Machine Reading Comprehension

2021-01-20 · Rishav Chakravarti, Avirup Sil

There has been considerable progress on academic benchmarks for the Reading Comprehension (RC) task with State-of-the-Art models closing the gap with human performance on extractive question answering. Datasets such as S…

answerability predictionExtractive Question-AnsweringMachine Reading ComprehensionPrediction+2

전체 9편 보기 →