paper-with-me

홈 › Papers

Evaluation of medium-large Language Models at zero-shot closed book generative question answering

2023-05-19 · René Peinl, Johannes Wirth

Large language models (LLMs) have garnered significant attention, but the definition of "large" lacks clarity. This paper focuses on medium-sized language models (MLMs), defined as having at least six billion parameters but less than 100 billion. The study evaluates MLMs regarding zero-shot generative question answering, which requires models to provide elaborate answers without external document retrieval. The paper introduces an own test dataset and presents results from human evaluation. Results show that combining the best answers from different MLMs yielded an overall correct answer rate of 82.7% which is better than the 60.9% of ChatGPT. The best MLM achieved 71.8% and has 33B parameters, which highlights the importance of using appropriate training data for fine-tuning rather than solely relying on the number of parameters. More fine-grained feedback should be used to further improve the quality of answers. The open source community is quickly closing the gap to the best commercial models.

📄 PDF Abstract BibTeX arXiv:2305.11991

Code (0)

등록된 구현이 없습니다.

Tasks

Generative Question AnsweringQuestion AnsweringRetrieval

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

On the Zero-Shot Generalization of Machine-Generated Text Detectors

2023-10-08 · Xiao Pu, Jingyu Zhang, Xiaochuang Han, Yulia Tsvetkov 외

The rampant proliferation of large language models, fluent enough to generate text indistinguishable from human-written language, gives unprecedented importance to the detection of machine-generated text. This work is mo…

Zero-shot Generalization

XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

2024-06-07 · Edresson Casanova, Kelly Davis, Eren Gölge, Görkem Göknar 외

Most Zero-shot Multi-speaker TTS (ZS-TTS) systems support only a single language. Although models like YourTTS, VALL-E X, Mega-TTS 2, and Voicebox explored Multilingual ZS-TTS they are limited to just a few high/medium r…

text-to-speechText to SpeechVoice CloningZero-Shot Multi-Speaker TTS

Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon

2024-02-03 · Fajri Koto, Tilman Beck, Zeerak Talat, Iryna Gurevych 외

Improving multilingual language models capabilities in low-resource languages is generally difficult due to the scarcity of large-scale data in those languages. In this paper, we relax the reliance on texts in low-resour…

SentenceSentiment Analysis

Developing and Evaluating Tiny to Medium-Sized Turkish BERT Models

2023-07-26 · Himmet Toprak Kesgin, Muzaffer Kaan Yuce, Mehmet Fatih Amasyali

This study introduces and evaluates tiny, mini, small, and medium-sized uncased Turkish BERT models, aiming to bridge the research gap in less-resourced languages. We trained these models on a diverse dataset encompassin…

ClassificationComputational EfficiencyNews ClassificationSentiment Analysis+2

Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation

2026-04-06 · Hanif Rahman arxiv

Pashto is spoken by approximately 60--80 million people but has no published benchmarks for multilingual automatic speech recognition (ASR) on any shared public test set. This paper reports the first reproducible multi-m…

Speech Recognition