paper-with-me

홈 › Papers

Focused Large Language Models are Stable Many-Shot Learners

2024-08-26 · Peiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Yueqi Zhang, Chuyi Tan, Boyuan Pan, HeDa Wang, Yao Hu, Kan Li

In-Context Learning (ICL) enables large language models (LLMs) to achieve rapid task adaptation by learning from demonstrations. With the increase in available context length of LLMs, recent experiments have shown that the performance of ICL does not necessarily scale well in many-shot (demonstration) settings. We theoretically and experimentally confirm that the reason lies in more demonstrations dispersing the model attention from the query, hindering its understanding of key content. Inspired by how humans learn from examples, we propose a training-free method FocusICL, which conducts triviality filtering to avoid attention being diverted by unimportant contents at token-level and operates hierarchical attention to further ensure sufficient attention towards current query at demonstration-level. We also design an efficient hyperparameter searching strategy for FocusICL based on model perplexity of demonstrations. Comprehensive experiments validate that FocusICL achieves an average performance improvement of 5.2% over vanilla ICL and scales well with many-shot demonstrations.

📄 PDF Abstract BibTeX arXiv:2408.13987

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context Learning

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Retcon -- a Prompt-Based Technique for Precise Control of LLMs in Conversations

2026-02-09 · David Kogan, Sam Nguyen, Masanori Suzuki, Feiyang Chen arxiv

Recent advances in Large Language Models (LLMs) allow agents to execute complex natural language tasks. Many LLM applications, such as support agents, teaching assistants, and interactive bots, involve multi-turn convers…

Post-Hoc Merging is Not Enough: Many-Shot Model Merging with Loss-Gap Balancing

2026-06-15 · Kyungjin Im, Miru Kim, Chanin Eom, Minhae Kwon arxiv

Model merging has become a practical post-training strategy for building a single multi-task large language model (LLM) by combining multiple task-specialized models. However, most existing approaches rely on post-hoc me…

Augmenters at SemEval-2023 Task 1: Enhancing CLIP in Handling Compositionality and Ambiguity for Zero-Shot Visual WSD through Prompt Augmentation and Text-To-Image Diffusion

2023-07-09 · Jie S. Li, Yow-Ting Shiue, Yong-Siang Shih, Jonas Geiping

This paper describes our zero-shot approaches for the Visual Word Sense Disambiguation (VWSD) Task in English. Our preliminary study shows that the simple approach of matching candidate images with the phrase using CLIP …

DescriptiveWord Sense Disambiguation

Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models

2024-08-08 · Fabio Pernisi, Dirk Hovy, Paul Röttger

As diverse linguistic communities and users adopt large language models (LLMs), assessing their safety across languages becomes critical. Despite ongoing efforts to make LLMs safe, they can still be made to behave unsafe…

MDIA: A Benchmark for Multilingual Dialogue Generation in 46 Languages

2022-08-27 · Qingyu Zhang, Xiaoyu Shen, Ernie Chang, Jidong Ge 외

Owing to the lack of corpora for low-resource languages, current works on dialogue generation have mainly focused on English. In this paper, we present mDIA, the first large-scale multilingual benchmark for dialogue gene…

ChatbotDialogue GenerationDiversity