paper-with-me

홈 › Papers

From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning

2024-07-01 · Nan Xu, Fei Wang, Sheng Zhang, Hoifung Poon, Muhao Chen

Motivated by in-context learning (ICL) capabilities of Large Language models (LLMs), multimodal LLMs with additional visual modality are also exhibited with similar ICL abilities when multiple image-text pairs are provided as demonstrations. However, relatively less work has been done to investigate the principles behind how and why multimodal ICL works. We conduct a systematic and principled evaluation of multimodal ICL for models of different scales on a broad spectrum of new yet critical tasks. Through perturbations over different modality information, we show that modalities matter differently across tasks in multimodal ICL. Guided by task-specific modality impact, we recommend modality-driven demonstration strategies to boost ICL performance. We also find that models may follow inductive biases from multimodal ICL even if they are rarely seen in or contradict semantic priors from pretraining data. Our principled analysis provides a comprehensive way of understanding the role of demonstrations in multimodal in-context learning, and sheds light on effectively improving multimodal ICL on a wide range of tasks.

📄 PDF Abstract BibTeX arXiv:2407.00902

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context Learning

Similar Papers 제목 키워드 기반

Me, Myself, and $π$ : Evaluating and Explaining LLM Introspection

2026-03-17 · Atharv Naphade, Samarth Bhargav, Sean Lim, Mcnair Shah arxiv

A hallmark of human intelligence is Introspection-the ability to assess and reason about one's own cognitive processes. Introspection has emerged as a promising but contested capability in large language models (LLMs). H…

Can BERT eat RuCoLA? Topological Data Analysis to Explain

2023-04-04 · Irina Proskurina, Irina Piontkovskaya, Ekaterina Artemova

This paper investigates how Transformer language models (LMs) fine-tuned for acceptability classification capture linguistic features. Our approach uses the best practices of topological data analysis (TDA) in NLP: we co…

CoLALinguistic AcceptabilityText ClassificationTopological Data Analysis

Eliciting Best Practices for Collaboration with Computational Notebooks

2022-02-15 · Luigi Quaranta, Fabio Calefato, Filippo Lanubile

Despite the widespread adoption of computational notebooks, little is known about best practices for their usage in collaborative contexts. In this paper, we fill this gap by eliciting a catalog of best practices for col…

Latent Introspection: Models Can Detect Prior Concept Injections

2026-02-23 · Theia Pearson-Vogel, Martin Vanek, Raymond Douglas, Jan Kulveit arxiv

We uncover a latent capacity for introspection in a Qwen 32B model, demonstrating that the model can detect when concepts have been injected into its earlier context and identify which concept was injected. While the mod…

Ethical by Design: Ethics Best Practices for Natural Language Processing

2017-04-01 · WS 2017 4 · Jochen L. Leidner, Vassilis Plachouras

Natural language processing (NLP) systems analyze and/or generate human language, typically on users{'} behalf. One natural and necessary question that needs to be addressed in this context, both in research projects and…

Ethics