Improving Large Vision and Language Models by Learning from a Panel of Peers
Traditional alignment methods for Large Vision and Language Models (LVLMs) primarily rely on human-curated preference data. Human-generated preference data is costly; machine-generated preference data is limited in quality; and self-supervised preference data often introduces hallucinations. To overcome these limitations, we propose a novel Panel-of-Peers learning framework inspired by collaborative learning among humans. This approach leverages a panel of LVLMs, each evaluating and learning from their collective outputs through an iterative self-improvement process. By simulating a peer review system, our models generate, assess, and refine outputs in response to a curated set of prompts, mimicking a classroom learning environment. We demonstrate that this methodology enhances model performance without requiring extensive human-labeled datasets. Our experiments show significant improvement across multiple benchmarks, demonstrating the potential of peer evaluations as a scalable alternative to self-supervised alignment. Notably, we show that Panel-of-Peers increases the average score on fifteen benchmarks from 48% to 57%
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Is the panel fair? Evaluating panel compositions through network analysis. The case of research assessments in Italy
Research evaluation is usually governed by panels of peers. Procedural fairness refers to the principles that ensures decisions are made through a fair and transparent process. It requires that the composition of panels …
FairnessTesting for Peer Effects without Specifying the Network Structure
This paper proposes an Anderson-Rubin (AR) test for the presence of peer effects in panel data without the need to specify the network structure. The unrestricted model of our test is a linear panel data model of social …
validFrom Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
There is a growing interest in developing strong biomedical vision-language models. A popular approach to achieve robust representations is to use web-scale scientific data. However, current biomedical vision-language pr…
Detecting Learning by Exporting and from Exporters
Existing literature at the nexus of firm productivity and export behavior mostly focuses on "learning by exporting," whereby firms can improve their performance by engaging in exports. Whereas, the secondary channel of l…
Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning
Many multimodal tasks depend on how visual elements are ordered and composed, not only on recognizing them in isolation. Internet memes are a compact case of this problem: their punchline often depends on a constrained r…
Explanation GenerationMultimodal Reasoning