paper-with-me

홈 › Papers

FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models

2024-07-16 · Pengxiang Li, Zhi Gao, Bofei Zhang, Tao Yuan, Yuwei Wu, Mehrtash Harandi, Yunde Jia, Song-Chun Zhu, Qing Li

Vision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn conversations that are derived from 27 source datasets, empowering VLMs to spontaneously refine their responses based on user feedback across diverse tasks. To scale up the data collection, FIRE is collected in two components: FIRE-100K and FIRE-1M, where FIRE-100K is generated by GPT-4V, and FIRE-1M is freely generated via models trained on FIRE-100K. Then, we build FIRE-Bench, a benchmark to comprehensively evaluate the feedback-refining capability of VLMs, which contains 11K feedback-refinement conversations as the test data, two evaluation settings, and a model to provide feedback for VLMs. We develop the FIRE-LLaVA model by fine-tuning LLaVA on FIRE-100K and FIRE-1M, which shows remarkable feedback-refining capability on FIRE-Bench and outperforms untrained VLMs by 50%, making more efficient user-agent interactions and underscoring the significance of the FIRE dataset.

📄 PDF Abstract BibTeX arXiv:2407.11522

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle

2025-11-21 · Mario Markov, Stefan Maria Ailuro, Luc Van Gool, Konrad Schindler 외 arxiv

Predicting wildfire risk is a reasoning-intensive spatial problem that requires the integration of visual, climatic, and geographic factors to infer continuous risk maps. Existing methods lack the causal reasoning and mu…

Reinforcement Learning

FiRe: Fine-grained Multimodal Reasoning for Enhanced Image Generation

2026-04-15 · Yongjin Kim, Yoonjin Oh, Yerin Kim, Hyomin Kim 외 arxiv

With the rapid progress of Multimodal Large Language Models (MLLMs), unified MLLMs that jointly perform image understanding and generation have advanced significantly. However, despite the inherent reasoning capabilities…

Text-to-Image GenerationReinforcement LearningMultimodal Reasoning

Exploring Reasoning Reward Model for Agents

2026-01-29 · Kaixuan Fan, Kaituo Feng, Manyuan Zhang, Tianshuo Peng 외 arxiv

Agentic Reinforcement Learning (Agentic RL) has achieved notable success in enabling agents to perform complex reasoning and tool use. However, most methods still relies on sparse outcome-based reward for training. Such …

Reinforcement Learning

FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training

2025-02-02 · Liangyu Xu, Xuemiao Zhang, Feiyu Duan, Sirui Wang 외

Selecting high-quality data can significantly improve the pre-training efficiency of large language models (LLMs). Existing methods often rely on heuristic techniques and single quality signals, limiting their ability to…

BioDisco: Multi-agent hypothesis generation with dual-mode evidence, iterative feedback and temporal evaluation

2025-08-02 · Yujing Ke, Kevin George, Kathan Pandya, David Blumenthal 외 arxiv

Identifying novel hypotheses is essential to scientific research, yet this process risks being overwhelmed by the sheer volume and complexity of available information. Existing automated methods often struggle to generat…

Knowledge Graphs