paper-with-me

Papers

GEMS: Agent-Native Multimodal Generation with Memory and Skills

2026-03-30 · Zefeng He, Siyuan Huang, Xiaoye Qu, Yafu Li, Tong Zhu, Yu Cheng, Yang Yang arxiv

Recent multimodal generation models have achieved remarkable progress on general-purpose generation tasks, yet continue to struggle with complex instructions and specialized downstream tasks. Inspired by the success of advanced agent frameworks such as Claude Code, we propose \textbf{GEMS} (Agent-Native Multimodal \textbf{GE}neration with \textbf{M}emory and \textbf{S}kills), a framework that pushes beyond the inherent limitations of foundational models on both general and downstream tasks. GEMS is built upon three core components. Agent Loop introduces a structured multi-agent framework that iteratively improves generation quality through closed-loop optimization. Agent Memory provides a persistent, trajectory-level memory that hierarchically stores both factual states and compressed experiential summaries, enabling a global view of the optimization process while reducing redundancy. Agent Skill offers an extensible collection of domain-specific expertise with on-demand loading, allowing the system to effectively handle diverse downstream applications. Across five mainstream tasks and four downstream tasks, evaluated on multiple generative backends, GEMS consistently achieves significant performance gains. Most notably, it enables the lightweight 6B model Z-Image-Turbo to surpass the state-of-the-art Nano Banana 2 on GenEval2, demonstrating the effectiveness of agent harness in extending model capabilities beyond their original limits.

📄 PDF Abstract BibTeX arXiv:2603.28088

Code (0)

등록된 구현이 없습니다.

Tasks

multimodal generation

Similar Papers 제목 키워드 기반

See and Remember: A Multimodal Agent for Web Traversal

2026-03-03 · Xinjun Wang, Shengyao Wang, Aimin Zhou, Hao Hao arxiv

Autonomous web navigation requires agents to perceive complex visual environments and maintain long-term context, yet current Large Language Model (LLM) based agents often struggle with spatial disorientation and navigat…

Visual Grounding

Generative Evolutionary Meta-Solver (GEMS): Scalable Surrogate-Free Multi-Agent Reinforcement Learning

2025-09-27 · Alakh Sharma, Gaurish Trivedi, Kartikey Singh Bhandari, Yash Sinha 외 arxiv

Scalable multi-agent reinforcement learning (MARL) remains a central challenge for AI. Existing population-based methods, like Policy-Space Response Oracles, PSRO, require storing explicit policy populations and construc…

Multi-agent Reinforcement Learning

t-gems: text-guided exit modules for decreasing clip image encoder

2026-05-17 · Alberto Presta, Grzegorz Stefanski, Michal Byra, Krzysztof Arendt arxiv

Multimodal deep neural networks enhance deep comprehension by integrating diverse data modalities. Data from different modalities are typically projected into a shared latent space for similarity computation, but this pr…

Gems: Group Emotion Profiling Through Multimodal Situational Understanding

2025-07-30 · Anubhav Kataria, Surbhi Madan, Shreya Ghosh, Tom Gedeon 외 arxiv

Understanding individual, group and event level emotions along with contextual information is crucial for analyzing a multi-person social situation. To achieve this, we frame emotion comprehension as the task of predicti…

GeMS: Efficient Gaussian Splatting for Extreme Motion Blur

2025-08-20 · Gopi Raju Matta, Trisha Reddypalli, Vemunuri Divya Madhuri, Kaushik Mitra arxiv

We introduce GeMS, a framework for 3D Gaussian Splatting (3DGS) designed to handle severely motion-blurred images. State-of-the-art deblurring methods for extreme blur, such as ExBluRF, as well as Gaussian Splatting-base…

Point Cloud GenerationCamera Pose EstimationPoint Clouds