paper-with-me

Papers Instruction Following

“Instruction Following” 태그가 달린 논문 1,603편 · 필터 해제

SenseNova-U1.5: Towards Native Unified Visual Intelligence

2026-09-10 · Haiwen Diao, Jiahao Wang, Chenjing Ding, Hanming Deng 외 hf

We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface throu…

Reinforcement LearningInstruction FollowingImage Editing

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

2026-09-10 · Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang 외 hf

We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video g…

Instruction FollowingVideo Generation

Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning

2026-09-09 · Mehrnaz Mofakhami, Ananya Sahu, Alejandro R. Salamanca, Daniel D'souza 외 arxiv

Reasoning language models have made substantial advances on a variety of complex tasks, yet their capabilities remain overwhelmingly English-centric: models primarily reason in English regardless of the language they are…

Instruction Following

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

2026-09-08 · NeoHorse Team, Guoliang Cao, Guohao Dai, Tianyu Guo 외 hf

Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-nat…

Instruction Following

MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control

2026-09-05 · Ting Huang, Yue Huang, Zeyu Zhang, Shuicheng Yan 외 hf

Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge for vision-language-action (VLA) systems on mobile robots, due to the persistent gap between high-level semanti…

Reinforcement LearningInstruction FollowingMultimodal ReasoningDecision Making

EuroAlpaca: Task-Preserving Localisation of Instruction Data for European Languages

2026-09-04 · Aleix Sant, Jordi Luque, Carlos Escolano arxiv

Machine translation (MT) offers a scalable way to extend English instruction-tuning data to multiple languages, but it can distort task-critical constraints and required outputs, creating corrupted training examples and …

Instruction FollowingMachine Translation

Learning 3D Editing without Paired Supervision via Generative Prior Distillation

2026-09-04 · Hao Wen, Weibin Yun, Hongxing Fan, Haotian Lu 외 arxiv

Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottleneck: the severe scarcity of high-quality paired training data. Existing approaches attempt to bypass this by …

Instruction FollowingImage Editing

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

2026-09-04 · Wenhui Chen, Shiwen Cheng, Hao Dong, Chenda Duan 외 hf

Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation across many model invocations. Because cost and latency accumulate over the full episode, their practic…

Instruction Following

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

2026-09-02 · Zhiwei Zhang, Zechen Sun, Fei Zhao, Kang Peng 외 hf

On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, withou…

Instruction Following

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix

2026-09-01 · Olga Tsymboi, Dmitrii Stoianov, Ramil Latypov, Danil Taranets 외 hf

Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate …

Instruction Following

Reading the News: Adapting Large Language Models to Swedish Journalism Through Continued Pre-Training

2026-08-31 · Lukas Borggren, Jenny Kunz, Marco Kuhlmann arxiv

Large language models are increasingly capable in general, but their utility can remain modest in niche or understudied areas. One approach to address this limitation is to specialise existing models through additional t…

parameter-efficient fine-tuningInstruction Following

VIBE: Video Instruction-aligned Background music gEneration

2026-08-31 · Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj, Gouthaman KV 외 arxiv

Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representational bottleneck of static cross-modal …

Instruction FollowingMusic Generation

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

2026-08-31 · Shaoan Wang, Aocheng Luo, Fei Huang, Jingyi Xu 외 hf

Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial pri…

Zero-shot GeneralizationReinforcement LearningInstruction FollowingSpatial Reasoning

DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging

2026-08-28 · Aaryan Ajay Sharma, Sai Nishanth Padala, Seganrasan Subramanian arxiv

Model merging combines multiple task-specific fine-tuned LLMs into a single multi-task model without additional training. However, merged models are known to suffer from representation bias: systematic drift between the …

Mathematical ReasoningInstruction FollowingCode Generation

Preserving General Capabilities during Domain Specialization with Uncertainty-Calibrated MOPD

2026-08-27 · Ziyuan Liu, Jiao Ou, Jian Liang, Ruiming Tang 외 arxiv

Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities such as reasoning, coding, instruction following, and creative writing. We study this domai…

Instruction Following

VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following

2026-08-26 · Min Zeng, Guanxin Tan, Libin Cen, Yawei Wen 외 arxiv

Multimodal instruction-following models require training data that is accurate, diverse, verifiable, and challenging. Existing synthesis pipelines typically follow a one-pass generate-and-filter paradigm, discarding feed…

Reinforcement LearningInstruction Following

SciMIF: Understanding Multimodal Instruction Following in Scientific Domains

2026-08-26 · Ye Shen, Yuting Zheng, Dun Pei, Zijian Chen 외 arxiv

Understanding instruction-following capabilities in scientific domains is essential for effectively leveraging Multimodal Large Language Models (MLLMs) to advance the development of scientific fields. In this work, we in…

Instruction Following

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization

2026-08-26 · Zaibin Zhang, Junlan Xiao, Zhongbo Zhang, Yifan Wang 외 arxiv

Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) models integrate perception, language, and control, but most represent language as a single global instr…

Instruction Following

V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning

2026-08-26 · Shulin Tian, Minglun Li, Yuhao Dong, Hao Ding 외 arxiv

Vision-language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single unsupported object, chart value, or intermediate inference can undermine an otherwise plausible response…

Reinforcement LearningInstruction Following

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

2026-08-26 · Hongbo Liu, Peixian Chen, Sihan Liu, Peiyuan Zhang 외 arxiv

Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires …

Instruction Following
1–20 / 1,603 다음 →