paper-with-me

Papers

Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective

2025-05-26 · Junnan Liu, Hongwei Liu, Linchen Xiao, Shudong Liu, Taolin Zhang, Zihan Ma, Songyang Zhang, Kai Chen

We propose a novel framework for comprehending the reasoning capabilities of large language models (LLMs) through the perspective of meta-learning. By conceptualizing reasoning trajectories as pseudo-gradient descent updates to the LLM's parameters, we identify parallels between LLM reasoning and various meta-learning paradigms. We formalize the training process for reasoning tasks as a meta-learning setup, with each question treated as an individual task, and reasoning trajectories serving as the inner loop optimization for adapting model parameters. Once trained on a diverse set of questions, the LLM develops fundamental reasoning capabilities that can generalize to previously unseen questions. Extensive empirical evaluations substantiate the strong connection between LLM reasoning and meta-learning, exploring several issues of significant interest from a meta-learning standpoint. Our work not only enhances the understanding of LLM reasoning but also provides practical insights for improving these models through established meta-learning techniques.

📄 PDF Abstract BibTeX arXiv:2505.19815

Code (1)

open-compass/raml 공식 구현 pytorch

Tasks

Meta-Learning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Deciphering Scientific Reasoning Steps from Outcome Data for Molecule Optimization

2026-03-13 · Zequn Liu, Kehan Wu, Shufang Xie, Zekun Guo 외 arxiv

Emerging reasoning models hold promise for automating scientific discovery. However, their training is hindered by a critical supervision gap: experimental outcomes are abundant, whereas intermediate reasoning steps are …

Drug Discovery

Unmanned Aerial Vehicle-Aided Communications: Joint Transmit Power and Trajectory Optimization

2018-01-14

This letter investigates the transmit power and trajectory optimization problem for unmanned aerial vehicle (UAV)-aided networks. Different from majority of the existing studies with fixed communication infrastructure, a…

Sense-Store-Send: Trajectory Optimization for a Buffer-aided Internet of UAVs

2020-09-15 · Yujie Jin, Hongliang Zhang, Shuhang Zhang, Zhu Han 외

In this letter, we study a buffer-aided Internet of unmanned aerial vehicles (UAVs) in which a UAV performs data sensing, stores the data, and sends it to the base station (BS) in cellular networks. To minimize the overa…

Deciphering Human Mobility: Inferring Semantics of Trajectories with Large Language Models

2024-05-30 · Yuxiao Luo, Zhongcai Cao, Xin Jin, Kang Liu 외

Understanding human mobility patterns is essential for various applications, from urban planning to public safety. The individual trajectory such as mobile phone location data, while rich in spatio-temporal information, …

Spotlight on Token Perception for Multimodal Reinforcement Learning

2025-10-10 · Siyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo 외 arxiv

While Reinforcement Learning with Verifiable Rewards (RLVR) has advanced the reasoning capabilities of Large Vision-Language Models (LVLMs), most existing methods in multimodal reasoning neglect the critical role of visu…

Reinforcement LearningMultimodal Reasoning