paper-with-me

홈 › Papers

AgentOhana: Design Unified Data and Training Pipeline for Effective Agent Learning

2024-02-23 · JianGuo Zhang, Tian Lan, Rithesh Murthy, Zhiwei Liu, Weiran Yao, Ming Zhu, Juntao Tan, Thai Hoang, Zuxin Liu, Liangwei Yang, Yihao Feng, Shirley Kokane, Tulika Awalgaonkar, Juan Carlos Niebles, Silvio Savarese, Shelby Heinecke, Huan Wang, Caiming Xiong

Autonomous agents powered by large language models (LLMs) have garnered significant research attention. However, fully harnessing the potential of LLMs for agent-based tasks presents inherent challenges due to the heterogeneous nature of diverse data sources featuring multi-turn trajectories. In this paper, we introduce \textbf{AgentOhana} as a comprehensive solution to address these challenges. \textit{AgentOhana} aggregates agent trajectories from distinct environments, spanning a wide array of scenarios. It meticulously standardizes and unifies these trajectories into a consistent format, streamlining the creation of a generic data loader optimized for agent training. Leveraging the data unification, our training pipeline maintains equilibrium across different data sources and preserves independent randomness across devices during dataset partitioning and model training. Additionally, we present \textbf{xLAM-v0.1}, a large action model tailored for AI agents, which demonstrates exceptional performance across various benchmarks. Begin the exploration at \url{https://github.com/SalesforceAIResearch/xLAM}.

📄 PDF Abstract BibTeX arXiv:2402.15506

Code (2)

SalesforceAIResearch/xLAM/tree/main/xLAM 공식 구현 pytorch
SalesforceAIResearch/xLAM pytorch

Similar Papers 제목 키워드 기반

RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models

2025-10-08 · Hongzhi Zang, Mingjie Wei, Si Xu, Yongji Wu 외 arxiv

Recent advances in vision-language-action (VLA) models have motivated the extension of their capabilities to embodied settings, where reinforcement learning (RL) offers a principled way to optimize task success through i…

Reinforcement Learning

Reconstructing the Image Stitching Pipeline: Integrating Fusion and Rectangling into a Unified Inpainting Model

2024-04-23 · Ziqi Xie, Weidong Zhao, Xianhui Liu, Jian Zhao 외

Deep learning-based image stitching pipelines are typically divided into three cascading stages: registration, fusion, and rectangling. Each stage requires its own network training and is tightly coupled to the others, l…

Image Stitching

UniTE: A Survey and Unified Pipeline for Pre-training Spatiotemporal Trajectory Embeddings

2024-07-17 · Yan Lin, Zeyu Zhou, Yicheng Liu, Haochen Lv 외

Spatiotemporal trajectories are sequences of timestamped locations, which enable a variety of analyses that in turn enable important real-world applications. It is common to map trajectories to vectors, called embeddings…

A Unified and Reproducible Experimentation Framework for Speech Understanding

2026-05-29 · Jing Peng, Junhao Du, Chenghao Wang, Hanqi Li 외 arxiv

Speech foundation models and Speech LLMs have advanced speech understanding, yet deployment-oriented model selection is hindered by non-comparable evaluations caused by mismatched post-processing, and by training results…

PosterText: Towards Unified Visual Text Generation and Editing for E-commerce Poster

2026-08-17 · Xiaoan Liu, Lichen Ma, Zipeng Guo, Yu He 외 arxiv

Automated e-commerce poster design requires both high-quality poster generation and flexible editing of existing designs. However, most existing methods either target end-to-end poster generation or follow multi-stage de…

Reinforcement LearningText Generation