paper-with-me

홈 › Papers

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

2025-08-21 · Yirong Sun, Yizhong Geng, Peidong Wei, Yanjun Chen, Jinghan Yang, Rongfei Chen, Wei Zhang, Xiaoyu Shen arxiv

The development of Large Speech-Language Models (LSLMs) has been slowed by fragmented architectures and a lack of transparency, hindering the systematic comparison and reproducibility of research. Unlike in the vision-language domain, the LSLM field suffers from the common practice of releasing model weights without their corresponding training data and configurations. To address these critical gaps, we introduce LLaSO, the first fully open, end-to-end framework for large-scale speech-language modeling. LLaSO provides the community with three essential resources: (1) LLaSO-Align, a 12M-instance speech-text alignment corpus; (2) LLaSO-Instruct, a 13.5M-instance multi-task instruction-tuning dataset; and (3) LLaSO-Eval, a reproducible benchmark for standardized evaluation. To validate our framework, we build and release LLaSO-Base, a 3.8B-parameter reference model trained exclusively on our public data. It achieves a normalized score of 0.72, establishing a strong, reproducible baseline that surpasses comparable models. Our analysis reveals that while broader training coverage enhances performance, significant generalization gaps persist on unseen tasks, particularly in pure audio scenarios. By releasing the complete stack of data, benchmarks, and models, LLaSO establishes a foundational open standard to unify research efforts and accelerate community-driven progress in LSLMs. We release the code, dataset, pretrained models, and results in https://github.com/EIT-NLP/LLaSO.

📄 PDF Abstract BibTeX arXiv:2508.15418

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LibContinual: A Comprehensive Library towards Realistic Continual Learning

2025-12-26 · Wenbin Li, Shangge Liu, Borui Kang, Yiyang Chen 외 arxiv

A fundamental challenge in Continual Learning (CL) is catastrophic forgetting, where adapting to new tasks degrades the performance on previous ones. While the field has evolved with diverse methods, this rapid surge in …

Continual Learning

OmniGenBench: A Modular Platform for Reproducible Genomic Foundation Models Benchmarking

2025-05-20 · Heng Yang, Jack Cole, Yuan Li, Renzhi Chen 외

The code of nature, embedded in DNA and RNA genomes since the origin of life, holds immense potential to impact both humans and ecosystems through genome modeling. Genomic Foundation Models (GFMs) have emerged as a trans…

Benchmarking

Machine Learning for Medicine Must Be Interpretable, Shareable, Reproducible and Accountable by Design

2025-08-22 · Ayyüce Begüm Bektaş, Mithat Gönen arxiv

This paper claims that machine learning models deployed in high stakes domains such as medicine must be interpretable, shareable, reproducible and accountable. We argue that these principles should form the foundational …

Federated Learning

A Large-Scale Dataset of MCP Implementations on GitHub

2026-07-11 · Benny Toeppe, Amine Barrak, Emna Ksontini arxiv

The rapid emergence of the Model Context Protocol (MCP) has introduced a new standard for connecting large language models to external tools and services. Despite its rapid adoption in open-source development, systematic…

L-ReLF: A Framework for Lexical Dataset Creation

2026-03-31 · Anass Sedrati, Mounir Afifi, Reda Benkhadra arxiv

This paper introduces the L-ReLF (Low-Resource Lexical Framework), a novel, reproducible methodology for creating high-quality, structured lexical datasets for underserved languages. The lack of standardized terminology,…

Machine Translation