paper-with-me

홈 › Papers

Staircase Streaming for Low-Latency Multi-Agent Inference

2025-10-06 · Junlin Wang, Jue Wang, Zhen, Xu, Ben Athiwaratkun, Bhuwan Dhingra, Ce Zhang, James Zou arxiv

Recent advances in large language models (LLMs) opened up new directions for leveraging the collective expertise of multiple LLMs. These methods, such as Mixture-of-Agents, typically employ additional inference steps to generate intermediate outputs, which are then used to produce the final response. While multi-agent inference can enhance response quality, it can significantly increase the time to first token (TTFT), posing a challenge for latency-sensitive applications and hurting user experience. To address this issue, we propose staircase streaming for low-latency multi-agent inference. Instead of waiting for the complete intermediate outputs from previous steps, we begin generating the final response as soon as we receive partial outputs from these steps. Experimental results demonstrate that staircase streaming reduces TTFT by up to 93% while maintaining response quality.

📄 PDF Abstract BibTeX arXiv:2510.05059

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SpeakStream: Streaming Text-to-Speech with Interleaved Data

2025-05-25 · Richard He Bai, Zijin Gu, Tatiana Likhomanenko, Navdeep Jaitly

The latency bottleneck of traditional text-to-speech (TTS) systems fundamentally hinders the potential of streaming large language models (LLMs) in conversational AI. These TTS systems, typically trained and inferenced o…

Decodertext-to-speechText to Speech

SHARP: Short-Window Streaming for Accurate and Robust Prediction in Motion Forecasting

2026-03-30 · Alexander Prutsch, Christian Fruhwirth-Reisinger, David Schinagl, Horst Possegger arxiv

In dynamic traffic environments, motion forecasting models must be able to accurately estimate future trajectories continuously. Streaming-based methods are a promising solution, but despite recent advances, their perfor…

Motion Forecasting

Adapting Offline Speech Translation Models for Streaming with Future-Aware Distillation and Inference

2023-03-14 · Biao Fu, Minpeng Liao, Kai Fan, Zhongqiang Huang 외

A popular approach to streaming speech translation is to employ a single offline model with a wait-k policy to support different latency requirements, which is simpler than training multiple online models with different …

FADTranslation

Dynamic latency speech recognition with asynchronous revision

2020-11-03 · Mingkun Huang, Meng Cai, Jun Zhang, Yang Zhang 외

In this work we propose an inference technique, asynchronous revision, to unify streaming and non-streaming speech recognition models. Specifically, we achieve dynamic latency with only one model by using arbitrary right…

Decoderspeech-recognitionSpeech Recognition

A low latency attention module for streaming self-supervised speech representation learning

2023-02-27 · Jianbo Ma, Siqi Pan, Deepak Chandran, Andrea Fanelli 외

The transformer is a fundamental building block in deep learning, and the attention mechanism is the transformer's core component. Self-supervised speech representation learning (SSRL) represents a popular use-case for t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionRepresentation Learning+4