paper-with-me

홈 › Papers

ANCHOR: Autoregressive Non-intrusive Chunk-Ordered Refinement for Joint Multi-Resolution Speech Quality Modeling

2026-06-08 · Zhuoyan Tao, Jiatong Shi, Hye-jin Shim, Shinji Watanabe arxiv

While speech quality is typically assessed on complete utterances, streaming and generative systems require incremental estimation from partial audio. Existing predictors assume full context, degrading on prefix-constrained inputs. Extending ARECHO, we propose ANCHOR, reformulating incremental assessment as a multi-resolution autoregressive task. It models chunk- and utterance-level quality within a single decoder using dual-resolution tokens and a resolution-aware hierarchy for coarse-to-fine refinement. Experiments show substantial robustness under partial input, including a 48% PLCMOS error reduction on 2-second prefixes. Convergence analysis reveals a 4-6 s effective perceptual context horizon. A stress test further isolates structured extrapolation biases under localized corruption. Results demonstrate that hierarchical supervision improves incremental prediction and elucidates how perceptual quality accumulates over time.

📄 PDF Abstract BibTeX arXiv:2606.10233

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OAT: Ordered Action Tokenization

2026-02-04 · Chaoqi Liu, Xiaoshen Han, Jiawei Gao, Yue Zhao 외 arxiv

Autoregressive policies offer a compelling foundation for scalable robot learning by enabling discrete abstraction, token-level reasoning, and flexible inference. However, applying autoregressive modeling to continuous r…

Monotonic Simultaneous Translation with Chunk-wise Reordering and Refinement

2021-10-18 · WMT (EMNLP) 2021 11 · Hyojung Han, Seokchan Ahn, Yoonjung Choi, Insoo Chung 외

Recent work in simultaneous machine translation is often trained with conventional full sentence translation corpora, leading to either excessive latency or necessity to anticipate as-yet-unarrived words, when dealing wi…

Machine TranslationSentenceTranslationWord Alignment

AdaState: Self-Evolving Anchors for Streaming Video Generation

2026-05-28 · Yusuf Dalva, Pinar Yanardag arxiv

Autoregressive video diffusion models generate streaming video by producing frames sequentially, conditioning each chunk on previously generated content. These models are structurally anchored to the first frame: its key…

Video Generation

WorldKV: Efficient World Memory with World Retrieval and Compression

2026-05-21 · Jung Yi, Minjae Kim, Paul Hyunbin Cho, Wooseok Jang 외 arxiv

Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisiting a previously seen viewpoint yields consistent content, remains a…

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

2026-08-04 · Yicheng Xiao, Wenxun Dai, Xinran Qin, Lin Song 외 hf

Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autore…