paper-with-me

Papers

Beyond Transcripts: A Renewed Perspective on Audio Chaptering

2026-02-09 · Fabian Retkowski, Maike Züfle, Thai Binh Nguyen, Jan Niehues, Alexander Waibel arxiv

Audio chaptering, the task of segmenting long-form audio into coherent sections, is increasingly important for navigating podcasts, lectures, and videos. Despite its relevance, research remains limited and text-based, leaving key questions unresolved about leveraging audio information, handling ASR errors, and transcript-free evaluation. We address these gaps through three contributions: (1) a systematic comparison between text-based models with acoustic features, a novel audio-only architecture (AudioSeg) operating on learned audio representations, and multimodal LLMs; (2) empirical analysis of factors affecting performance, including transcript quality, acoustic features, duration, and speaker composition; and (3) formalized evaluation protocols contrasting transcript-dependent text-space protocols with transcript-invariant time-space protocols. Our experiments on YTSeg reveal that AudioSeg substantially outperforms text-based approaches, pauses provide the largest acoustic gains, and MLLMs remain limited by context length and weak instruction following, yet MLLMs are promising on shorter audio.

📄 PDF Abstract BibTeX arXiv:2602.08979

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs

2025-03-31 · CVPR 2025 1 · Lucas Ventura, Antoine Yang, Cordelia Schmid, Gül Varol

We address the task of video chaptering, i.e., partitioning a long video timeline into semantic units and generating corresponding chapter titles. While relatively underexplored, automatic chaptering has the potential to…

Large Language ModelVideo Chaptering

ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries

2025-11-18 · Junfu Pu, Teng Wang, Yixiao Ge, Yuying Ge 외 arxiv

The proliferation of hour-long videos (e.g., lectures, podcasts, documentaries) has intensified demand for efficient content structuring. However, existing approaches are constrained by small-scale training with annotati…

Dense Video CaptioningSemantic SimilarityVideo Chaptering

Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech

2026-05-17 · Kaavya Chaparala, Thomas Thebaud, Jesús Villalba López, Laureano Moro-Velazquez 외 arxiv

There are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed systemic errors and bias into datasets. We test ten annotation workf…

TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics

2025-09-30 · Yi-Cheng Lin, Yu-Hua Chen, Jia-Kai Dong, Yueh-Hsuan Huang 외 arxiv

Large audio-language models are advancing rapidly, yet most evaluations emphasize speech or globally sourced sounds, overlooking culturally distinctive cues. This gap raises a critical question: can current models genera…

Question Generation

From Text Segmentation to Smart Chaptering: A Novel Benchmark for Structuring Video Transcriptions

2024-02-27 · Fabian Retkowski, Alexander Waibel

Text segmentation is a fundamental task in natural language processing, where documents are split into contiguous sections. However, prior research in this area has been constrained by limited datasets, which are either …

Headline GenerationSegmentationText Segmentation