paper-with-me

홈 › Papers

Categorize Early, Integrate Late: Divergent Processing Strategies in Automatic Speech Recognition

2026-01-11 · Nathan Roll, Pranav Bhalerao, Martijn Bartelds, Arjun Pawar, Yuka Tatsumi, Tolulope Ogunremi, Chen Shani, Calbert Graham, Meghan Sumner, Dan Jurafsky arxiv

In speech language modeling, two architectures dominate the frontier: the Transformer and the Conformer. However, it remains unknown whether their comparable performance stems from convergent processing strategies or distinct architectural inductive biases. We introduce Architectural Fingerprinting, a probing framework that isolates the effect of architecture on representation, and apply it to a controlled suite of 24 pre-trained encoders (39M-3.3B parameters). Our analysis reveals divergent hierarchies: Conformers implement a "Categorize Early" strategy, resolving phoneme categories 29% earlier in depth and speaker gender by 16% depth. In contrast, Transformers "Integrate Late," deferring phoneme, accent, and duration encoding to deep layers (49-57%). These fingerprints suggest design heuristics: Conformers' front-loaded categorization may benefit low-latency streaming, while Transformers' deep integration may favor tasks requiring rich context and cross-utterance normalization.

📄 PDF Abstract BibTeX arXiv:2601.06972

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

From Shots to Stories: LLM-Assisted Video Editing with Unified Language Representations

2025-05-18 · Yuzhi Li, Haojun Xu, Fang Tian

Large Language Models (LLMs) and Vision-Language Models (VLMs) have demonstrated remarkable reasoning and generalization capabilities in video understanding; however, their application in video editing remains largely un…

Video EditingVideo Understanding

Inclusive Practices for Child-Centered AI Design and Testing

2024-04-09 · Emani Dotch, Vitica Arnold

We explore ideas and inclusive practices for designing and testing child-centered artificially intelligent technologies for neurodivergent children. AI is promising for supporting social communication, self-regulation, a…

Characterizing and Taming Model Instability Across Edge Devices

2020-10-18 · Eyal Cidon, Evgenya Pergament, Zain Asgar, Asaf Cidon 외

The same machine learning model running on different edge devices may produce highly-divergent outputs on a nearly-identical input. Possible reasons for the divergence include differences in the device sensors, the devic…

Convergent World Representations and Divergent Tasks

2026-01-31 · Core Francisco Park arxiv

While neural representations are central to modern deep learning, the conditions governing their geometry and their roles in downstream adaptability remain poorly understood. We develop a framework clearly separating the…

The most controversial topics in Wikipedia: A multilingual and geographical analysis

2013-05-23 · Taha Yasseri, Anselm Spoerri, Mark Graham, János Kertész

We present, visualize and analyse the similarities and differences between the controversial topics related to "edit wars" identified in 10 different language versions of Wikipedia. After a brief review of the related wo…

Articles