paper-with-me

Papers

Better Intermediates Improve CTC Inference

2022-04-01 · Tatsuya Komatsu, Yusuke Fujita, Jaesong Lee, Lukas Lee, Shinji Watanabe, Yusuke Kida

This paper proposes a method for improved CTC inference with searched intermediates and multi-pass conditioning. The paper first formulates self-conditioned CTC as a probabilistic model with an intermediate prediction as a latent representation and provides a tractable conditioning framework. We then propose two new conditioning methods based on the new formulation: (1) Searched intermediate conditioning that refines intermediate predictions with beam-search, (2) Multi-pass conditioning that uses predictions of previous inference for conditioning the next inference. These new approaches enable better conditioning than the original self-conditioned CTC during inference and improve the final performance. Experiments with the LibriSpeech dataset show relative 3%/12% performance improvement at the maximum in test clean/other sets compared to the original self-conditioned CTC.

📄 PDF Abstract BibTeX arXiv:2204.00176

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Leap: molecular synthesisability scoring with intermediates

2024-03-14 · Antonia Calvi, Théophile Gaudin, Dominik Miketa, Dominique Sydow 외

Assessing whether a molecule can be synthesised is a primary task in drug discovery. It enables computational chemists to filter for viable compounds or bias molecular generative models. The notion of synthesisability is…

Drug Discovery

Gen-VCoT: Generative Visual Chain-of-Thought Reasoning via Diffusion-Based RGB Intermediate Representations

2026-06-15 · Zhiqiang Zhou, Junliang Dai, Xu ling arxiv

Multimodal large language models (MLLMs) excel at visual reasoning but rely on text-based chain-of-thought (CoT), lacking interpretable visual intermediates. Existing methods use opaque tokens or external tools, missing …

Multimodal ReasoningVisual GroundingVisual Reasoning

Fast-MD: Fast Multi-Decoder End-to-End Speech Translation with Non-Autoregressive Hidden Intermediates

2021-09-27 · Hirofumi Inaguma, Siddharth Dalmia, Brian Yan, Shinji Watanabe

The multi-decoder (MD) end-to-end speech translation model has demonstrated high translation quality by searching for better intermediate automatic speech recognition (ASR) decoder states as hidden intermediates (HI). It…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)CPUDecoder+6

A Two-Step Biopolymer Nucleation Model Shows a Nonequilibrium Critical Point

2019-12-17

Biopolymer self-assembly pathways are central to biological activity, but are complicated by the ability of the monomeric subunits of biopolymers to adopt different conformational states. As a result, biopolymer nucleati…

Vocal Bursts Valence Prediction

Searchable Hidden Intermediates for End-to-End Models of Decomposable Sequence Tasks

2021-05-02 · NAACL 2021 4 · Siddharth Dalmia, Brian Yan, Vikas Raunak, Florian Metze 외

End-to-end approaches for sequence tasks are becoming increasingly popular. Yet for complex sequence tasks, like speech translation, systems that cascade several models trained on sub-tasks have shown to be superior, sug…

Decoderspeech-recognitionSpeech RecognitionTranslation