paper-with-me

Papers

A Multimodal Simultaneous Interpretation Prototype: Who Said What

2022-09-01 · AMTA 2022 9 · Xiaolin Wang, Masao Utiyama, Eiichiro Sumita

“Who said what” is essential for users to understand video streams that have more than one speaker, but conventional simultaneous interpretation systems merely present “what was said” in the form of subtitles. Because the translations unavoidably have delays and errors, users often find it difficult to trace the subtitles back to speakers. To address this problem, we propose a multimodal SI system that presents users “who said what”. Our system takes audio-visual approaches to recognize the speaker of each sentence, and then annotates its translation with the textual tag and face icon of the speaker, so that users can quickly understand the scenario. Furthermore, our system is capable of interpreting video streams in real-time on a single desktop equipped with two Quadro RTX 4000 GPUs owing to an efficient sentence-based architecture.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceTAGTranslation

Similar Papers 제목 키워드 기반

M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models

2025-10-22 · Yejin Kwon, Taewoo Kang, Hyunsoo Yoon, Changouk Kim arxiv

We present M3-SLU, a new multimodal large language model (MLLM) benchmark for evaluating multi-speaker, multi-turn spoken language understanding. While recent models show strong performance in speech and text comprehensi…

Spoken Language UnderstandingQuestion Answering

What's Left Unsaid? Detecting and Correcting Misleading Omissions in Multimodal News Previews

2026-01-09 · Fanxiao Li, Jiaying Wu, Tingchao Fu, Dayang Li 외 arxiv

Even when factually correct, social-media news previews (image-headline pairs) can induce interpretation drift: by selectively omitting crucial context, they lead readers to form judgments that diverge from what the full…

A Prototype Automatic Simultaneous Interpretation System

2016-12-01 · COLING 2016 12 · Xiaolin Wang, Andrew Finch, Masao Utiyama, Eiichiro Sumita

Simultaneous interpretation allows people to communicate spontaneously across language boundaries, but such services are prohibitively expensive for the general public. This paper presents a fully automatic simultaneous …

Decoupling Deep Learning for Interpretable Image Recognition

2022-10-15 · Yitao Peng, Yihang Liu, Longzhen Yang, Lianghua He

The interpretability of neural networks has recently received extensive attention. Previous prototype-based explainable networks involved prototype activation in both reasoning and interpretation processes, requiring spe…

Decision MakingDecoderDeep Learning

Interpreting systems as solving POMDPs: a step towards a formal understanding of agency

2022-09-04 · Martin Biehl, Nathaniel Virgo

Under what circumstances can a system be said to have beliefs and goals, and how do such agency-related features relate to its physical state? Recent work has proposed a notion of interpretation map, a function that maps…