Training a Multilingual Sportscaster: Using Perceptual Context to Learn Language
We present a novel framework for learning to interpret and generate language using only perceptual context as supervision. We demonstrate its capabilities by developing a system that learns to sportscast simulated robot soccer games in both English and Korean without any language-specific prior knowledge. Training employs only ambiguous supervision consisting of a stream of descriptive textual comments and a sequence of events extracted from the simulation trace. The system simultaneously establishes correspondences between individual comments and the events that they describe while building a translation model that supports both parsing and generation. We also present a novel algorithm for learning which events are worth describing. Human evaluations of the generated commentaries indicate they are of reasonable quality and in some cases even on par with those produced by humans for our limited domain.
Code (0)
등록된 구현이 없습니다.
Tasks
DescriptiveTranslationSimilar Papers 제목 키워드 기반
Learning to Make Inferences in a Semantic Parsing Task
We introduce a new approach to training a semantic parser that uses textual entailment judgements as supervision. These judgements are based on high-level inferences about whether the meaning of one sentence follows from…
Machine TranslationNatural Language InferenceQuestion AnsweringRTE+2Do Speech Emphasis Models Generalize across Languages and Emotions?
Prosodic emphasis varies across languages, emotions, and speaking styles, yet existing emphasis detection models are largely trained and evaluated on monolingual neutral read speech. We introduce MMEE (Multilingual Multi…
A Multimodal Recaptioning Framework to Account for Perceptual Diversity in Multilingual Vision-Language Modeling
There are many ways to describe, name, and group objects when captioning an image. Differences are evident when speakers come from diverse cultures due to the unique experiences that shape perception. Machine translation…
DiversityImage RetrievalLanguage ModelingLanguage Modelling+2Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
Dysarthria, a motor speech disorder, severely impacts voice quality, pronunciation, and prosody, leading to diminished speech intelligibility and reduced quality of life. Accurate assessment is crucial for effective trea…
Decision MakingQuantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval
There is a scarcity of multilingual vision-language models that properly account for the perceptual differences that are reflected in image captions across languages and cultures. In this work, through a multimodal, mult…
Image CaptioningRetrieval