paper-with-me

Papers

Worldly Wise (WoW) - Cross-Lingual Knowledge Fusion for Fact-based Visual Spoken-Question Answering

2021-06-01 · NAACL 2021 4 · Kiran Ramnath, Leda Sari, Mark Hasegawa-Johnson, Chang Yoo

Although Question-Answering has long been of research interest, its accessibility to users through a speech interface and its support to multiple languages have not been addressed in prior studies. Towards these ends, we present a new task and a synthetically-generated dataset to do Fact-based Visual Spoken-Question Answering (FVSQA). FVSQA is based on the FVQA dataset, which requires a system to retrieve an entity from Knowledge Graphs (KGs) to answer a question about an image. In FVSQA, the question is spoken rather than typed. Three sub-tasks are proposed: (1) speech-to-text based, (2) end-to-end, without speech-to-text as an intermediate component, and (3) cross-lingual, in which the question is spoken in a language different from that in which the KG is recorded. The end-to-end and cross-lingual tasks are the first to require world knowledge from a multi-relational KG as a differentiable layer in an end-to-end spoken language understanding task, hence the proposed reference implementation is called Worldly-Wise (WoW).WoW is shown to perform end-to-end cross-lingual FVSQA at same levels of accuracy across 3 languages - English, Hindi, and Turkish.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge GraphsQuestion AnsweringSpeech-to-TextSpoken Language UnderstandingWorld Knowledge

Similar Papers 제목 키워드 기반

From task structures to world models: What do LLMs know?

2023-10-06 · Ilker Yildirim, L. A. Paul

In what sense does a large language model have knowledge? The answer to this question extends beyond the capabilities of a particular AI system, and challenges our assumptions about the nature of knowledge and intelligen…

Language ModelingLanguage ModellingLarge Language Model

From Multilingual Complexity to Emotional Clarity: Leveraging Commonsense to Unveil Emotions in Code-Mixed Dialogues

2023-10-19 · Shivani Kumar, Ramaneswaran S, Md Shad Akhtar, Tanmoy Chakraborty

Understanding emotions during conversation is a fundamental aspect of human communication, driving NLP research for Emotion Recognition in Conversation (ERC). While considerable research has focused on discerning emotion…

Dialogue UnderstandingEmotional IntelligenceEmotion RecognitionEmotion Recognition in Conversation+1

Conversational Negation using Worldly Context in Compositional Distributional Semantics

2021-05-12 · ACL (SemSpace, IWCS) 2021 6 · Benjamin Rodatz, Razin A. Shaikh, Lia Yeh

We propose a framework to model an operational conversational negation by applying worldly context (prior knowledge) to logical negation in compositional distributional semantics. Given a word, our framework can create i…

Negation

Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval

2025-06-11 · Shubhashis Roy Dipta, Francis Ferraro

Recent approaches have shown impressive proficiency in extracting and leveraging parametric knowledge from Large-Language Models (LLMs) and Vision-Language Models (VLMs). In this work, we consider how we can improve the …

RetrievalText to Video RetrievalVideo Retrieval

Multilingual Knowledge Graph Embeddings for Cross-lingual Knowledge Alignment

2016-11-12 · Muhao Chen, Yingtao Tian, Mohan Yang, Carlo Zaniolo

Many recent works have demonstrated the benefits of knowledge graph embeddings in completing monolingual knowledge graphs. Inasmuch as related knowledge bases are built in several different languages, achieving cross-lin…

Entity AlignmentKnowledge Graph EmbeddingsKnowledge GraphsTranslation