paper-with-me

홈 › Papers

Linguistic unit discovery from multi-modal inputs in unwritten languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop

2018-02-14 · Odette Scharenborg, Laurent Besacier, Alan Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stueker, Pierre Godard, Markus Mueller, Lucas Ondel, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang, Emmanuel Dupoux

We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translated text in a well-resourced language to help unsupervised discovery from raw speech.

📄 PDF Abstract BibTeX arXiv:1802.05092

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multimodal Generalized Category Discovery

2024-09-18 · Yuchang Su, Renping Zhou, Siyu Huang, Xingjian Li 외

Generalized Category Discovery (GCD) aims to classify inputs into both known and novel categories, a task crucial for open-world scientific discoveries. However, current GCD methods are limited to unimodal data, overlook…

Contrastive Learning

Unsupervised Multimodal Word Discovery based on Double Articulation Analysis with Co-occurrence cues

2022-01-18 · Akira Taniguchi, Hiroaki Murakami, Ryo Ozaki, Tadahiro Taniguchi

Human infants acquire their verbal lexicon with minimal prior knowledge of language based on the statistical properties of phonological distributions and the co-occurrence of other sensory stimuli. This study proposes a …

XAI for In-hospital Mortality Prediction via Multimodal ICU Data

2023-12-29 · Xingqiao Li, Jindong Gu, Zhiyong Wang, Yancheng Yuan 외

Predicting in-hospital mortality for intensive care unit (ICU) patients is key to final clinical outcomes. AI has shown advantaged accuracy but suffers from the lack of explainability. To address this issue, this paper p…

Decision MakingMortality Prediction

Automating construction safety inspections using a multi-modal vision-language RAG framework

2025-10-05 · Chenxin Wang, Elyas Asadi Shamsabadi, Zhaohui Chen, Luming Shen 외 arxiv

Conventional construction safety inspection methods are often inefficient as they require navigating through large volume of information. Recent advances in large vision-language models (LVLMs) provide opportunities to a…

Information Retrieval

VCD: Visual Causality Discovery for Cross-Modal Question Reasoning

2023-04-17 · Yang Liu, Ying Tan, Jingzhou Luo, Weixing Chen

Existing visual question reasoning methods usually fail to explicitly discover the inherent causal mechanism and ignore jointly modeling cross-modal event temporality and causality. In this paper, we propose a visual que…