paper-with-me

Papers

Multimodal Word Distributions

2017-04-27 · ACL 2017 7 · Ben Athiwaratkun, Andrew Gordon Wilson

Word embeddings provide point representations of words containing useful semantic information. We introduce multimodal word distributions formed from Gaussian mixtures, for multiple word meanings, entailment, and rich uncertainty information. To learn these distributions, we propose an energy-based max-margin objective. We show that the resulting approach captures uniquely expressive semantic information, and outperforms alternatives, such as word2vec skip-grams, and Gaussian embeddings, on benchmark datasets such as word similarity and entailment.

📄 PDF Abstract BibTeX arXiv:1704.08424

Code (2)

benathi/word2gm 공식 구현 tf
benathi/multisense-prob-fasttext

Tasks

Word EmbeddingsWord Similarity

Similar Papers 제목 키워드 기반

Where did the ambiguity go? Examining how multimodal models interpret polysemous words

2026-08-01 · Jasin Cekinmez, Addison J. Wu, Raja Marjieh, Thomas L. Griffiths arxiv

Human language is highly polysemous. Many common words (e.g., 'bank' or 'palm') carry several distinct meanings that shape what humans communicate and imagine. Large language models (LLMs) have been shown to understand t…

Unsupervised Multimodal Word Discovery based on Double Articulation Analysis with Co-occurrence cues

2022-01-18 · Akira Taniguchi, Hiroaki Murakami, Ryo Ozaki, Tadahiro Taniguchi

Human infants acquire their verbal lexicon with minimal prior knowledge of language based on the statistical properties of phonological distributions and the co-occurrence of other sensory stimuli. This study proposes a …

MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware

2026-07-30 · Xianpeng Zhang, Jiahua Yang, Dongyu Chen, Lei zhang 외 arxiv

Multimodal long documents are core carriers of professional knowledge, where critical evidence is sparsely distributed across paragraphs and modalities. This easily causes key information omission and cross-modal halluci…

Document Summarization

CEMTM: Contextual Embedding-based Multimodal Topic Modeling

2025-09-14 · Amirhossein Abaskohi, Raymond Li, Chuyuan Li, Shafiq Joty 외 arxiv

We introduce CEMTM, a context-enhanced multimodal topic model designed to infer coherent and interpretable topic structures from both short and long documents containing text and images. CEMTM builds on fine-tuned large …

A Bayesian Approach to Multimodal Visual Dictionary Learning

2013-06-01 · CVPR 2013 6 · Go Irie, Dong Liu, Zhenguo Li, Shih-Fu Chang

nary learning methods rely on image descriptors alone or together with class labels. However, Web images are often associated with text data which may carry substantial information regarding image semantics, and may be e…

Bayesian InferenceClusteringDictionary LearningImage Categorization+1