Chahta Anumpa: A multimodal corpus of the Choctaw Language
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Exploring a Choctaw Language Corpus with Word Vectors and Minimum Distance Length
This work introduces additions to the corpus ChoCo, a multimodal corpus for the American indigenous language Choctaw. Using texts from the corpus, we develop new computational resources by using two off-the-shelf tools: …
EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture
Despite recent advances in AI, multimodal culturally diverse datasets are still limited, particularly for regions in the Middle East and Africa. In this paper, we introduce EgMM-Corpus, a multimodal dataset dedicated to …
MultiNews: A Web collection of an Aligned Multimodal and Multilingual Corpus
Integrating Natural Language Processing (NLP) and computer vision is a promising effort. However, the applicability of these methods directly depends on the availability of a specific multimodal data that includes images…
ArticlesContent-Based Image RetrievalImage RetrievalMachine Translation+1LAMP: A Multimodal Web Platform for Collaborative Linguistic Analysis
This paper describes the underlying software platform used to develop and publish annotations for the Quranic Arabic Corpus (QAC). The QAC (Dukes, Atwell and Habash, 2011) is a multimodal language resource that integrate…
Part-Of-Speech TaggingTranslationOmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studie…
In-Context Learning