paper-with-me

홈 › Papers

Frame-Based Annotation of Multimodal Corpora: Tracking (A)Synchronies in Meaning Construction

2020-05-01 · LREC 2020 5 · Frederico Belcavello, Marcelo Viridiano, Alex Diniz da Costa, re, Ely Edison da Silva Matos, Tiago Timponi Torrent

Multimodal aspects of human communication are key in several applications of Natural Language Processing, such as Machine Translation and Natural Language Generation. Despite recent advances in integrating multimodality into Computational Linguistics, the merge between NLP and Computer Vision techniques is still timid, especially when it comes to providing fine-grained accounts for meaning construction. This paper reports on research aiming to determine appropriate methodology and develop a computational tool to annotate multimodal corpora according to a principled structured semantic representation of events, relations and entities: FrameNet. Taking a Brazilian television travel show as corpus, a pilot study was conducted to annotate the frames that are evoked by the audio and the ones that are evoked by visual elements. We also implemented a Multimodal Annotation tool which allows annotators to choose frames and locate frame elements both in the text and in the images, while keeping track of the time span in which those elements are active in each modality. Results suggest that adding a multimodal domain to the linguistic layer of annotation and analysis contributes both to enrich the kind of information that can be tagged in a corpus, and to enhance FrameNet as a model of linguistic cognition.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationText GenerationTranslation

Methods 이 논문이 사용한 방법론

Travel 설명 없음

Similar Papers 제목 키워드 기반

A Model-Based Approach to Synthetic Data Set Generation for Patient-Ventilator Waveforms for Machine Learning and Educational Use

2021-03-29 · A. van Diepen, T. H. G. F. Bakkes, A. J. R. De Bie, S. Turco 외

Although mechanical ventilation is a lifesaving intervention in the ICU, it has harmful side-effects, such as barotrauma and volutrauma. These harms can occur due to asynchronies. Asynchronies are defined as a mismatch b…

Charon: a FrameNet Annotation Tool for Multimodal Corpora

2022-05-24 · LREC (LAW) 2022 6 · Frederico Belcavello, Marcelo Viridiano, Ely Edison Matos, Tiago Timponi Torrent

This paper presents Charon, a web tool for annotating multimodal corpora with FrameNet categories. Annotation can be made for corpora containing both static images and video sequences paired - or not - with text sequence…

Divert More Attention to Vision-Language Object Tracking

2023-07-19 · Mingzhe Guo, Zhipeng Zhang, Liping Jing, Haibin Ling 외

Multimodal vision-language (VL) learning has noticeably pushed the tendency toward generic intelligence owing to emerging large foundation models. However, tracking, as a fundamental vision problem, surprisingly enjoys l…

AttributeObjectObject Tracking

Learning Tracking Representations from Single Point Annotations

2024-04-15 · Qiangqiang Wu, Antoni B. Chan

Existing deep trackers are typically trained with largescale video frames with annotated bounding boxes. However, these bounding boxes are expensive and time-consuming to annotate, in particular for large scale datasets.…

Contrastive LearningVisual Tracking

Transformer RGBT Tracking with Spatio-Temporal Multimodal Tokens

2024-01-03 · Dengdi Sun, Yajie Pan, Andong Lu, Chenglong Li 외

Many RGBT tracking researches primarily focus on modal fusion design, while overlooking the effective handling of target appearance changes. While some approaches have introduced historical frames or fuse and replace ini…

Rgb-T TrackingTemplate Matching