paper-with-me

Papers

Cross-stitched Multi-modal Encoders

2022-04-20 · Karan Singla, Daniel Pressel, Ryan Price, Bhargav Srinivas Chinnari, Yeon-Jun Kim, Srinivas Bangalore

In this paper, we propose a novel architecture for multi-modal speech and text input. We combine pretrained speech and text encoders using multi-headed cross-modal attention and jointly fine-tune on the target problem. The resultant architecture can be used for continuous token-level classification or utterance-level prediction acting on simultaneous text and speech. The resultant encoder efficiently captures both acoustic-prosodic and lexical information. We compare the benefits of multi-headed attention-based fusion for multi-modal utterance-level classification against a simple concatenation of pre-pooled, modality-specific representations. Our model architecture is compact, resource efficient, and can be trained on a single consumer GPU card.

📄 PDF Abstract BibTeX arXiv:2204.09227

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGPU

Similar Papers 제목 키워드 기반

Cross-stitched Multi-modal Encoders

2022-01-16 · ACL ARR January 2022 1 · Anonymous

In this paper, we propose a novel architecture for multi-modal speech and text input. We combine pretrained speech and text encoders using multi-headed cross-modal attention and jointly fine-tune on the target problem. T…

GPU

Interpreting the linear structure of vision-language model embedding spaces

2025-04-16 · Isabel Papadimitriou, Huangyuan Su, Thomas Fel, Sham Kakade 외

Vision-language models encode images and text in a joint space, minimizing the distance between corresponding image and text pairs. How are language and images organized in this joint space, and how do the models encode …

Language ModelingLanguage Modelling

Modification Takes Courage: Seamless Image Stitching via Reference-Driven Inpainting

2024-11-15 · Ziqi Xie, Xiao Lai, Weidong Zhao, Xianhui Liu 외

Current image stitching methods often produce noticeable seams in challenging scenarios such as uneven hue and large parallax. To tackle this problem, we propose the Reference-Driven Inpainting Stitcher (RDIStitcher), wh…

Image Stitching

ESC: Evolutionary Stitched Camera Calibration in the Wild

2024-04-19 · Grzegorz Rypeść, Grzegorz Kurzejamski

This work introduces a novel end-to-end approach for estimating extrinsic parameters of cameras in multi-camera setups on real-life sports fields. We identify the source of significant calibration errors in multi-camera …

Camera CalibrationImage SegmentationSemantic Segmentation

Image Quality Assessment for Omnidirectional Cross-reference Stitching

2019-04-10 · Kaiwen Yu, Jia Li, Yu Zhang, Yifan Zhao 외

Along with the development of virtual reality (VR), omnidirectional images play an important role in producing multimedia content with immersive experience. However, despite various existing approaches for omnidirectiona…

Image Quality AssessmentImage Stitching