paper-with-me

Papers

COSA: Concatenated Sample Pretrained Vision-Language Foundation Model

2023-06-15 · Sihan Chen, Xingjian He, Handong Li, Xiaojie Jin, Jiashi Feng, Jing Liu

Due to the limited scale and quality of video-text training corpus, most vision-language foundation models employ image-text datasets for pretraining and primarily focus on modeling visually semantic representations while disregarding temporal semantic representations and correlations. To address this issue, we propose COSA, a COncatenated SAmple pretrained vision-language foundation model. COSA jointly models visual contents and event-level temporal cues using only image-text corpora. We achieve this by sequentially concatenating multiple image-text pairs as inputs for pretraining. This transformation effectively converts existing image-text corpora into a pseudo long-form video-paragraph corpus, enabling richer scene transformations and explicit event-description correspondence. Extensive experiments demonstrate that COSA consistently improves performance across a broad range of downstream tasks, including long-form/short-form video-text tasks and image-text tasks such as retrieval, captioning, and question answering. Notably, COSA achieves state-of-the-art results on various competitive benchmarks. Code and model are released at https://github.com/TXH-mercury/COSA.

📄 PDF Abstract BibTeX arXiv:2306.09085

Code (1)

txh-mercury/cosa 공식 구현 pytorch

Tasks

FormmodelQuestion AnsweringRetrievalTGIF-FrameVideo CaptioningVideo Captioning on MSR-VTTVideo Question AnsweringVideo RetrievalVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Annotation-Free MIDI-to-Audio Synthesis via Concatenative Synthesis and Generative Refinement

2024-10-22 · Osamu Take, Taketo Akama

Recent MIDI-to-audio synthesis methods using deep neural networks have successfully generated high-quality, expressive instrumental tracks. However, these methods require MIDI annotations for supervised training, limitin…

Audio SynthesisDiversity

CoSA: Compressed Sensing-Based Adaptation of Large Language Models

2026-02-05 · Songtao Wei, Yi Li, Bohan Zhang, Zhichun Guo 외 arxiv

Parameter-Efficient Fine-Tuning (PEFT) has emerged as a practical paradigm for adapting large language models (LLMs) without updating all parameters. Most existing approaches, such as LoRA and PiSSA, rely on low-rank dec…

parameter-efficient fine-tuningNatural Language Understanding

Learning Structured Compressed Sensing with Automatic Resource Allocation

2024-10-24 · Han Wang, Eduardo Pérez, Iris A. M. Huijben, Hans van Gorp 외

Multidimensional data acquisition often requires extensive time and poses significant challenges for hardware and software regarding data storage and processing. Rather than designing a single compression matrix as in co…

compressed sensing

CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection

2024-10-10 · Guankun Wang, Han Xiao, Huxin Gao, Renrui Zhang 외

submucosal dissection (ESD) enables rapid resection of large lesions, minimizing recurrence rates and improving long-term overall survival. Despite these advantages, ESD is technically challenging and carries high risks …

Instruction Following

Orientation-aware Semantic Segmentation on Icosahedron Spheres

2019-07-30 · ICCV 2019 10 · Chao Zhang, Stephan Liwicki, William Smith, Roberto Cipolla

We address semantic segmentation on omnidirectional images, to leverage a holistic understanding of the surrounding scene for applications like autonomous driving systems. For the spherical domain, several methods recent…

Autonomous DrivingSemantic Segmentation