paper-with-me

홈 › Papers

Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment

2025-03-12 · Xiaowei Bi, Zheyuan Xu

Long Video Question Answering (LVQA) is challenging due to the need for temporal reasoning and large-scale multimodal data processing. Existing methods struggle with retrieving cross-modal information from long videos, especially when relevant details are sparsely distributed. We introduce UMaT (Unified Multi-modal as Text), a retrieval-augmented generation (RAG) framework that efficiently processes extremely long videos while maintaining cross-modal coherence. UMaT converts visual and auditory data into a unified textual representation, ensuring semantic and temporal alignment. Short video clips are analyzed using a vision-language model, while automatic speech recognition (ASR) transcribes dialogue. These text-based representations are structured into temporally aligned segments, with adaptive filtering to remove redundancy and retain salient details. The processed data is embedded into a vector database, enabling precise retrieval of dispersed yet relevant content. Experiments on a benchmark LVQA dataset show that UMaT outperforms existing methods in multimodal integration, long-form video understanding, and sparse information retrieval. Its scalability and interpretability allow it to process videos over an hour long while maintaining semantic and temporal coherence. These findings underscore the importance of structured retrieval and multimodal synchronization for advancing LVQA and long-form AI systems.

📄 PDF Abstract BibTeX arXiv:2503.09081

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Information RetrievalQuestion AnsweringRAGRetrievalRetrieval-augmented GenerationSparse Information Retrievalspeech-recognitionSpeech RecognitionVideo Question AnsweringVideo Understanding

Similar Papers 제목 키워드 기반

Slavic Forest, Norwegian Wood

2017-04-01 · WS 2017 4 · Rudolf Rosa, Daniel Zeman, David Mare{\v{c}}ek, Zden{\v{e}}k {\v{Z}}abokrtsk{\'y}

We once had a corp, or should we say, it once had us They showed us its tags, isn{'}t it great, unified tags They asked us to parse and they told us to use everything So we looked around and we noticed there was near not…

Dependency ParsingMachine TranslationWord Alignment

Edit Everything: A Text-Guided Generative System for Images Editing

2023-04-27 · Defeng Xie, Ruichen Wang, Jian Ma, Chen Chen 외

We introduce a new generative system called Edit Everything, which can take image and text inputs and produce image outputs. Edit Everything allows users to edit images using simple text instructions. Our system designs …

Scaling Laws for Transfer

2021-02-02 · Danny Hernandez, Jared Kaplan, Tom Henighan, Sam McCandlish

We study empirical scaling laws for transfer learning between distributions in an unsupervised, fine-tuning setting. When we train increasingly large neural networks from-scratch on a fixed-size dataset, they eventually …

Transfer Learning

Unsupervised Summarization by Jointly Extracting Sentences and Keywords

2020-09-16 · Zongyi Li, Xiaoqing Zheng, Jun He

We present RepRank, an unsupervised graph-based ranking model for extractive multi-document summarization in which the similarity between words, sentences, and word-to-sentence can be estimated by the distances between t…

DiversityDocument SummarizationMulti-Document SummarizationSentence+1

Internet of Everything in the 6G Era: Paradigms, Enablers, Potentials and Future Directions

2026-04-27 · Driss Choukri, Essaid Sabir, Elmahdi Driouch, Abdelkrim Haqiq arxiv

The Internet of Everything (IoE) represents an evolution of the Internet of Things (IoT) by integrating people, data, processes, and things into a unified intelligent ecosystem. IoE aims to enhance automation, decision-m…