paper-with-me

홈 › Papers

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding

2025-06-09 · Nianbo Zeng, Haowen Hou, Fei Richard Yu, Si Shi, Ying Tiffany He

Despite recent advances in retrieval-augmented generation (RAG) for video understanding, effectively understanding long-form video content remains underexplored due to the vast scale and high complexity of video data. Current RAG approaches typically segment videos into fixed-length chunks, which often disrupts the continuity of contextual information and fails to capture authentic scene boundaries. Inspired by the human ability to naturally organize continuous experiences into coherent scenes, we present SceneRAG, a unified framework that leverages large language models to segment videos into narrative-consistent scenes by processing ASR transcripts alongside temporal metadata. SceneRAG further sharpens these initial boundaries through lightweight heuristics and iterative correction. For each scene, the framework fuses information from both visual and textual modalities to extract entity relations and dynamically builds a knowledge graph, enabling robust multi-hop retrieval and generation that account for long-range dependencies. Experiments on the LongerVideos benchmark, featuring over 134 hours of diverse content, confirm that SceneRAG substantially outperforms prior baselines, achieving a win rate of up to 72.5 percent on generation tasks.

📄 PDF Abstract BibTeX arXiv:2506.07600

Code (0)

등록된 구현이 없습니다.

Tasks

RAGRetrievalRetrieval-augmented GenerationVideo Understanding

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Attention 설명 없음
WordPiece 설명 없음
Weight Decay 설명 없음
Multi-Head Attention 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…

Similar Papers 제목 키워드 기반

Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning

2025-11-08 · Fei Yu, Quan Deng, Shengeng Tang, Yuehua Li 외 arxiv

Understanding 3D scenes in open-world settings poses fundamental challenges for vision and robotics, particularly due to the limitations of closed-vocabulary supervision and static annotations. To address this, we propos…

Scene Graph GenerationScene UnderstandingQuestion AnsweringVisual Grounding

SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation

2024-11-29 · Wenjia Wang, Liang Pan, Zhiyang Dou, Jidong Mei 외

Simulating stylized human-scene interactions (HSI) in physical environments is a challenging yet fascinating task. Prior works emphasize long-term execution but fall short in achieving both diverse style and physical pla…

Motion PlanningRAGRetrievalRetrieval-augmented Generation+2

From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG

2026-05-14 · Guanhua Chen, Chuyue Huang, Yutong Yao, Shudong Liu 외 arxiv

Multimodal Retrieval-Augmented Generation (RAG) systems retrieve evidence at coarse granularities (entire images or scenes), creating a mismatch with fine-grained user queries and making failures unverifiable. We introdu…

SGR3 Model: Scene Graph Retrieval-Reasoning Model in 3D

2026-03-04 · Zirui Wang, Ruiping Liu, Yufan Chen, Junwei Zheng 외 arxiv

3D scene graphs provide a structured representation of object entities and their relationships, enabling high-level interpretation and reasoning for robots while remaining intuitively understandable to humans. Existing a…

Scene Graph GenerationRelational Reasoning3D Reconstruction

ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation

2026-08-31 · Jiawei Zhang, Hongsong Wang, Pan Zhou arxiv

Text-driven 3D indoor scene generation has advanced from dataset-bound layout modeling to open-vocabulary synthesis with large language and vision-language models. Yet existing methods remain limited: one-pass generators…

Scene Generation