paper-with-me

홈 › Papers

VDCook:DIY video data cook your MLLMs

2026-03-04 · Chengwei Wu arxiv

We introduce VDCook: a self-evolving video data operating system, a configurable video data construction platform for researchers and vertical domain teams. Users initiate data requests via natural language queries and adjustable parameters (scale, retrieval-synthesis ratio, quality threshold). The system automatically performs query optimization, concurrently running real video retrieval and controlled synthesis modules. It ultimately generates in-domain data packages with complete provenance and metadata, along with reproducible Notebooks. Unlike traditional static, one-time-built datasets, VDCook enables continuous updates and domain expansion through its automated data ingestion mechanism based on MCP (Model Context Protocol)\cite{mcp2024anthropic}, transforming datasets into dynamically evolving open ecosystems. The system also provides multi-dimensional metadata annotation (scene segmentation, motion scoring, OCR ratio, automatic captioning, etc.), laying the foundation for flexible subsequent data `cooking' and indexing\cite{vlogger}. This platform aims to significantly lower the barrier to constructing specialized video training datasets through infrastructure-level solutions, while supporting community contributions and a governance-enabled data expansion paradigm. \textbf{Project demo:} https://screenapp.io/app/v/WP0SvffgsH

📄 PDF Abstract BibTeX arXiv:2603.05539

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language QueriesScene SegmentationVideo Retrieval

Similar Papers 제목 키워드 기반

A Systematic Evaluation of Positional Bias in Multi-Video Summarization with MLLMs

2026-06-03 · Huangchen Xu, Yuan Wu, Yi Chang arxiv

Multimodal Large Language Models (MLLMs) are increasingly used for video understanding, yet their reliability under multi-video inputs remains poorly understood. We study positional bias in multi-video summarization, whe…

Video Summarization

Eating Your Own Cooking: Automatically Linking Wordnet Synsets of Two Languages

2012-12-01 · COLING 2012 12 · Salil Joshi, Arindam Chatterjee, Arun Karthikeyan Karra, Pushpak Bhattacharyya
Vocal Bursts Valence Prediction

GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video

2026-07-03 · Fang Liu, Jinpeng Chen, Ke Xu, Yuhao Liu 외 arxiv

While multimodal Large Language Models (MLLMs) excel at offline video understanding, an interesting question of how far they are from serving as a real-time procedural coach remains unknown. Such a role typically require…

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

2025-09-16 · Tianyu Yu, Zefan Wang, Chongyi Wang, Fuwei Huang 외 arxiv

Multimodal Large Language Models (MLLMs) are undergoing rapid progress and represent the frontier of AI development. However, their training and inference efficiency have emerged as a core bottleneck in making MLLMs more…

Reinforcement Learning

EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering

2025-08-14 · Yanjun Li, Yuqian Fu, Tianwen Qian, Qi'ao Xu 외 arxiv

Recent advances in Multimodal Large Language Models (MLLMs) have significantly pushed the frontier of egocentric video question answering (EgocentricQA). However, existing benchmarks and studies are mainly limited to com…

Video Question AnsweringReinforcement LearningDomain Generalization