Extracting Latent Attributes from Video Scenes Using Text as Background Knowledge
Code (0)
등록된 구현이 없습니다.
Tasks
Coreference ResolutionInformation RetrievalSimilar Papers 제목 키워드 기반
Complex Event Detection via Multi-source Video Attributes
Complex events essentially include human, scenes, objects and actions that can be summarized by visual attributes, so leveraging relevant attributes properly could be helpful for event detection. Many works have exploite…
Event DetectionCoNeRF: Controllable Neural Radiance Fields
We extend neural 3D representations to allow for intuitive and interpretable user control beyond novel view rendering (i.e. camera control). We allow the user to annotate which part of the scene one wishes to control wit…
3D Face Modelling3D ReconstructionAttributeFew-Shot LearningExploring MLLM-Diffusion Information Transfer with MetaCanvas
Multimodal learning has rapidly advanced visual understanding, largely via multimodal large language models (MLLMs) that use powerful LLMs as cognitive cores. In visual generation, however, these powerful core models are…
Text-to-Image GenerationVideo GenerationPre-training Contextualized World Models with In-the-wild Videos for Reinforcement Learning
Unsupervised pre-training methods utilizing large and diverse datasets have achieved tremendous success across a range of domains. Recent work has investigated such unsupervised pre-training methods for model-based reinf…
Autonomous DrivingDecoderModel-based Reinforcement LearningTransfer Learning+2SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video Decomposition
To help agents reason about scenes in terms of their building blocks, we wish to extract the compositional structure of any given scene (in particular, the configuration and characteristics of objects comprising the scen…
Instance SegmentationObjectSemantic Segmentation