paper-with-me

Papers

TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

2024-10-08 · Jeremy Andrew Irvin, Emily Ruoyu Liu, Joyce Chuyi Chen, Ines Dormoy, Jinyoung Kim, Samar Khanna, Zhuo Zheng, Stefano Ermon

Large vision and language assistants have enabled new capabilities for interpreting natural images. These approaches have recently been adapted to earth observation data, but they are only able to handle single image inputs, limiting their use for many real-world tasks. In this work, we develop a new vision and language assistant called TEOChat that can engage in conversations about temporal sequences of earth observation data. To train TEOChat, we curate an instruction-following dataset composed of many single image and temporal tasks including building change and damage assessment, semantic change detection, and temporal scene classification. We show that TEOChat can perform a wide variety of spatial and temporal reasoning tasks, substantially outperforming previous vision and language assistants, and even achieving comparable or better performance than several specialist models trained to perform specific tasks. Furthermore, TEOChat achieves impressive zero-shot performance on a change detection and change question answering dataset, outperforms GPT-4o and Gemini 1.5 Pro on multiple temporal tasks, and exhibits stronger single image capabilities than a comparable single image instruction-following model on scene classification, visual question answering, and captioning. We publicly release our data, model, and code at https://github.com/ermongroup/TEOChat .

📄 PDF Abstract BibTeX arXiv:2410.06234

Code (1)

ermongroup/teochat 공식 구현 pytorch

Tasks

Change DetectionEarth ObservationInstruction FollowingQuestion AnsweringScene ClassificationTemporal SequencesVisual Question Answering

Similar Papers 제목 키워드 기반

VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation

2026-03-28 · Jihwan Hong, Jaeyoung Do arxiv

Referring Video Object Segmentation (RVOS) aims to segment target objects in videos based on natural language descriptions. However, fixed keyframe-based approaches that couple a vision language model with a separate pro…

Referring Video Object Segmentation

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

2024-11-21 · Weiheng Lu, Jian Li, An Yu, Ming-Ching Chang 외

Multimodal Large Language Models (MLLMs) are widely used for visual perception, understanding, and reasoning. However, long video processing and precise moment retrieval remain challenging due to LLMs' limited context si…

Moment RetrievalNatural Language Moment RetrievalRetrieval

A Foundational Multimodal Vision Language AI Assistant for Human Pathology

2023-12-13 · Ming Y. Lu, Bowen Chen, Drew F. K. Williamson, Richard J. Chen 외

The field of computational pathology has witnessed remarkable progress in the development of both task-specific predictive models and task-agnostic self-supervised vision encoders. However, despite the explosive growth o…

Decision MakingDiagnosticLanguage ModellingLarge Language Model+1

Valley: Video Assistant with Large Language model Enhanced abilitY

2023-06-12 · Ruipu Luo, Ziwang Zhao, Min Yang, Junwei DOng 외

Large language models (LLMs), with their remarkable conversational capabilities, have demonstrated impressive performance across various applications and have emerged as formidable AI assistants. In view of this, it rais…

Action RecognitionInstruction FollowingLanguage ModelingLanguage Modelling+3

Visuospatial Cognitive Assistant

2025-05-18 · Qi Feng

Video-based spatial cognition is vital for robotics and embodied AI but challenges current Vision-Language Models (VLMs). This paper makes two key contributions. First, we introduce ViCA (Visuospatial Cognitive Assistant…

Spatial Reasoning