Power-LLaVA: Large Language and Vision Assistant for Power Transmission Line Inspection
The inspection of power transmission line has achieved notable achievements in the past few years, primarily due to the integration of deep learning technology. However, current inspection approaches continue to encounter difficulties in generalization and intelligence, which restricts their further applicability. In this paper, we introduce Power-LLaVA, the first large language and vision assistant designed to offer professional and reliable inspection services for power transmission line by engaging in dialogues with humans. Moreover, we also construct a large-scale and high-quality dataset specialized for the inspection task. By employing a two-stage training strategy on the constructed dataset, Power-LLaVA demonstrates exceptional performance at a comparatively low training cost. Extensive experiments further prove the great capabilities of Power-LLaVA within the realm of power transmission line inspection. Code shall be released.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
Conversational generative AI has demonstrated remarkable promise for empowering biomedical practitioners, but current investigations focus on unimodal text. Multimodal conversational AI has seen rapid progress by leverag…
Image ClassificationInstruction FollowingLanguage ModellingQuestion Answering+3STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering
Large Vision-Language Models (LVLMs) have shown significant potential in assisting medical diagnosis by leveraging extensive biomedical datasets. However, the advancement of medical image understanding and reasoning crit…
Medical DiagnosisMedical Question AnsweringMedical Visual Question AnsweringQuestion Answering+2LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
Visual Spatial Description (VSD) aims to generate texts that describe the spatial relationships between objects within images. Traditional visual spatial relationship classification (VSRC) methods typically output the sp…
DiversityInstruction FollowingLanguage ModelingLanguage Modelling+2LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Multimodal Large Language Models (MLLMs) are widely used for visual perception, understanding, and reasoning. However, long video processing and precise moment retrieval remain challenging due to LLMs' limited context si…
Moment RetrievalNatural Language Moment RetrievalRetrievalRLLaVA: An RL-central Framework for Language and Vision Assistants
We present an RL-central framework for Language and Vision Assistants (RLLaVA) with its formulation of Markov decision process (MDP). RLLaVA decouples RL algorithmic logic from model architecture and distributed executio…