paper-with-me

Papers

The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

2023-08-03 · Weiyun Wang, Min Shi, Qingyun Li, Wenhai Wang, Zhenhang Huang, Linjie Xing, Zhe Chen, Hao Li, Xizhou Zhu, Zhiguo Cao, Yushi Chen, Tong Lu, Jifeng Dai, Yu Qiao

We present the All-Seeing (AS) project: a large-scale data and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates human feedback and efficient models in the loop, we create a new dataset (AS-1B) with over 1 billion regions annotated with semantic tags, question-answering pairs, and detailed captions. It covers a wide range of 3.5 million common and rare concepts in the real world, and has 132.2 billion tokens that describe the concepts and their attributes. Leveraging this new dataset, we develop the All-Seeing model (ASM), a unified framework for panoptic visual recognition and understanding. The model is trained with open-ended language prompts and locations, which allows it to generalize to various vision and language tasks with remarkable zero-shot performance, including region-text retrieval, region recognition, captioning, and question-answering. We hope that this project can serve as a foundation for vision-language artificial general intelligence research. Models and the dataset shall be released at https://github.com/OpenGVLab/All-Seeing, and demo can be seen at https://huggingface.co/spaces/OpenGVLab/all-seeing.

📄 PDF Abstract BibTeX arXiv:2308.01907

Code (1)

opengvlab/all-seeing 공식 구현 pytorch

Tasks

AllQuestion AnsweringRetrievalText Retrieval

Similar Papers 제목 키워드 기반

JRDB-PanoTrack: An Open-world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments

2024-04-02 · CVPR 2024 1 · Duy-Tho Le, Chenhui Gou, Stavya Datta, Hengcan Shi 외

Autonomous robot systems have attracted increasing research attention in recent years, where environment understanding is a crucial step for robot navigation, human-robot interaction, and decision. Real-world robot syste…

Decision MakingPanoptic SegmentationRobot Navigation

Can we cover navigational perception needs of the visually impaired by panoptic segmentation?

2020-07-20 · Wei Mao, Jiaming Zhang, Kailun Yang, Rainer Stiefelhagen

Navigational perception for visually impaired people has been substantially promoted by both classic and deep learning based segmentation methods. In classic visual recognition methods, the segmentation models are mostly…

Deep LearningInstance SegmentationPanoptic SegmentationSegmentation+1

In-Place Panoptic Radiance Field Segmentation with Perceptual Prior for 3D Scene Understanding

2024-10-06 · Shenghao Li

Accurate 3D scene representation and panoptic understanding are essential for applications such as virtual reality, robotics, and autonomous driving. However, challenges persist with existing methods, including precise 2…

2D Panoptic SegmentationAutonomous DrivingPanoptic SegmentationScene Understanding

Visual Room 2.0: Seeing is Not Understanding for MLLMs

2025-11-17 · Haokun Li, Yazhou Zhang, Jizhi Ding, Qiuchi Li 외 arxiv

Can multi-modal large language models (MLLMs) truly understand what they can see? Extending Searle's Chinese Room into the multi-modal domain, this paper proposes the Visual Room argument: MLLMs may describe every visual…

Scene Understanding

Panoptic Lintention Network: Towards Efficient Navigational Perception for the Visually Impaired

2021-03-06 · Wei Mao, Jiaming Zhang, Kailun Yang, Rainer Stiefelhagen

Classic computer vision algorithms, instance segmentation, and semantic segmentation can not provide a holistic understanding of the surroundings for the visually impaired. In this paper, we utilize panoptic segmentation…

Instance SegmentationPanoptic SegmentationSegmentationSemantic Segmentation