paper-with-me

Papers

A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering

2023-05-12 · Chaoning Zhang, Joseph Cho, Fachrina Dewi Puspitasari, Sheng Zheng, Chenghao Li, Yu Qiao, Taegoo Kang, Xinru Shan, Chenshuang Zhang, Caiyan Qin, Francois Rameau, Lik-Hang Lee, Sung-Ho Bae, Choong Seon Hong

The Segment Anything Model (SAM), developed by Meta AI Research, represents a significant breakthrough in computer vision, offering a robust framework for image and video segmentation. This survey provides a comprehensive exploration of the SAM family, including SAM and SAM 2, highlighting their advancements in granularity and contextual understanding. Our study demonstrates SAM's versatility across a wide range of applications while identifying areas where improvements are needed, particularly in scenarios requiring high granularity and in the absence of explicit prompts. By mapping the evolution and capabilities of SAM models, we offer insights into their strengths and limitations and suggest future research directions, including domain-specific adaptations and enhanced memory and propagation mechanisms. We believe that this survey comprehensively covers the breadth of SAM's applications and challenges, setting the stage for ongoing advancements in segmentation technology.

📄 PDF Abstract BibTeX arXiv:2306.06211

Code (0)

등록된 구현이 없습니다.

Tasks

Edge DetectionmodelPrompt EngineeringSurveyVideo SegmentationVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
SAM 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

A Comprehensive Survey on Segment Anything Model for Vision and Beyond

2023-05-14 · Chunhui Zhang, Li Liu, Yawen Cui, Guanjie Huang 외

Artificial intelligence (AI) is evolving towards artificial general intelligence, which refers to the ability of an AI system to perform a wide range of tasks and exhibit a level of intelligence similar to that of a huma…

Segment Anything Model (SAM) Meets Glass: Mirror and Transparent Objects Cannot Be Easily Detected

2023-04-29 · Dongsheng Han, Chaoning Zhang, Yu Qiao, Maryam Qamar 외

Meta AI Research has recently released SAM (Segment Anything Model) which is trained on a large segmentation dataset of over 1 billion masks. As a foundation model in the field of computer vision, SAM (Segment Anything M…

SegmentationSemantic SegmentationTransparent objects

WeakSAM: Segment Anything Meets Weakly-supervised Instance-level Recognition

2024-02-22 · Lianghui Zhu, Junwei Zhou, Yan Liu, Xin Hao 외

Weakly supervised visual recognition using inexact supervision is a critical yet challenging learning problem. It significantly reduces human labeling costs and traditionally relies on multi-instance learning and pseudo-…

Image-level Supervised Instance Segmentationobject-detectionObject DetectionSegmentation+2

Detect Any Deepfakes: Segment Anything Meets Face Forgery Detection and Localization

2023-06-29 · Yingxin Lai, Zhiming Luo, Zitong Yu

The rapid advancements in computer vision have stimulated remarkable progress in face forgery techniques, capturing the dedicated attention of researchers committed to detecting forgeries and precisely localizing manipul…

DeepFake DetectionFace Swapping

Segment Anything for Videos: A Systematic Survey

2024-07-31 · Chunhui Zhang, Yawen Cui, Weilin Lin, Guanjie Huang 외

The recent wave of foundation models has witnessed tremendous success in computer vision (CV) and beyond, with the segment anything model (SAM) having sparked a passion for exploring task-agnostic visual foundation model…

Image SegmentationRobot Manipulation GeneralizationSemantic SegmentationSurvey+4