paper-with-me

홈 › Papers

SLiMe: Segment Like Me

2023-09-06 · Aliasghar Khani, Saeid Asgari Taghanaki, Aditya Sanghi, Ali Mahdavi Amiri, Ghassan Hamarneh

Significant strides have been made using large vision-language models, like Stable Diffusion (SD), for a variety of downstream tasks, including image editing, image correspondence, and 3D shape generation. Inspired by these advancements, we explore leveraging these extensive vision-language models for segmenting images at any desired granularity using as few as one annotated sample by proposing SLiMe. SLiMe frames this problem as an optimization task. Specifically, given a single training image and its segmentation mask, we first extract attention maps, including our novel "weighted accumulated self-attention map" from the SD prior. Then, using the extracted attention maps, the text embeddings of Stable Diffusion are optimized such that, each of them, learn about a single segmented region from the training image. These learned embeddings then highlight the segmented region in the attention maps, which in turn can then be used to derive the segmentation map. This enables SLiMe to segment any real-world image during inference with the granularity of the segmented region in the training image, using just one example. Moreover, leveraging additional training data when available, i.e. few-shot, improves the performance of SLiMe. We carried out a knowledge-rich set of experiments examining various design factors and showed that SLiMe outperforms other existing one-shot and few-shot segmentation methods.

📄 PDF Abstract BibTeX arXiv:2309.03179

Code (1)

aliasgharkhani/slime 공식 구현 jax

Tasks

3D Shape GenerationSegmentation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate

2025-06-26 · Hankun Wang, Yiwei Guo, Chongtian Shao, Bohan Li 외

Neural speech codecs have been widely used in audio compression and various downstream tasks. Current mainstream codecs are fixed-frame-rate (FFR), which allocate the same number of tokens to every equal-duration slice. …

Audio Compression

SLIME: Stabilized Likelihood Implicit Margin Enforcement for Preference Optimization

2026-02-02 · Maksim Afanasyev, Illarion Iov arxiv

Direct preference optimization methods have emerged as a computationally efficient alternative to Reinforcement Learning from Human Feedback (RLHF) for aligning Large Language Models (LLMs). Latest approaches have stream…

Reinforcement Learning

SLIMER-IT: Zero-Shot NER on Italian Language

2024-09-24 · Andrew Zamai, Leonardo Rigutini, Marco Maggini, Andrea Zugarini

Traditional approaches to Named Entity Recognition (NER) frame the task into a BIO sequence labeling problem. Although these systems often excel in the downstream task at hand, they require extensive annotated data and s…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

EINCASM: Emergent Intelligence in Neural Cellular Automaton Slime Molds

2023-05-22 · Aidan Barbieux, Rodrigo Canaan

This paper presents EINCASM, a prototype system employing a novel framework for studying emergent intelligence in organisms resembling slime molds. EINCASM evolves neural cellular automata with NEAT to maximize cell grow…

Exposing Pink Slime Journalism: Linguistic Signatures and Robust Detection Against LLM-Generated Threats

2025-12-05 · Sadat Shahriar, Navid Ayoobi, Arjun Mukherjee, Mostafa Musharrat 외 arxiv

The local news landscape, a vital source of reliable information for 28 million Americans, faces a growing threat from Pink Slime Journalism, a low-quality, auto-generated articles that mimic legitimate local reporting. …