paper-with-me

홈 › Papers

Bridging the gap to real-world language-grounded visual concept learning

2025-10-24 · Whie Jung, Semin Kim, Junee Kim, Seunghoon Hong arxiv

Human intelligence effortlessly interprets visual scenes along a rich spectrum of semantic dimensions. However, existing approaches to language-grounded visual concept learning are limited to a few predefined primitive axes, such as color and shape, and are typically explored in synthetic datasets. In this work, we propose a scalable framework that adaptively identifies image-related concept axes and grounds visual concepts along these axes in real-world scenes. Leveraging a pretrained vision-language model and our universal prompting strategy, our framework identifies a diverse image-related axes without any prior knowledge. Our universal concept encoder adaptively binds visual features to the discovered axes without introducing additional model parameters for each concept. To ground visual concepts along the discovered axes, we optimize a compositional anchoring objective, which ensures that each axis can be independently manipulated without affecting others. We demonstrate the effectiveness of our framework on subsets of ImageNet, CelebA-HQ, and AFHQ, showcasing superior editing capabilities across diverse real-world concepts that are too varied to be manually predefined. Our method also exhibits strong compositional generalization, outperforming existing visual concept learning and text-based editing methods. The code is available at https://github.com/whieya/Language-grounded-VCL.

📄 PDF Abstract BibTeX arXiv:2510.21412

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models

2026-05-28 · Zhongxi Chen, Yifan Han, Yanming Shao, Huanming Liu 외 arxiv

Vision-Language-Action (VLA) models have emerged as a promising paradigm for grounding visual-language understanding into real-world robotic manipulation. However, dexterous manipulation remains challenging for VLA polic…

Reinforcement Learning

Scan, Materialize, Simulate: A Generalizable Framework for Physically Grounded Robot Planning

2025-05-20 · Amine Elhafsi, Daniel Morton, Marco Pavone

Autonomous robots must reason about the physical consequences of their actions to operate effectively in unstructured, real-world environments. We present Scan, Materialize, Simulate (SMS), a unified framework that combi…

Semantic Segmentation

Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench

2025-10-30 · Fenfen Lin, Yesheng Liu, Haiyu Xu, Chen Yue 외 arxiv

Reading measurement instruments is effortless for humans and requires relatively little domain expertise, yet it remains surprisingly challenging for current vision-language models (VLMs) as we find in preliminary evalua…

NewtPhys: Do Foundation Models Understand Newtonian Physics?

2026-06-02 · Sebastian Cavada, Soumava Paul, Tuan-Hung Vu, Andrei Bursuc 외 arxiv

Previous work has evaluated physics reasoning in foundation models using synthetic or semi-synthetic scenes and visual question-answering tasks. However, these benchmarks emphasize high-level events and lack the visual f…

Bridging the Gap: Using Deep Acoustic Representations to Learn Grounded Language from Percepts and Raw Speech

2021-12-27 · Gaoussou Youssouf Kebe, Luke E. Richards, Edward Raff, Francis Ferraro 외

Learning to understand grounded language, which connects natural language to percepts, is a critical research area. Prior work in grounded language acquisition has focused primarily on textual inputs. In this work we dem…

Language Acquisitionspeech-recognitionSpeech Recognition