paper-with-me

Papers Object

“Object” 태그가 달린 논문 10,696편 · 필터 해제

SeC: Advancing Complex Video Object Segmentation via Progressive Concept Construction

2025-07-21 · Zhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Songxin He 외

Video Object Segmentation (VOS) is a core task in computer vision, requiring models to track and segment target objects across video frames. Despite notable advances with recent efforts, current techniques still lag behi…

ObjectSegmentationSemantic SegmentationVideo Object Segmentation+1

AutoPartGen: Autogressive 3D Part Generation and Discovery

2025-07-17 · Minghao Chen, Jianyuan Wang, Roman Shapovalov, Tom Monnier 외

We introduce AutoPartGen, a model that generates objects composed of 3D parts in an autoregressive manner. This model can take as input an image of an object, 2D masks of the object's parts, or an existing 3D object, and…

3D Generation3D ReconstructionObject

A Real-Time System for Egocentric Hand-Object Interaction Detection in Industrial Domains

2025-07-17 · Antonio Finocchiaro, Alessandro Sebastiano Catinello, Michele Mazzamuto, Rosario Leonardi 외

Hand-object interaction detection remains an open challenge in real-time applications, where intuitive user experiences depend on fast and accurate detection of interactions with surrounding objects. We propose an effici…

Action RecognitionHand-Object Interaction DetectionMambaObject+2

A Multi-Level Similarity Approach for Single-View Object Grasping: Matching, Planning, and Fine-Tuning

2025-07-16 · Hao Chen, Takuya Kiyokawa, Zhengtao Hu, Weiwei Wan 외

Grasping unknown objects from a single view has remained a challenging topic in robotics due to the uncertainty of partial observation. Recent advances in large-scale models have led to benchmark solutions such as GraspN…

ObjectPoint Cloud Registration

RoHOI: Robustness Benchmark for Human-Object Interaction Detection

2025-07-12 · Di Wen, Kunyu Peng, Kailun Yang, Yufan Chen 외

Human-Object Interaction (HOI) detection is crucial for robot-human assistance, enabling context-aware support. However, models trained on clean datasets degrade in real-world conditions due to unforeseen corruptions, le…

Human-Object Interaction DetectionObject

Car Object Counting and Position Estimation via Extension of the CLIP-EBC Framework

2025-07-11 · Seoik Jung, Taekyung Song

In this paper, we investigate the applicability of the CLIP-EBC framework, originally designed for crowd counting, to car object counting using the CARPK dataset. Experimental results show that our model achieves second-…

ClusteringCrowd CountingObjectObject Counting+1

MUVOD: A Novel Multi-view Video Object Segmentation Dataset and A Benchmark for 3D Segmentation

2025-07-10 · Bangning Wei, Joshua Maraval, Meriem Outtas, Kidiyo Kpalma 외

The application of methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3D GS) have steadily gained popularity in the field of 3D object segmentation in static scenes. These approaches demonstrate ef…

NeRFObjectScene UnderstandingSegmentation+4

EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow

2025-07-08 · Yixiang Chen, Peiyan Li, Yan Huang, Jiabing Yang 외

Current language-guided robotic manipulation systems often require low-level action-labeled datasets for imitation learning. While object-centric flow prediction methods mitigate this issue, they remain limited to scenar…

Deformable Object ManipulationImitation LearningObject

ECORE: Energy-Conscious Optimized Routing for Deep Learning Models at the Edge

2025-07-08 · Daghash K. Alqahtani, Maria A. Rodriguez, Muhammad Aamir Cheema, Hamid Rezatofighi 외

Edge computing enables data processing closer to the source, significantly reducing latency an essential requirement for real-time vision-based analytics such as object detection in surveillance and smart city environmen…

Edge-computingObjectobject-detectionObject Detection+1

DreamGrasp: Zero-Shot 3D Multi-Object Reconstruction from Partial-View Images for Robotic Manipulation

2025-07-08 · Young Hun Kim, Seungyeon Kim, Yonghyeon LEE, Frank Chongwoo Park

Partial-view 3D recognition -- reconstructing 3D geometry and identifying object instances from a few sparse RGB images -- is an exceptionally challenging yet practically essential task, particularly in cluttered, occlud…

3D geometry3D ReconstructionContrastive LearningInstance Segmentation+3

Dyn-O: Building Structured World Models with Object-Centric Representations

2025-07-04 · Zizhao Wang, Kaixin Wang, Li Zhao, Peter Stone 외

World models aim to capture the dynamics of the environment, enabling agents to predict and plan for future states. In most scenarios of interest, the dynamics are highly centered on interactions among objects within the…

Object

LMPNet for Weakly-supervised Keypoint Discovery

2025-07-03 · Pei Guo, Ryan Farrell

In this work, we explore the task of semantic object keypoint discovery weakly-supervised by only category labels. This is achieved by transforming discriminatively-trained intermediate layer filters into keypoint detect…

ObjectPose Estimation

NOCTIS: Novel Object Cyclic Threshold based Instance Segmentation

2025-07-02 · Max Gandyra, Alessandro Santonicola, Michael Beetz

Instance segmentation of novel objects instances in RGB images, given some example images for each object, is a well known problem in computer vision. Designing a model general enough to be employed, for all kinds of nov…

Instance SegmentationObjectSegmentationSemantic Segmentation

Refine Any Object in Any Scene

2025-06-30 · Ziwei Chen, Ziling Liu, Zitong Huang, Mingqi Gao 외

Viewpoint missing of objects is common in scene reconstruction, as camera paths typically prioritize capturing the overall scene structure rather than individual objects. This makes it highly challenging to achieve high-…

Novel View SynthesisObject

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

2025-06-30 · Xiangtai Li, Tao Zhang, Yanwei Li, Haobo Yuan 외

Multimodal Large Language Models (MLLMs) demonstrate a complex understanding of scenes, benefiting from large-scale and high-quality datasets. Most existing caption datasets lack the ground locations and relations for vi…

Caption GenerationObjectVisual Grounding

Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval

2025-06-28 · Li-Cheng Shen, Jih-Kang Hsieh, Wei-Hua Li, Chu-Song Chen

Text-to-image retrieval (TIR) aims to find relevant images based on a textual query, but existing approaches are primarily based on whole-image captions and lack interpretability. Meanwhile, referring expression segmenta…

Cross-Modal RetrievalImage CaptioningImage RetrievalLarge Language Model+9

Deterministic Object Pose Confidence Region Estimation

2025-06-28 · Jinghao Wang, Zhang Li, Zi Wang, Banglei Guan 외

6D pose confidence region estimation has emerged as a critical direction, aiming to perform uncertainty quantification for assessing the reliability of estimated poses. However, current sampling-based approach suffers fr…

Conformal PredictionObjectPose EstimationUncertainty Quantification

Towards Reliable Detection of Empty Space: Conditional Marked Point Processes for Object Detection

2025-06-26 · Tobias J. Riedlinger, Kira Maag, Hanno Gottschalk

Deep neural networks have set the state-of-the-art in computer vision tasks such as bounding box detection and semantic segmentation. Object detectors and segmentation models assign confidence scores to predictions, refl…

Objectobject-detectionObject DetectionPoint Processes+1

PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object Modeling

2025-06-26 · Hao Zhang, Haolan Xu, Chun Feng, Varun Jampani 외

Skinning and rigging are fundamental components in animation, articulated object reconstruction, motion transfer, and 4D generation. Existing approaches predominantly rely on Linear Blend Skinning (LBS), due to its simpl…

ObjectObject ReconstructionPose Transfer

SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification

2025-06-26 · Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran, Trung-Nghia Le

Retrieving 3D objects in complex indoor environments using only a masked 2D image and a natural language description presents significant challenges. The ROOMELSA challenge limits access to full 3D scene context, complic…

3D Object RetrievalObjectRe-RankingRetrieval
1–20 / 10,696 다음 →