paper-with-me

홈 › Papers

Counting Everyday Objects in Everyday Scenes

2016-04-12 · CVPR 2017 7 · Prithvijit Chattopadhyay, Ramakrishna Vedantam, Ramprasaath R. Selvaraju, Dhruv Batra, Devi Parikh

We are interested in counting the number of instances of object classes in natural, everyday images. Previous counting approaches tackle the problem in restricted domains such as counting pedestrians in surveillance videos. Counts can also be estimated from outputs of other vision tasks like object detection. In this work, we build dedicated models for counting designed to tackle the large variance in counts, appearances, and scales of objects found in natural scenes. Our approach is inspired by the phenomenon of subitizing - the ability of humans to make quick assessments of counts given a perceptual signal, for small count values. Given a natural scene, we employ a divide and conquer strategy while incorporating context across the scene to adapt the subitizing idea to counting. Our approach offers consistent improvements over numerous baseline approaches for counting on the PASCAL VOC 2007 and COCO datasets. Subsequently, we study how counting can be used to improve object detection. We then show a proof of concept application of our counting methods to the task of Visual Question Answering, by studying the `how many?' questions in the VQA and COCO-QA datasets.

📄 PDF Abstract BibTeX arXiv:1604.03505

Code (1)

prithv1/cvpr2017_counting 공식 구현 torch

Tasks

ObjectObject Countingobject-detectionObject DetectionQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Linking WordNet to 3D Shapes

2018-01-01 · GWC 2018 1 · Angel X Chang, Rishi Mago, Pranav Krishna, Manolis Savva 외

We describe a project to link the Princeton WordNet to 3D representations of real objects and scenes. The goal is to establish a dataset that helps us to understand how people categorize everyday common objects via their…

NeRF-enabled Analysis-Through-Synthesis for ISAR Imaging of Small Everyday Objects with Sparse and Noisy UWB Radar Data

2024-10-14 · Md Farhan Tasnim Oshim, Albert Reed, Suren Jayasuriya, Tauhidur Rahman

Inverse Synthetic Aperture Radar (ISAR) imaging presents a formidable challenge when it comes to small everyday objects due to their limited Radar Cross-Section (RCS) and the inherent resolution constraints of radar syst…

NeRF

In the sight of my wearable camera: Classifying my visual experience

2013-04-26 · Alessandro Perina, Nebojsa Jojic

We introduce and we analyze a new dataset which resembles the input to biological vision systems much more than most previously published ones. Our analysis leaded to several important conclusions. First, it is possible …

EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation

2025-11-07 · Samarth Chopra, Alex McMoil, Ben Carnovale, Evan Sokolson 외 arxiv

While Vision-Language-Action (VLA) models map visual inputs and language instructions directly to robot actions, they often rely on costly hardware and struggle in novel or cluttered scenes. We introduce EverydayVLA, a 6…

Playing Text-Based Games with Common Sense

2020-12-04 · Sahith Dambekodi, Spencer Frazier, Prithviraj Ammanabrolu, Mark O. Riedl

Text based games are simulations in which an agent interacts with the world purely through natural language. They typically consist of a number of puzzles interspersed with interactions with common everyday objects and l…

Common Sense ReasoningDeep Reinforcement LearningLanguage ModelingLanguage Modelling+1