paper-with-me

홈 › Papers

BEHAVIOR Vision Suite: Customizable Dataset Generation via Simulation

2024-05-15 · CVPR 2024 1 · Yunhao Ge, Yihe Tang, Jiashu Xu, Cem Gokmen, Chengshu Li, Wensi Ai, Benjamin Jose Martinez, Arman Aydin, Mona Anvari, Ayush K Chakravarthy, Hong-Xing Yu, Josiah Wong, Sanjana Srivastava, Sharon Lee, Shengxin Zha, Laurent Itti, Yunzhu Li, Roberto Martín-Martín, Miao Liu, Pengchuan Zhang, Ruohan Zhang, Li Fei-Fei, Jiajun Wu

The systematic evaluation and understanding of computer vision models under varying conditions require large amounts of data with comprehensive and customized labels, which real-world vision datasets rarely satisfy. While current synthetic data generators offer a promising alternative, particularly for embodied AI tasks, they often fall short for computer vision tasks due to low asset and rendering quality, limited diversity, and unrealistic physical properties. We introduce the BEHAVIOR Vision Suite (BVS), a set of tools and assets to generate fully customized synthetic data for systematic evaluation of computer vision models, based on the newly developed embodied AI benchmark, BEHAVIOR-1K. BVS supports a large number of adjustable parameters at the scene level (e.g., lighting, object placement), the object level (e.g., joint configuration, attributes such as "filled" and "folded"), and the camera level (e.g., field of view, focal length). Researchers can arbitrarily vary these parameters during data generation to perform controlled experiments. We showcase three example application scenarios: systematically evaluating the robustness of models across different continuous axes of domain shift, evaluating scene understanding models on the same set of images, and training and evaluating simulation-to-real transfer for a novel vision task: unary and binary state prediction. Project website: https://behavior-vision-suite.github.io/

📄 PDF Abstract BibTeX arXiv:2405.09546

Code (0)

등록된 구현이 없습니다.

Tasks

Dataset GenerationScene Understanding

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Patchview: LLM-Powered Worldbuilding with Generative Dust and Magnet Visualization

2024-08-07 · John Joon Young Chung, Max Kreminski

Large language models (LLMs) can help writers build story worlds by generating world elements, such as factions, characters, and locations. However, making sense of many generated elements can be overwhelming. Moreover, …

Jochre 3 and the Yiddish OCR corpus

2025-01-14 · Assaf Urieli, Amber Clooney, Michelle Sigiel, Grisha Leyfer

We describe the construction of a publicly available Yiddish OCR Corpus, and describe and evaluate the open source OCR tool suite Jochre 3, including an Alto editor for corpus annotation, OCR software for Alto OCR layer …

Optical Character Recognition (OCR)

LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

2024-06-07 · Lukas Helff, Felix Friedrich, Manuel Brack, Kristian Kersting 외

This paper introduces LlavaGuard, a suite of VLM-based vision safeguards that address the critical need for reliable guardrails in the era of large-scale data and models. To this end, we establish a novel open framework,…

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents

2025-06-05 · Alexander Huang-Menders, Xinhang Liu, Andy Xu, Yuyao Zhang 외

SmartAvatar is a vision-language-agent-driven framework for generating fully rigged, animation-ready 3D human avatars from a single photo or textual prompt. While diffusion-based methods have made progress in general 3D …

Attribute

Tricolore: Multi-Behavior User Profiling for Enhanced Candidate Generation in Recommender Systems

2025-05-04 · Xiao Zhou, Zhongxiang Zhao, Hanze Guo

Online platforms aggregate extensive user feedback across diverse behaviors, providing a rich source for enhancing user engagement. Traditional recommender systems, however, typically optimize for a single target behavio…

Recommendation Systems