paper-with-me

홈 › Papers

B-RIGHT: Benchmark Re-evaluation for Integrity in Generalized Human-Object Interaction Testing

2025-01-28 · Yoojin Jang, Junsu Kim, Hayeon Kim, Eun-ki Lee, Eun-Sol Kim, Seungryul Baek, Jaejun Yoo

Human-object interaction (HOI) is an essential problem in artificial intelligence (AI) which aims to understand the visual world that involves complex relationships between humans and objects. However, current benchmarks such as HICO-DET face the following limitations: (1) severe class imbalance and (2) varying number of train and test sets for certain classes. These issues can potentially lead to either inflation or deflation of model performance during evaluation, ultimately undermining the reliability of evaluation scores. In this paper, we propose a systematic approach to develop a new class-balanced dataset, Benchmark Re-evaluation for Integrity in Generalized Human-object Interaction Testing (B-RIGHT), that addresses these imbalanced problems. B-RIGHT achieves class balance by leveraging balancing algorithm and automated generation-and-filtering processes, ensuring an equal number of instances for each HOI class. Furthermore, we design a balanced zero-shot test set to systematically evaluate models on unseen scenario. Re-evaluating existing models using B-RIGHT reveals substantial the reduction of score variance and changes in performance rankings compared to conventional HICO-DET. Our experiments demonstrate that evaluation under balanced conditions ensure more reliable and fair model comparisons.

📄 PDF Abstract BibTeX arXiv:2501.16724

Code (1)

hellog2n/b-right 공식 구현

Tasks

Human-Object Interaction Detection

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages

2026-02-28 · Jason Lucas, Matt Murtagh-White, Adaku Uchendu, Ali Al-Lawati 외 arxiv

Multilingual falsehoods threaten information integrity worldwide, yet detection benchmarks remain confined to English or a few high-resource languages, leaving low-resource linguistic communities without robust defense t…

NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles

2026-05-03 · Xiao Jia arxiv

Outcome-only evaluation under-specifies whether an evaluated agent profile preserves the commitments required to solve a multi-turn task coherently. NeuroState-Bench is a human-calibrated benchmark that operationalizes c…

Solving Copyright Infringement on Short Video Platforms: Novel Datasets and an Audio Restoration Deep Learning Pipeline

2025-04-30 · Minwoo Oh, Minsu Park, Eunil Park

Short video platforms like YouTube Shorts and TikTok face significant copyright compliance challenges, as infringers frequently embed arbitrary background music (BGM) to obscure original soundtracks (OST) and evade conte…

Music Source SeparationVideo Restoration

Minding rights: Mapping ethical and legal foundations of 'neurorights'

2023-02-13 · Sjors Ligthart, Marcello Ienca, Gerben Meynen, Fruzsina Molnar-Gabor 외

The rise of neurotechnologies, especially in combination with AI-based methods for brain data analytics, has given rise to concerns around the protection of mental privacy, mental integrity and cognitive liberty - often …

It's TIME: Towards the Next Generation of Time Series Forecasting Benchmarks

2026-02-12 · Zhongzheng Qiao, Sheng Pan, Anni Wang, Viktoriya Zhukova 외 arxiv

Time series foundation models (TSFMs) are revolutionizing the forecasting landscape from specific dataset modeling to generalizable task evaluation. However, we contend that existing benchmarks exhibit common limitations…

Time Series Forecasting