paper-with-me

Papers

On the Reliability of Cue Conflict and Beyond

2026-03-11 · Pum Jun Kim, Seung-Ah Lee, Seongho Park, Dongyoon Han, Jaejun Yoo arxiv

Understanding how neural networks rely on visual cues offers a human-interpretable view of their internal decision processes. The cue-conflict benchmark has been influential in probing shape-texture preference and in motivating the insight that stronger, human-like shape bias is often associated with improved in-domain performance. However, we find that the current stylization-based instantiation can yield unstable and ambiguous bias estimates. Specifically, stylization may not reliably instantiate perceptually valid and separable cues nor control their relative informativeness, ratio-based bias can obscure absolute cue sensitivity, and restricting evaluation to preselected classes can distort model predictions by ignoring the full decision space. Together, these factors can confound preference with cue validity, cue balance, and recognizability artifacts. We introduce REFINED-BIAS, an integrated dataset and evaluation framework for reliable and interpretable shape-texture bias diagnosis. REFINED-BIAS constructs balanced, human- and model- recognizable cue pairs using explicit definitions of shape and texture, and measures cue-specific sensitivity over the full label space via a ranking-based metric, enabling fairer cross-model comparisons. Across diverse training regimes and architectures, REFINED-BIAS enables fairer cross-model comparison, more faithful diagnosis of shape and texture biases, and clearer empirical conclusions, resolving inconsistencies that prior cue-conflict evaluations could not reliably disambiguate.

📄 PDF Abstract BibTeX arXiv:2603.10834

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Radial-Angular Geometry for Reliable Update Diagnosis in Noisy-Label Learning

2026-05-17 · Ningkang Peng, Jingyang Mao, Xiaoqian Peng, Weiguang Qu 외 arxiv

Noisy-label methods often estimate sample reliability from forward-space signals such as loss, confidence, or entropy. These signals indicate whether a sample is difficult to predict, but they do not directly test whethe…

Certainly Bot Or Not? Trustworthy Social Bot Detection via Robust Multi-Modal Neural Processes

2025-03-11 · Qi Wu, Yingguang Yang, Hao liu, Hao Peng 외

Social bot detection is crucial for mitigating misinformation, online manipulation, and coordinated inauthentic behavior. While existing neural network-based detectors perform well on benchmarks, they struggle with gener…

Misinformation

Context-specific Credibility-aware Multimodal Fusion with Conditional Probabilistic Circuits

2026-03-27 · Pranuthi Tenali, Sahil Sidheekh, Saurabh Mathur, Erik Blasch 외 arxiv

Multimodal fusion requires integrating information from multiple sources that may conflict depending on context. Existing fusion approaches typically rely on static assumptions about source reliability, limiting their ab…

Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability

2026-06-24 · Yang Tian, Zhengpeng Shi, Yu Zhou, Bo Zhao arxiv

Large language models are increasingly deployed as agents that solve tasks by interacting with external tool environments. Although recent tool-use benchmarks increasingly cover complex task settings, they still largely …

Sensor Data Fusion in Top-View Grid Maps using Evidential Reasoning with Advanced Conflict Resolution

2022-04-19 · Sven Richter, Frank Bieder, Sascha Wirges, Christoph Stiller

We present a new method to combine evidential top-view grid maps estimated based on heterogeneous sensor sources. Dempster's combination rule that is usually applied in this context provides undesired results with highly…