paper-with-me

홈 › Papers

Moving Beyond More Views: Redundancy-Aware Ego-Exo Fusion for Proficiency Estimation

2026-08-26 · Xu Dong, Wanqing Li, Anthony Adeyemi-Ejeye, Andrew Gilbert arxiv

EgoExo proficiency estimation aims to assess action quality by integrating fine-grained motion cues from egocentric (1st-person) views with spatial context from multiple exocentric (3rd-person) views. Simply adding more exocentric views degrades EgoExo performance, as redundant or noisy perspectives dilute useful motion cues. Our analysis identifies two key causes: (1) Multiview redundancy - From the data perspective, certain views provide limited or noisy information, diluting discriminative cues; (2) Overfitting - From the feature perspective, conventional fusion increases representational complexity, causing the model to memorise view-specific patterns rather than learn generalisable representations. To address these issues, we propose two complementary modules: AdaMVS, which adaptively identifies and fuses the most informative view tokens under weak supervision from the data perspective, and VIB-GB, which combines Gradient Blending and Variational Information Bottleneck regularisation from the feature perspective to compress redundant signals and suppress overfitting during training. Experiments on EgoExo-4D and EgoExo-Fitness demonstrate that our method learns both which view to look at and how to fuse them, achieving new state-of-the-art results. Our source code is available at https://github.com/dx199771/AdaMVS

📄 PDF Abstract BibTeX arXiv:2608.25736

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Sentiment: A Multi-Agent Pipeline for Actionable Business Advice from Reviews

2026-01-17 · Kartikey Singh Bhandari, Tanish Jain, Archit Agrawal, Dhruv Kumar 외 arxiv

Customer reviews contain valuable signals about service quality, but converting large-scale review corpora into actionable business recommendations remains difficult. Standard sentiment/aspect analysis is largely descrip…

A Spatial RNN Codec for End-to-End Image Compression

2020-06-01 · CVPR 2020 6 · Chaoyi Lin, Jiabao Yao, Fangdong Chen, Li Wang

Recently, deep learning has been explored as a promising direction for image compression. Removing the spatial redundancy of the image is crucial for image compression and most learning based methods focus on removing th…

Image CompressionMS-SSIMSSIM

Think Clearly: Improving Reasoning via Redundant Token Pruning

2025-06-17 · Daewon Choi, JiMin Lee, Jihoon Tack, Woomin Song 외

Recent large language models have shown promising capabilities in long-form reasoning, following structured chains of thought before arriving at a final answer. However, we observe that these reasoning paths tend to incl…

RAL:Redundancy-Aware Lipreading Model Based on Differential Learning with Symmetric Views

2024-09-09 · Zejun Gu, Junxia jiang

Lip reading involves interpreting a speaker's speech by analyzing sequences of lip movements. Currently, most models regard the left and right halves of the lips as a symmetrical whole, lacking a thorough investigation o…

LipreadingLip Reading

PixelGaussian: Generalizable 3D Gaussian Reconstruction from Arbitrary Views

2024-10-24 · Xin Fei, Wenzhao Zheng, Yueqi Duan, Wei Zhan 외

We propose PixelGaussian, an efficient feed-forward framework for learning generalizable 3D Gaussian reconstruction from arbitrary views. Most existing methods rely on uniform pixel-wise Gaussian representations, which l…