paper-with-me

Papers

PCIE_Interaction Solution for Ego4D Social Interaction Challenge

2025-05-30 · Kanokphan Lertniphonphan, Feng Chen, Junda Xu, Fengbu Lan, Jun Xie, Tao Zhang, Zhepeng Wang

This report presents our team's PCIE_Interaction solution for the Ego4D Social Interaction Challenge at CVPR 2025, addressing both Looking At Me (LAM) and Talking To Me (TTM) tasks. The challenge requires accurate detection of social interactions between subjects and the camera wearer, with LAM relying exclusively on face crop sequences and TTM combining speaker face crops with synchronized audio segments. In the LAM track, we employ face quality enhancement and ensemble methods. For the TTM task, we extend visual interaction analysis by fusing audio and visual cues, weighted by a visual quality score. Our approach achieved 0.81 and 0.71 mean average precision (mAP) on the LAM and TTM challenges leader board. Code is available at https://github.com/KanokphanL/PCIE_Ego4D_Social_Interaction

📄 PDF Abstract BibTeX arXiv:2505.24404

Code (1)

kanokphanl/pcie_ego4d_social_interaction 공식 구현

Similar Papers 제목 키워드 기반

Project Tracyn: Generative Artificial Intelligence based Peripherals Trace Synthesizer

2024-11-10 · Zhibai Huang, Yihan Shen, Yongchen Xie, Zhixiang Wei 외

Peripheral Component Interconnect Express (PCIe) is the de facto interconnect standard for high-speed peripherals and CPUs. Prototyping and optimizing PCIe devices for emerging scenarios is an ongoing challenge. Since Tr…

CPU

PCIE_LAM Solution for Ego4D Looking At Me Challenge

2024-06-18 · Kanokphan Lertniphonphan, Jun Xie, Yaqing Meng, Shijing Wang 외

This report presents our team's 'PCIE_LAM' solution for the Ego4D Looking At Me Challenge at CVPR2024. The main goal of the challenge is to accurately determine if a person in the scene is looking at the camera wearer, b…

PCIE_EgoHandPose Solution for EgoExo4D Hand Pose Challenge

2024-06-18 · Feng Chen, Ling Ding, Kanokphan Lertniphonphan, Jian Li 외

This report presents our team's 'PCIE_EgoHandPose' solution for the EgoExo4D Hand Pose Challenge at CVPR2024. The main goal of the challenge is to accurately estimate hand poses, which involve 21 3D joints, using an RGB …

Position

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers

2025-09-28 · Enda Yu, Dezun Dong, Zhaoning Zhang, Zhe Bai 외 arxiv

Mixture-of-Experts (MoE) models face memory and PCIe latency bottlenecks when deployed on commodity hardware. Offloading expert weights to CPU memory results in PCIe transfer latency that exceeds GPU computation by sever…

Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations

2024-03-04 · CVPR 2024 1 · Sangmin Lee, Bolin Lai, Fiona Ryan, Bikram Boote 외

Understanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However, most prior works on multimodal social cues focus predominantly on single-p…

coreference-resolutionCoreference Resolution