paper-with-me

Papers

CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views

2026-07-07 · Alexey Gavryushin, Dingxi Zhang, Zhao Huang, Alexandros Delitzas, Jiaqi Chen, Ben Ellis, Cedric Zöllner, Manthan Patel, Manuel Kaufmann, Marc Pollefeys, Xi Wang arxiv

Human-human collaboration is a fundamental aspect of everyday life, essential to success in a wide range of goal-directed activities from household tasks to professional teamwork. While much research has focused on modeling coordination and task execution, the cognitive processes that support such collaboration, particularly Theory of Mind (the ability to infer the mental states of others), remain difficult to study in natural settings. To address this gap, we introduce a novel egocentric and exocentric video dataset capturing real-world collaboration in cooking scenarios. The dataset integrates multi-perspective video, high-quality audio, gaze tracking, and 3D scene and object scans, with annotations for shared attention to objects, social cues and interactions between agents, as well as agent-object interactions. We establish benchmarks for Joint Attention Estimation, Socially Conditioned Object Interaction Anticipation, and Collaborative Handover Prediction, enabling research on multimodal perception, proactive assistance, and collaborative planning. By providing temporally aligned, richly annotated multimodal data, CoMind facilitates the development and evaluation of AI systems capable of modeling complex social interactions and reasoning about human behaviors in collaborative environments. Our dataset and benchmarks are made available at https://comind.ethz.ch/.

📄 PDF Abstract BibTeX arXiv:2607.06691

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Community-Driven Agents for Machine Learning Engineering

2025-06-25 · Sijie Li, Weiwei Sun, Shanda Li, Ameet Talwalkar 외

Large language model-based machine learning (ML) agents have shown great promise in automating ML research. However, existing agents typically operate in isolation on a given research problem, without engaging with the b…

Language ModelingLanguage ModellingLarge Language Model

RecoMind: A Reinforcement Learning Framework for Optimizing In-Session User Satisfaction in Recommendation Systems

2025-07-31 · Mehdi Ben Ayed, Fei Feng, Jay Adams, Vishwakarma Singh 외 arxiv

Existing web-scale recommendation systems commonly use supervised learning methods that prioritize immediate user feedback. Although reinforcement learning (RL) offers a solution to optimize longer-term goals, such as in…

Reinforcement LearningRecommendation Systems

A Nonparametric Model for Multimodal Collaborative Activities Summarization

2017-09-04 · Guy Rosman, John W. Fisher III, Daniela Rus

Ego-centric data streams provide a unique opportunity to reason about joint behavior by pooling data across individuals. This is especially evident in urban environments teeming with human activities, but which suffer fr…

Action DetectionActivity Detection

TTF: Temporal Token Fusion for Efficient Video-Language Model

2026-05-08 · Simin Huo, Ning LI arxiv

Video-language models (VLMs) face rapid inference costs as visual token counts scale with video length. For example, 32 frames at $448{\times}448$ resolution already yield >8,000 visual tokens in Qwen3-VL, making LLM pre…

Periodic RoPE for Infinite Context LLMs

2026-05-27 · Simin Huo arxiv

The ability to process ultra-long contexts is crucial for large language models (LLMs) to perform long-horizon tasks. While recent efforts have extended context windows to 1M and beyond, model performance degrades when s…