paper-with-me

홈 › Papers

Learning Multimodal Cues of Children's Uncertainty

2024-10-17 · Qi Cheng, Mert İnan, Rahma Mbarki, Grace Grmek, Theresa Choi, Yiming Sun, Kimele Persaud, Jenny Wang, Malihe Alikhani

Understanding uncertainty plays a critical role in achieving common ground (Clark et al.,1983). This is especially important for multimodal AI systems that collaborate with users to solve a problem or guide the user through a challenging concept. In this work, for the first time, we present a dataset annotated in collaboration with developmental and cognitive psychologists for the purpose of studying nonverbal cues of uncertainty. We then present an analysis of the data, studying different roles of uncertainty and its relationship with task difficulty and performance. Lastly, we present a multimodal machine learning model that can predict uncertainty given a real-time video clip of a participant, which we find improves upon a baseline multimodal transformer model. This work informs research on cognitive coordination between human-human and human-AI and has broad implications for gesture understanding and generation. The anonymized version of our data and code will be publicly available upon the completion of the required consent forms and data sheets.

📄 PDF Abstract BibTeX arXiv:2410.14050

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Enhanced Multimodal Content Moderation of Children's Videos using Audiovisual Fusion

2024-05-09 · Syed Hammad Ahmed, Muhammad Junaid Khan, Gita Sukthankar

Due to the rise in video content creation targeted towards children, there is a need for robust content moderation schemes for video hosting platforms. A video that is visually benign may include audio content that is in…

Prompt LearningRobust classification

MMASD: A Multimodal Dataset for Autism Intervention Analysis

2023-06-14 · Jicheng Li, Vuthea Chheang, Pinar Kullu, Eli Brignac 외

Autism spectrum disorder (ASD) is a developmental disorder characterized by significant social communication impairments and difficulties perceiving and presenting communication cues. Machine learning techniques have bee…

Action Quality AssessmentOptical Flow EstimationPrivacy Preserving

Deep Bayesian Network for Visual Question Generation

2020-01-23 · Badri N. Patro, Vinod K. Kurmi, Sandeep Kumar, Vinay P. Namboodiri

Generating natural questions from an image is a semantic task that requires using vision and language modalities to learn multimodal representations. Images can have multiple visual and language cues such as places, capt…

Natural QuestionsQuestion GenerationQuestion-Generation

MEWL: Few-shot multimodal word learning with referential uncertainty

2023-06-01 · Guangyuan Jiang, Manjie Xu, Shiji Xin, Wei Liang 외

Without explicit feedback, humans can rapidly learn the meaning of words. Children can acquire a new word after just a few passive exposures, a process known as fast mapping. This word learning capability is believed to …

GPT-3-driven pedagogical agents for training children's curious question-asking skills

2022-11-25 · Rania Abdelghani, Yen-Hsiang Wang, Xingdi Yuan, Tong Wang 외

In order to train children's ability to ask curiosity-driven questions, previous research has explored designing specific exercises relying on providing semantic and linguistic cues to help formulate such questions. But …

Language ModellingLarge Language Model