paper-with-me

홈 › Papers

Supporting Experts with a Multimodal Machine-Learning-Based Tool for Human Behavior Analysis of Conversational Videos

2024-02-17 · Riku Arakawa, Kiyosu Maeda, Hiromu Yakura

Multimodal scene search of conversations is essential for unlocking valuable insights into social dynamics and enhancing our communication. While experts in conversational analysis have their own knowledge and skills to find key scenes, a lack of comprehensive, user-friendly tools that streamline the processing of diverse multimodal queries impedes efficiency and objectivity. To solve it, we developed Providence, a visual-programming-based tool based on design considerations derived from a formative study with experts. It enables experts to combine various machine learning algorithms to capture human behavioral cues without writing code. Our study showed its preferable usability and satisfactory output with less cognitive load imposed in accomplishing scene search tasks of conversations, verifying the importance of its customizability and transparency. Furthermore, through the in-the-wild trial, we confirmed the objectivity and reusability of the tool transform experts' workflow, suggesting the advantage of expert-AI teaming in a highly human-contextual domain.

📄 PDF Abstract BibTeX arXiv:2402.11145

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MM-DeepResearch: A Simple and Effective Multimodal Agentic Search Baseline

2026-03-01 · Huanjin Yao, Qixiang Yin, Min Yang, Ziwang Zhao 외 arxiv

We aim to develop a multimodal research agent capable of explicit reasoning and planning, multi-tool invocation, and cross-modal information synthesis, enabling it to conduct deep research tasks. However, we observe thre…

Reinforcement Learning

VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool

2024-07-07 · Yan Wang, Yawen Zeng, Jingsheng Zheng, Xiaofen Xing 외

Multimodal large language models (MLLMs) are flourishing, but mainly focus on images with less attention than videos, especially in sub-fields such as prompt engineering, video chain-of-thought (CoT), and instruction tun…

Active LearningHallucinationPrompt Engineering

SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification

2025-06-18 · Chengye Wang, Yifei Shen, Zexi Kuang, Arman Cohan 외

We introduce SciVer, the first benchmark specifically designed to evaluate the ability of foundation models to verify claims within a multimodal scientific context. SciVer consists of 3,000 expert-annotated examples over…

Claim VerificationRAGRetrieval-augmented Generation

CollabKG: A Learnable Human-Machine-Cooperative Information Extraction Toolkit for (Event) Knowledge Graph Construction

2023-07-03 · Xiang Wei, Yufeng Chen, Ning Cheng, Xingyu Cui 외

In order to construct or extend entity-centric and event-centric knowledge graphs (KG and EKG), the information extraction (IE) annotation toolkit is essential. However, existing IE toolkits have several non-trivial prob…

Event Extractiongraph constructionKnowledge Graphsnamed-entity-recognition+3

Understanding Human Judgments of Causality

2019-12-19 · Masahiro Kazama, Yoshihiko Suhara, Andrey Bogomolov, Alex `Sandy' Pentland

Discriminating between causality and correlation is a major problem in machine learning, and theoretical tools for determining causality are still being developed. However, people commonly make causality judgments and ar…

AttributeBIG-bench Machine Learning