paper-with-me

홈 › Papers

A Multimodal Social Agent

2024-12-11 · Athina Bikaki, Ioannis A. Kakadiaris

In recent years, large language models (LLMs) have demonstrated remarkable progress in common-sense reasoning tasks. This ability is fundamental to understanding social dynamics, interactions, and communication. However, the potential of integrating computers with these social capabilities is still relatively unexplored. However, the potential of integrating computers with these social capabilities is still relatively unexplored. This paper introduces MuSA, a multimodal LLM-based agent that analyzes text-rich social content tailored to address selected human-centric content analysis tasks, such as question answering, visual question answering, title generation, and categorization. It uses planning, reasoning, acting, optimizing, criticizing, and refining strategies to complete a task. Our approach demonstrates that MuSA can automate and improve social content analysis, helping decision-making processes across various applications. We have evaluated our agent's capabilities in question answering, title generation, and content categorization tasks. MuSA performs substantially better than our baselines.

📄 PDF Abstract BibTeX arXiv:2501.06189

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningDecision MakingQuestion AnsweringVisual Question Answering

Similar Papers 제목 키워드 기반

CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

2026-07-29 · Zheng Zhang, Nanjie Yao, Jiarui He, Deheng Ye 외 arxiv

Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents. These games require complex social skills such as reasoning, deception, and collaboration. While recent advances in large lan…

Reinforcement Learning

Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation

2026-06-13 · Shijun Wan, Xuehai Wu, Jiwen Zhang, Siyuan Wang 외 arxiv

Social interaction depends on both language and visible social signals, such as facial expressions, posture, gaze, and emotional shifts. Yet existing social-agent benchmarks are largely text-based and rarely test whether…

A model to generate adaptive multimodal job interviews with a virtual recruiter

2014-05-01 · LREC 2014 5 · Zoraida Callejas, Brian Ravenet, Magalie Ochs, Catherine Pelachaud

This paper presents an adaptive model of multimodal social behavior for embodied conversational agents. The context of this research is the training of youngsters for job interviews in a serious game where the agent play…

Multimodal Safety Evaluation in Generative Agent Social Simulations

2025-10-09 · Alhim Vera, Karen Sanchez, Carlos Hinojosa, Haidar Bin Hamid 외 arxiv

Can generative agents be trusted in multimodal environments? Despite advances in large language and vision-language models that enable agents to act autonomously and pursue goals in rich settings, their ability to reason…

MV-Debate: Multi-view Agent Debate with Dynamic Reflection Gating for Multimodal Harmful Content Detection in Social Media

2025-08-07 · Rui Lu, Jinhe Bi, Yunpu Ma, Feng Xiao 외 arxiv

Social media has evolved into a complex multimodal environment where text, images, and other signals interact to shape nuanced meanings, often concealing harmful intent. Identifying such intent, whether sarcasm, hate spe…

Intent Detection