paper-with-me

홈 › Papers

TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation

2025-08-18 · Keunwoo Choi, Seungheon Doh, Juhan Nam arxiv

We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline. In the proposed pipeline, multiple large language model (LLM) agents are created under various roles with specialized prompts and access to different parts of information, and the chat data is acquired by logging the conversation between the Listener LLM and the Recsys LLM. To cover various conversation scenarios, for each conversation, the Listener LLM is conditioned on a finetuned conversation goal. Finally, all the LLMs are multimodal with audio and images, allowing a simulation of multimodal recommendation and conversation. In the LLM-as-a-judge and subjective evaluation experiments, TalkPlayData 2 achieved the proposed goal in various aspects related to training a generative recommendation model for music. TalkPlayData 2 and its generation code are released at https://talkpl-ai.github.io.

📄 PDF Abstract BibTeX arXiv:2509.09685

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Recommendation

Similar Papers 제목 키워드 기반

Agentic AI Microservice Framework for Deepfake and Document Fraud Detection in KYC Pipelines

2026-01-09 · Chandra Sekhar Kubam arxiv

The rapid proliferation of synthetic media, presentation attacks, and document forgeries has created significant vulnerabilities in Know Your Customer (KYC) workflows across financial services, telecommunications, and di…

DeepFake DetectionFraud Detection

InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search

2026-05-08 · Bohan Hou, Jiuning Gu, Jiayan Guo, Ronghao Dang 외 arxiv

Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input or treated as an answer endpoint rather than part of an interleaved se…

Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation

2026-04-12 · Fangda Ye, Zhifei Xie, Yuxin Hu, Yihang Yin 외 arxiv

Recent agentic search frameworks enable deep research via iterative planning and retrieval, reducing hallucinations and enhancing factual grounding. However, they remain text-centric, overlooking the multimodal evidence …

multimodal generation

Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis

2026-03-31 · Shuang Chen, Quanxin Shou, Hangting Chen, Yucheng Zhou 외 arxiv

Unified multimodal models provide a natural and promising architecture for understanding diverse and complex real-world knowledge while generating high-quality images. However, they still rely primarily on frozen paramet…

Image Generation

Intern-S2-Preview: Scientific Agentic Foundation Model

2026-08-13 · Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen 외 arxiv

Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons.…

Reinforcement Learning