paper-with-me

Papers

Beyond Text: Unveiling Multimodal Proficiency of Large Language Models with MultiAPI Benchmark

2023-11-21 · Xiao Liu, Jianfeng Lin, Jiawei Zhang

The proliferation of Large Language Models like ChatGPT has significantly advanced language understanding and generation, impacting a broad spectrum of applications. However, these models predominantly excel in text-based tasks, overlooking the complexity of real-world multimodal information. This study introduces MultiAPI, a pioneering comprehensive large-scale API benchmark dataset aimed at expanding LLMs' proficiency in multimodal contexts. Developed collaboratively through ChatGPT, MultiAPI consists of 235 diverse API calls and 2,038 contextual prompts, offering a unique platform evaluation of tool-augmented LLMs handling multimodal tasks. Through comprehensive experiments, our findings reveal that while LLMs demonstrate proficiency in API call decision-making, they face challenges in domain identification, function selection, and argument generation. What's more, we surprisingly notice that auxiliary context can actually impair the performance. An in-depth error analysis paves the way for a new paradigm to address these challenges, suggesting a potential direction for future LLM research.

📄 PDF Abstract BibTeX arXiv:2311.13053

Code (1)

haroldliuj/multiapi 공식 구현 pytorch

Tasks

Decision Making

Similar Papers 제목 키워드 기반

Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation

2025-05-20 · Jiankun Zhang, Shenglai Zeng, Jie Ren, Tianqi Zheng 외

Multimodal Retrieval-Augmented Generation (MRAG) systems enhance LMMs by integrating external multimodal databases, but introduce unexplored privacy vulnerabilities. While text-based RAG privacy risks have been studied, …

Privacy PreservingRAGRetrievalRetrieval-augmented Generation

MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models

2025-06-05 · Gio Paik, Geewook Kim, Jinbae Im

This paper introduces MMRefine, a MultiModal Refinement benchmark designed to evaluate the error refinement capabilities of Multimodal Large Language Models (MLLMs). As the emphasis shifts toward enhancing reasoning duri…

Unveiling the Competitive Dynamics: A Comparative Evaluation of American and Chinese LLMs

2024-05-09 · Zhenhui Jiang, Jiaxin Li, Yang Liu

The strategic significance of Large Language Models (LLMs) in economic expansion, innovation, societal development, and national security has been increasingly recognized since the advent of ChatGPT. This study provides …

Unveiling and Manipulating Prompt Influence in Large Language Models

2024-05-20 · Zijian Feng, Hanzhang Zhou, Zixiao Zhu, Junlang Qian 외

Prompts play a crucial role in guiding the responses of Large Language Models (LLMs). However, the intricate role of individual tokens in prompts, known as input saliency, in shaping the responses remains largely underex…

Language ModellingText Generation

Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models

2025-12-17 · Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 외 arxiv

Multimodal Large Language Models (MLLMs) have recently demonstrated impressive capabilities in connecting vision and language, yet their proficiency in fundamental visual reasoning tasks remains limited. This limitation …

Multimodal ReasoningVisual Reasoning