paper-with-me

Papers

Beyond Specialization: Assessing the Capabilities of MLLMs in Age and Gender Estimation

2024-03-04 · Maksim Kuprashevich, Grigorii Alekseenko, Irina Tolstykh

Multimodal Large Language Models (MLLMs) have recently gained immense popularity. Powerful commercial models like ChatGPT-4V and Gemini, as well as open-source ones such as LLaVA, are essentially general-purpose models and are applied to solve a wide variety of tasks, including those in computer vision. These neural networks possess such strong general knowledge and reasoning abilities that they have proven capable of working even on tasks for which they were not specifically trained. We compared the capabilities of the most powerful MLLMs to date: ShareGPT4V, ChatGPT, LLaVA-Next in a specialized task of age and gender estimation with our state-of-the-art specialized model, MiVOLO. We also updated MiVOLO and provide details and new metrics in this article. This comparison has yielded some interesting results and insights about the strengths and weaknesses of the participating models. Furthermore, we attempted various ways to fine-tune the ShareGPT4V model for this specific task, aiming to achieve state-of-the-art results in this particular challenge. Although such a model would not be practical in production, as it is incredibly expensive compared to a specialized model like MiVOLO, it could be very useful in some tasks, like data annotation.

📄 PDF Abstract BibTeX arXiv:2403.02302

Code (1)

wildchlamydia/mivolo 공식 구현 pytorch

Tasks

Age And Gender ClassificationAge and Gender EstimationAge EstimationFacial Attribute ClassificationGender PredictionGeneral Knowledge

Similar Papers 제목 키워드 기반

CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?

2025-02-16 · Aashish Anantha Ramakrishnan, Aadarsh Anantha Ramakrishnan, Dongwon Lee

Multimodal Large Language Models (MLLMs) are renowned for their superior instruction-following and reasoning capabilities across diverse problem domains. However, existing benchmarks primarily focus on assessing factual …

Instruction Following

From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities

2024-01-26 · Chaochao Lu, Chen Qian, Guodong Zheng, Hongxing Fan 외

Multi-modal Large Language Models (MLLMs) have shown impressive abilities in generating reasonable responses with respect to multi-modal contents. However, there is still a wide gap between the performance of recent MLLM…

Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations

2026-02-01 · Sheng-Lun Wei, Yu-Ling Liao, Yen-Hua Chang, Hen-Hsen Huang 외 arxiv

This work presents the first systematic investigation of speech bias in multilingual MLLMs. We construct and release the BiasInEar dataset, a speech-augmented benchmark based on Global MMLU Lite, spanning English, Chines…

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

2025-10-12 · Caorui Li, Yu Chen, Yiyan Ji, Jin Xu 외 arxiv

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities…

Causal InferenceVisual Reasoning

H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding

2025-03-31 · Qi Wu, Quanlong Zheng, Yanhao Zhang, Junlin Xie 외

With the rapid development of multimodal models, the demand for assessing video understanding capabilities has been steadily increasing. However, existing benchmarks for evaluating video understanding exhibit significant…

Video Understanding