paper-with-me

MMR total 벤치마크

MMR total on MRR-Benchmark

14개 결과 · ⬇ CSV · JSON

Total Column Score

139 220 301 382 463 2023-04 2026-09 LLaVA-NEXT-34B — 412.0 (2023-04-17) LLaVA-NEXT-13B — 335.0 (2023-04-17) LLaVA-1.5-13B — 243.0 (2023-04-17) Idefics-80B — 139.0 (2023-06-21) Qwen-vl-max — 366.0 (2023-08-24) Qwen-vl-plus — 310.0 (2023-08-24) GPT-4V — 415.0 (2023-09-29) Monkey-Chat-7B — 214.0 (2023-11-11) InternVL2-8B — 368.0 (2023-12-21) InternVL2-1B — 237.0 (2023-12-21) Phi-3-Vision — 397.0 (2024-04-22) Idefics-2-8B — 256.0 (2024-05-03) GPT-4o — 457.0 (2024-06-14) Claude 3.5 Sonnet — 463.0 (2024-06-24) LLaVA-NEXT-34B — 412.0 (2023-04-17) GPT-4V — 415.0 (2023-09-29) GPT-4o — 457.0 (2024-06-14) Claude 3.5 Sonnet — 463.0 (2024-06-24)
RankModel Total Column Score Extra Training Data PaperCodeYear
1 Claude 3.5 Sonnet 463 Claude 3.5 Sonnet Model Card Addendum 2024
2 GPT-4o 457 GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding 2024
3 GPT-4V 415 The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) qi-zhangyang/gemini-vs-gpt4v · vista-h/gpt-4v_social_media 2023
4 LLaVA-NEXT-34B 412 Visual Instruction Tuning huggingface/transformers · haotian-liu/LLaVA · LLaVA-VL/LLaVA-NeXT · +10 2023
5 Phi-3-Vision 397 Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone 2024
6 InternVL2-8B 368 InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks opengvlab/internvl · opengvlab/internvl-mmdetseg 2023
7 Qwen-vl-max 366 Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond qwenlm/qwen-vl · brandon3964/multimodal-task-vector 2023
8 LLaVA-NEXT-13B 335 Visual Instruction Tuning huggingface/transformers · haotian-liu/LLaVA · LLaVA-VL/LLaVA-NeXT · +10 2023
9 Qwen-vl-plus 310 Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond qwenlm/qwen-vl · brandon3964/multimodal-task-vector 2023
10 Idefics-2-8B 256 What matters when building vision-language models? 2024
11 LLaVA-1.5-13B 243 Visual Instruction Tuning huggingface/transformers · haotian-liu/LLaVA · LLaVA-VL/LLaVA-NeXT · +10 2023
12 InternVL2-1B 237 InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks opengvlab/internvl · opengvlab/internvl-mmdetseg 2023
13 Monkey-Chat-7B 214 Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models yuliang-liu/monkey 2023
14 Idefics-80B 139 OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents huggingface/obelics · MindSpore-scientific-2/code-14 2023
1–14 / 14 페이지당 10 20 50 100