paper-with-me

홈 › Papers

OlympicArena Medal Ranks: Who Is the Most Intelligent AI So Far?

2024-06-24 · Zhen Huang, Zengzhi Wang, Shijie Xia, PengFei Liu

In this report, we pose the following question: Who is the most intelligent AI model to date, as measured by the OlympicArena (an Olympic-level, multi-discipline, multi-modal benchmark for superintelligent AI)? We specifically focus on the most recently released models: Claude-3.5-Sonnet, Gemini-1.5-Pro, and GPT-4o. For the first time, we propose using an Olympic medal Table approach to rank AI models based on their comprehensive performance across various disciplines. Empirical results reveal: (1) Claude-3.5-Sonnet shows highly competitive overall performance over GPT-4o, even surpassing GPT-4o on a few subjects (i.e., Physics, Chemistry, and Biology). (2) Gemini-1.5-Pro and GPT-4V are ranked consecutively just behind GPT-4o and Claude-3.5-Sonnet, but with a clear performance gap between them. (3) The performance of AI models from the open-source community significantly lags behind these proprietary models. (4) The performance of these models on this benchmark has been less than satisfactory, indicating that we still have a long way to go before achieving superintelligence. We remain committed to continuously tracking and evaluating the performance of the latest powerful models on this benchmark (available at https://github.com/GAIR-NLP/OlympicArena).

📄 PDF Abstract BibTeX arXiv:2406.16772

Code (1)

gair-nlp/olympicarena 공식 구현

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

2024-06-18 · Zhen Huang, Zengzhi Wang, Shijie Xia, Xuefeng Li 외

The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showcasing potential cognitive reasoning abil…

Benchmarkingscientific discovery

Forecasting the Olympic medal distribution during a pandemic: a socio-economic machine learning model

2020-12-08 · Christoph Schlembach, Sascha L. Schmidt, Dominik Schreyer, Linus Wunderlich

Forecasting the number of Olympic medals for each nation is highly relevant for different stakeholders: Ex ante, sports betting companies can determine the odds while sponsors and media companies can allocate their resou…

BIG-bench Machine Learning

Neural Feature Learning From Relational Database

2018-01-16 · Hoang Thanh Lam, Tran Ngoc Minh, Mathieu Sinn, Beat Buesser 외

Feature engineering is one of the most important but most tedious tasks in data science. This work studies automation of feature learning from relational database. We first prove theoretically that finding the optimal fe…

Feature Engineering

Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025

2025-09-01 · Jiahao Qiu, Jingzhe Shi, Xinzhe Juan, Zelin Zhao 외 arxiv

Physics provides fundamental laws that describe and predict the natural world. AI systems aspiring toward more general, real-world intelligence must therefore demonstrate strong physics problem-solving abilities: to form…

Medal S: Spatio-Textual Prompt Model for Medical Segmentation

2025-11-17 · Pengcheng Shi, Jiawei Chen, Jiaqi Liu, Xinglin Zhang 외 arxiv

We introduce Medal S, a medical segmentation foundation model that supports native-resolution spatial and textual prompts within an end-to-end trainable framework. Unlike text-only methods lacking spatial awareness, Meda…

Data Augmentation