paper-with-me

홈 › Papers

FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions

2025-09-21 · Bowen Qin, Chen Yue, Fang Yin, Hui Wang, JG Yao, Jiakang Liu, Jing-Shu Zheng, Miguel Hu Chen, Richeng Xuan, Shibei Meng, Shiqi Zhou, Teng Dai, Tong-Shuai Ren, Wei Cui, Xi Yang, Xialin Du, Xiaojing Xu, Xue Sun, Xuejing Li, Yaming Liu, Yesheng Liu, Ying Liu, Yonghua Lin, Yu Zhao, Yunduo Zhang, Yuwen Luo, Zheqi He, Zhiyuan He, Zhongyuan Wang arxiv

We conduct a moderate-scale contamination-free (to some extent) evaluation of current large reasoning models (LRMs) with some preliminary findings. We also release ROME, our evaluation benchmark for vision language models intended to test reasoning from visual clues. We attach links to the benchmark, evaluation data, and other updates on this website: https://flageval-baai.github.io/LRM-Eval/

📄 PDF Abstract BibTeX arXiv:2509.17177

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation

2025-06-10 · Zheqi He, Yesheng Liu, Jing-shu Zheng, Xuejing Li 외

We present FlagEvalMM, an open-source evaluation framework designed to comprehensively assess multimodal models across a diverse range of vision-language understanding and generation tasks, such as visual question answer…

Image-text RetrievalQuestion AnsweringText RetrievalVideo Generation+1

Preliminary WMT24 Ranking of General MT Systems and LLMs

2024-07-29 · Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondrej Bojar 외

This is the preliminary ranking of WMT24 General MT systems based on automatic metrics. The official ranking will be a human evaluation, which is superior to the automatic ranking and supersedes it. The purpose of this r…

Chest X-ray Report Generation through Fine-Grained Label Learning

2020-07-27 · Tanveer Syeda-Mahmood, Ken C. L. Wong, Yaniv Gur, Joy T. Wu 외

Obtaining automated preliminary read reports for common exams such as chest X-rays will expedite clinical workflows and improve operational efficiencies in hospitals. However, the quality of reports generated by current …

From Prompt Optimization to Multi-Dimensional Credibility Evaluation: Enhancing Trustworthiness of Chinese LLM-Generated Liver MRI Reports -- with Preliminary Extension to Lung Cancer

2025-10-27 · Qiuli Wang, Xinhuang Sun, Yonglin Chen, Jie Cheng 외 arxiv

Large language models (LLMs) have demonstrated promising performance in generating diagnostic conclusions from imaging findings, thereby supporting radiology reporting, trainee education, and quality control. However, sy…

Is ChatGPT A Good Keyphrase Generator? A Preliminary Study

2023-03-23 · Mingyang Song, Haiyun Jiang, Shuming Shi, Songfang Yao 외

The emergence of ChatGPT has recently garnered significant attention from the computational linguistics community. To demonstrate its capabilities as a keyphrase generator, we conduct a preliminary evaluation of ChatGPT …

Diversitydocument understandingKeyphrase Generation