paper-with-me

홈 › Papers

Preliminary WMT24 Ranking of General MT Systems and LLMs

2024-07-29 · Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondrej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Kenton Murray, Masaaki Nagata, Martin Popel, Maja Popovic, Mariya Shmatova, Steinþór Steingrímsson, Vilém Zouhar

This is the preliminary ranking of WMT24 General MT systems based on automatic metrics. The official ranking will be a human evaluation, which is superior to the automatic ranking and supersedes it. The purpose of this report is not to interpret any findings but only provide preliminary results to the participants of the General MT task that may be useful during the writing of the system submission.

📄 PDF Abstract BibTeX arXiv:2407.19884

Code (1)

wmt-conference/wmt-collect-translations 공식 구현

Similar Papers 제목 키워드 기반

Preliminary Ranking of WMT25 General Machine Translation Systems

2025-08-11 · Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar 외 arxiv

We present the preliminary rankings of machine translation (MT) systems submitted to the WMT25 General Machine Translation Shared Task, as determined by automatic evaluation metrics. Because these rankings are derived fr…

Machine Translation

The Ranking Blind Spot: Decision Hijacking in LLM-based Text Ranking

2025-09-23 · Yaoyao Qian, Yifan Zeng, Yuchao Jiang, Chelsi Jain 외 arxiv

Large Language Models (LLMs) have demonstrated strong performance in information retrieval tasks like passage ranking. Our research examines how instruction-following capabilities in LLMs interact with multi-document com…

Information RetrievalPassage Ranking

Prompting for Performance: Exploring LLMs for Configuring Software

2025-07-13 · Helge Spieker, Théo Matricon, Nassim Belmecheri, Jørn Eirik Betten 외

Software systems usually provide numerous configuration options that can affect performance metrics such as execution time, memory usage, binary size, or bitrate. On the one hand, making informed decisions is challenging…

Large Language Models are Zero-Shot Rankers for Recommender Systems

2023-05-15 · Yupeng Hou, Junjie Zhang, Zihan Lin, Hongyu Lu 외

Recently, large language models (LLMs) (e.g., GPT-4) have demonstrated impressive general-purpose task-solving abilities, including the potential to approach recommendation tasks. Along this line of research, this work a…

Recommendation Systems

CRS Arena: Crowdsourced Benchmarking of Conversational Recommender Systems

2024-12-13 · Nolwenn Bernard, Hideaki Joko, Faegheh Hasibi, Krisztian Balog

We introduce CRS Arena, a research platform for scalable benchmarking of Conversational Recommender Systems (CRS) based on human feedback. The platform displays pairwise battles between anonymous conversational recommend…

BenchmarkingRecommendation Systems