paper-with-me

홈 › Papers

Accounting for Language Effect in the Evaluation of Cross-lingual AMR Parsers

2022-10-01 · COLING 2022 10 · Shira Wein, Nathan Schneider

Cross-lingual Abstract Meaning Representation (AMR) parsers are currently evaluated in comparison to gold English AMRs, despite parsing a language other than English, due to the lack of multilingual AMR evaluation metrics. This evaluation practice is problematic because of the established effect of source language on AMR structure. In this work, we present three multilingual adaptations of monolingual AMR evaluation metrics and compare the performance of these metrics to sentence-level human judgments. We then use our most highly correlated metric to evaluate the output of state-of-the-art cross-lingual AMR parsers, finding that Smatch may still be a useful metric in comparison to gold English AMRs, while our multilingual adaptation of S2match (XS2match) is best for comparison with gold in-language AMRs.

📄 PDF Abstract BibTeX

Code (1)

shirawein/crossling-amr-eval 공식 구현 tf

Tasks

Abstract Meaning RepresentationSentence

Similar Papers 제목 키워드 기반

MUST-VQA: MUltilingual Scene-text VQA

2022-09-14 · Emanuele Vivoli, Ali Furkan Biten, Andres Mafla, Dimosthenis Karatzas 외

In this paper, we present a framework for Multilingual Scene Text Visual Question Answering that deals with new languages in a zero-shot fashion. Specifically, we consider the task of Scene Text Visual Question Answering…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety

2026-07-11 · Chigozirim Ifebi, Brent Kong, Ayushi Mehrotra arxiv

Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful compliance in non-English and low-resource settings. We introduce \textsc{Minionese}, a m…

Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs

2024-10-17 · Sumanth Doddapaneni, Mohammed Safi Ur Rahman Khan, Dilip Venkatesh, Raj Dabre 외

Evaluating machine-generated text remains a significant challenge in NLP, especially for non-English languages. Current methodologies, including automated metrics, human assessments, and LLM-based evaluations, predominan…

Benchmarking

Accounting Reasoning in Large Language Models: Concepts, Evaluation, and Empirical Analysis

2025-12-27 · Jie Zhou, Xin Chen, Jie Zhang, Zhe Li arxiv

Large language models (LLMs) are increasingly reshaping learning paradigms, cognitive processes, and research methodologies across diverse domains. As their adoption expands, effectively integrating LLMs into professiona…

Prompt Engineering

Disentangling Speaker and Language Effects in Cross-Lingual Speaker Verification for Iberian Languages

2026-07-01 · Pol Buitrago, Javier Hernando arxiv

Cross-lingual speaker verification (SV) systems typically exhibit performance degradation when enrollment and test utterances are spoken in different languages. However, standard evaluation protocols confound language mi…

Cross-Lingual TransferSpeaker Verification