paper-with-me

홈 › Papers

Beyond Agreement: Diagnosing the Rationale Alignment of Automated Essay Scoring Methods based on Linguistically-informed Counterfactuals

2024-05-29 · Yupei Wang, Renfen Hu, Zhe Zhao

While current Automated Essay Scoring (AES) methods demonstrate high scoring agreement with human raters, their decision-making mechanisms are not fully understood. Our proposed method, using counterfactual intervention assisted by Large Language Models (LLMs), reveals that BERT-like models primarily focus on sentence-level features, whereas LLMs such as GPT-3.5, GPT-4 and Llama-3 are sensitive to conventions & accuracy, language complexity, and organization, indicating a more comprehensive rationale alignment with scoring rubrics. Moreover, LLMs can discern counterfactual interventions when giving feedback on essays. Our approach improves understanding of neural AES methods and can also apply to other domains seeking transparency in model-driven decisions.

📄 PDF Abstract BibTeX arXiv:2405.19433

Code (1)

yplarrywang/beyond-agreement-aes-2024 공식 구현

Tasks

Automated Essay ScoringcounterfactualDecision MakingSentence

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Weight Decay 설명 없음
Position-Wise Feed-Forward Layer 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…

Similar Papers 제목 키워드 기반

A Finetuned SpeechLLM for Joint Multi-Granular L2 Assessment and Natural-Language Rationales

2026-06-08 · Aditya Kamlesh Parikh, Cristian Tejedor-Garcia, Catia Cucchiarini, Helmer Strik arxiv

Automated L2 speech assessment can assign proficiency labels, but often lacks interpretability. We propose a rubric-guided SpeechLLM for multi-aspect, multi-granular assessment, trained with a hybrid objective combining …

Persona Prompting as a Lens on LLM Social Reasoning

2026-01-28 · Jing Yang, Moritz Hechtbauer, Elisabeth Khalilov, Evelyn Luise Brinkmann 외 arxiv

For socially sensitive tasks like hate speech detection, the quality of explanations from Large Language Models (LLMs) is crucial for factors like user trust and model alignment. While Persona prompting (PP) is increasin…

Hate Speech Detection

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models

2026-04-27 · Weixing Wang, Liudvikas Zekas, Anton Hackl, Constantin Alexander Auga 외 arxiv

Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evaluation protocols assess these two capabilities independently and do no…

Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations

2026-05-19 · Somnath Banerjee, Pranav Jha, Rima Hazra, Animesh Mukherjee arxiv

LLMs deployed multilingually are often audited via English explanations for non-English inputs. We evaluate extractive explanations ''where the model identifies input token spans as evidence alongside a generated rationa…

Fine-Grained Perspectives: Modeling Explanations with Annotator-Specific Rationales

2026-04-23 · Olufunke O. Sarumi, Charles Welch, Daniel Braun arxiv

Beyond exploring disaggregated labels for modeling perspectives, annotator rationales provide fine-grained signals of individual perspectives. In this work, we propose a framework for jointly modeling annotator-specific …

Natural Language InferenceExplanation Generation