paper-with-me

홈 › Papers

Do LLMs Reliably Identify Correct Information Units in Aphasic Discourse?

2026-04-10 · Jason M Pittman, Yesenia Medina-Santos, Anton Phillips, Brielle C. Stark arxiv

Correct Information Units (CIUs) are central to discourse assessment in aphasia because they quantify communicative informativeness rather than linguistic form alone. However, CIU scoring is time intensive and requires trained raters. This study examined whether instruction-tuned large language models (LLMs) can reliably perform token-level CIU classification from aphasic discourse transcripts. Sixteen picture-description transcripts elicited with the Cat Rescue stimulus were annotated for CIU status according to Nicholas and Brookshire (1993). The sample spanned four severity strata: control, mild, moderate, and severe aphasia. Four publicly available instruction-tuned LLMs were benchmarked under zero-shot and two few-shot prompting conditions across five stratified random seeds. Performance was evaluated against consensus human labels using accuracy, precision, recall, F1, and Cohen's kappa. Zero-shot prompting was insufficient across models. In contrast, few-shot prompting yielded substantial gains and produced competitive performance for three viable models. Mean few-shot F1 scores ranged from 0.776 to 0.817 across Llama-3.1-8B, Qwen2.5-7B, and Mistral-7B, with no significant differences between fixed global and per-chunk local example selection. Phi-3-mini was unstable and did not yield reliable performance. Viable models showed high recall but lower precision, suggesting systematic over-classification of tokens as CIUs. Performance also varied by discourse severity, with the weakest results in more severe aphasia. Few-shot LLM prompting can support automated CIU identification without gradient-based task training, but agreement with human annotation remains insufficient for fully autonomous use. These findings support LLM-based CIU scoring as a promising human-in-the-loop component of discourse assessment systems.

📄 PDF Abstract BibTeX arXiv:2606.15696

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Different types of syntactic agreement recruit the same units within large language models

2025-12-03 · Daria Kryvosheieva, Andrea de Varda, Evelina Fedorenko, Greta Tuckute arxiv

Large language models (LLMs) can reliably distinguish grammatical from ungrammatical sentences, but how grammatical knowledge is represented within the models remains an open question. We investigate whether different sy…

Estimating Correctness Without Oracles in LLM-Based Code Generation

2025-06-26 · Thomas Valentin, Ardi Madadi, Gaetano Sapia, Marcel Böhme

Generating code from natural language specifications is one of the most successful applications of Large Language Models (LLMs). Yet, they hallucinate: LLMs produce outputs that may be grammatically correct but are factu…

Code Generation

Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations

2026-07-14 · Monica Munnangi, Saiph Savage arxiv

Patients seeking medical information often ask questions that embed incorrect assumptions or misconceptions. In such cases, safe medical communication requires not only answering the question, but identifying and correct…

RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation

2025-08-19 · Tianyi Niu, Jaemin Cho, Elias Stengel-Eskin, Mohit Bansal arxiv

We investigate to what extent Multimodal Large Language Models (MLLMs) can accurately identify the orientation of input images rotated 0°, 90°, 180°, and 270°. This task demands robust visual reasoning capabilities to de…

Spatial ReasoningVisual Reasoning

From Early Encoding to Late Suppression: Interpreting LLMs on Character Counting Tasks

2026-04-01 · Ayan Datta, Mounika Marreddy, Alexander Mehler, Zhixue Zhao 외 arxiv

Large language models (LLMs) exhibit failures on elementary symbolic tasks such as character counting in a word, despite excelling on complex benchmarks. Although this limitation has been noted, the internal reasons rema…