paper-with-me

홈 › Papers

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

2026-08-06 · Arya Labroo, Mengjie Qian, Kate Knill arxiv

Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age. Transformer-based foundation models have improved the accuracy of these L2 speaking graders, but their black-box representations make fairness and interpretability analysis more difficult. Building on prior work that used Concept Activation Vectors (CAVs) to detect bias towards unwanted attributes (`concepts') in feature-based graders, we extend CAV-based analysis to two neural speaking assessment systems: a text-based BERT grader and a speech-and-text multimodal grader based on Whisper. CAVs represent human-interpretable concepts as directions in a model's activation space, allowing us to distinguish between whether a concept is encoded in a model's internal representations and whether it influences the predicted score, the latter quantified using a gradient-based sensitivity metric. Since CAVs rely on linear separability, which is less likely in complex neural embedding spaces, we also investigate whether sparse autoencoders (SAEs) provide cleaner concept directions by learning CAVs in a sparse latent space and mapping them back to activation space. Our analysis shows that concept recoverability depends strongly on the representation and architecture being probed, rather than on the concept alone. Sensitivity to concepts is also architecture-dependent. SAEs make concepts more linearly recoverable, but attenuate the original activation-space sensitivity, especially in low-dimensional layers. These findings highlight the need to distinguish concept recoverability from concept influence when auditing bias in speaking assessment systems.

📄 PDF Abstract BibTeX arXiv:2608.06300

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Advancing Automated Speaking Assessment Leveraging Multifaceted Relevance and Grammar Information

2025-06-19 · Hao-Chien Lu, Jhen-Ke Lin, Hong-Yun Lin, Chung-Chun Wang 외

Current automated speaking assessment (ASA) systems for use in multi-aspect evaluations often fail to make full use of content relevance, overlooking image or exemplar cues, and employ superficial grammar analysis that l…

Fairness on Synthetic Visual and Thermal Mask Images

2022-09-19 · Kenneth Lai, Vlad Shmerko, Svetlana Yanushkevich

In this paper, we study performance and fairness on visual and thermal images and expand the assessment to masked synthetic images. Using the SpeakingFace and Thermal-Mask dataset, we propose a process to assess fairness…

DiversityFairness

Mitigating Data Imbalance in Automated Speaking Assessment

2025-09-03 · Fong-Chun Tsai, Kuan-Tang Huang, Bi-Cheng Yan, Tien-Hong Lo 외 arxiv

Automated Speaking Assessment (ASA) plays a crucial role in evaluating second-language (L2) learners proficiency. However, ASA models often suffer from class imbalance, leading to biased predictions. To address this, we …

Enhancing Public Speaking Skills in Engineering Students Through AI

2025-11-07 · Amol Harsh, Brainerd Prince, Siddharth Siddharth, Deepan Raj Prabakar Muthirayan 외 arxiv

This research-to-practice full paper was inspired by the persistent challenge in effective communication among engineering students. Public speaking is a necessary skill for future engineers as they have to communicate t…

Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration

2026-07-08 · Weicheng Ma, John Guerrerio, Soroush Vosoughi arxiv

Research on stereotypes in large language models (LLMs) has largely focused on English-speaking contexts, due to the lack of datasets in other languages and the high cost of manual annotation in underrepresented cultures…