paper-with-me

홈 › Papers

Exploring Human-AI Conceptual Alignment through the Prism of Chess

2025-10-29 · Semyon Lomasov, Judah Goldfeder, Mehmet Hamza Erol, Matthew So, Yao Yan, Addison Howard, Nathan Kutz, Ravid Shwartz Ziv arxiv

Do AI systems truly understand human concepts or merely mimic surface patterns? We investigate this through chess, where human creativity meets precise strategic concepts. Analyzing a 270M-parameter transformer that achieves grandmaster-level play, we uncover a striking paradox: while early layers encode human concepts like center control and knight outposts with up to 85\% accuracy, deeper layers, despite driving superior performance, drift toward alien representations, dropping to 50-65\% accuracy. To test conceptual robustness beyond memorization, we introduce the first Chess960 dataset: 240 expert-annotated positions across 6 strategic concepts. When opening theory is eliminated through randomized starting positions, concept recognition drops 10-20\% across all methods, revealing the model's reliance on memorized patterns rather than abstract understanding. Our layer-wise analysis exposes a fundamental tension in current architectures: the representations that win games diverge from those that align with human thinking. These findings suggest that as AI systems optimize for performance, they develop increasingly alien intelligence, a critical challenge for creative AI applications requiring genuine human-AI collaboration. Dataset and code are available at: https://github.com/slomasov/ChessConceptsLLM.

📄 PDF Abstract BibTeX arXiv:2510.26025

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Semantic Prompting: Agentic Incremental Narrative Refinement through Spatial Semantic Interaction

2026-04-21 · Xuxin Tang, Ibrahim Tahmid, Eric Krokos, Kirsten Whitley 외 arxiv

Interactive spatial layouts empower users to synthesize information and organize findings for sensemaking. While Large Language Models (LLMs) can automate narrative generation from spatial layouts, current collage-based …

The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models

2024-04-24 · Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean 외

Human feedback is central to the alignment of Large Language Models (LLMs). However, open questions remain about methods (how), domains (where), people (who) and objectives (to what end) of feedback processes. To navigat…

DiversityNavigate

PRISM: Learning Design Knowledge from Data for Stylistic Design Improvement

2026-01-16 · Huaxiaoyue Wang, Sunav Choudhary, Franck Dernoncourt, Yu Shen 외 arxiv

Graphic design often involves exploring different stylistic directions, which can be time-consuming for non-experts. We address this problem of stylistically improving designs based on natural language instructions. Whil…

A Taxonomy of Conceptual Alignment in Human-Robot Dialogue

2026-06-21 · Shengchen Zhang, Xiaohua Sun, Weiwei Guo arxiv

Successful conversations require speakers to align on the meaning of concepts, a challenging but crucial task for human-robot interaction. Understanding the process of establishing such alignment is hindered by competing…

Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM

2026-03-19 · Zizhao Hu, Mohammad Rostami, Jesse Thomason arxiv

Persona prompting can steer LLM generation towards a domain-specific tone and pattern. This behavior enables use cases in multi-agent systems where diverse interactions are crucial and human-centered tasks require high-l…