paper-with-me

Papers

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation

2026-05-08 · Adib Sakhawat, Tahsin Islam, Takia Farhin, Syed Rifat Raiyan, Hasan Mahmud, Md Kamrul Hasan arxiv

The evaluation of Large Language Models (LLMs) faces a critical challenge in construct validity, where fragmented benchmarks and ad hoc metrics frequently conflate method variance, such as prompt sensitivity, with true latent capabilities. Concurrently, emerging research suggests that LLM capabilities and outputs can be modeled as continuous geometric manifolds. In this Systematization of Knowledge (SoK), we bridge these paradigms by proposing a generalized Multi-Trait Multi-Method (MTMM) framework for LLM evaluation. We formalize and unify nine evaluation metrics, including Paraphrase Instability, Drift Score, Overton Width, and Pluralism Score, interpreting them not as isolated scalar values but as geometric measurements within a shared latent coordinate space. This spatial unification factorizes model behavior into three orthogonal latent dimensions: (1) Instability and Sensitivity, (2) Position and Alignment, and (3) Coverage and Expressiveness. By systematically separating task-irrelevant perturbations from true capability spans, the framework provides a theoretically grounded and domain-agnostic taxonomy for robust and empirically stable benchmark design.

📄 PDF Abstract BibTeX arXiv:2605.08522

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution

2026-05-26 · Adib Sakhawat, Syed Rifat Raiyan, Tahsin Islam, Takia Farhin 외 arxiv

We argue, with systematic empirical evidence, that a large language model's political ideology is not a fixed point, but a conditional distribution $\mathbb{P}($position$\mid$context$)$ over a real political space. We ev…

XRBench: An Extended Reality (XR) Machine Learning Benchmark Suite for the Metaverse

2022-11-16 · Hyoukjun Kwon, Krishnakumar Nair, Jamin Seo, Jason Yik 외

Real-time multi-task multi-model (MTMM) workloads, a new form of deep learning inference workloads, are emerging for applications areas like extended reality (XR) to support metaverse use cases. These workloads combine u…

MTMMC: A Large-Scale Real-World Multi-Modal Camera Tracking Benchmark

2024-03-29 · CVPR 2024 1 · Sanghyun Woo, KwanYong Park, Inkyu Shin, Myungchul Kim 외

Multi-target multi-camera tracking is a crucial task that involves identifying and tracking individuals over time using video streams from multiple cameras. This task has practical applications in various fields, such as…

Anomaly DetectionHuman DetectionMultiple Object TrackingObject Tracking

Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation

2026-04-02 · Chongjie Ye, Cheng Cao, Chuanyu Pan, Yiming Hao 외 arxiv

Recent multimodal large language models have achieved strong performance in unified text and image understanding and generation, yet extending such native capability to 3D remains challenging due to limited data. Compare…

3D Generation

HoGS: Unified Near and Far Object Reconstruction via Homogeneous Gaussian Splatting

2025-03-25 · CVPR 2025 1 · Xinpeng Liu, Zeyi Huang, Fumio Okura, Yasuyuki Matsushita

Novel view synthesis has demonstrated impressive progress recently, with 3D Gaussian splatting (3DGS) offering efficient training time and photorealistic real-time rendering. However, reliance on Cartesian coordinates li…

3DGSNovel View SynthesisObject Reconstruction