paper-with-me

홈 › Papers

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes

2026-08-08 · Yu Zhang arxiv

Mixture-of-Experts (MoE) models have outgrown accelerator memory, and offloading expert weights to host memory is now standard. This makes expert cache management an attractive lever: a policy that raised the hit rate would cut expert traffic per token. Evaluating that is a measurement problem, and we find the measurement fragile. With a trace-driven, event-atomic simulator over three MoE models (40, 64, 128 experts), we isolate three evaluation axes that change conclusions, not just numbers. Replay semantics: under a fused-event traffic contract, an inconsistent per-access replay inflates recency-based policies by 27-29% while leaving frequency-based and static ones within 4%, inverting the policy ranking. Workload contamination: probe sets using one instruction template per category produce verbatim-identical generation prefixes; a matched-pair rendering intervention moves the measured early-window effect by 19.4-31.9 points and reverses which workloads look most cache-friendly. Operating regimes: normalized miss fractions do not transfer across models, so the per-step expert union relative to per-layer capacity must be reported -- yet permuting only the temporal order of an identical event stream moves the offline-optimal gap from 44.9% to 30.8%, so it is not sufficient. Corrected, a stable gap to the offline optimum remains (44.2-45.9% over 13 frozen workload compositions). A forced-admission oracle attributes 84.3-96.6% of it to knowing which resident expert is used furthest in the future. A causal next-use predictor, used as an eviction rule, recovers -11.4% of the gap; it picks an optimal victim 3.4% of the time, against 2.4% for a random resident block and 20.6-22.1% for LRU and LFRU. Our position is narrow: in our evaluated settings a large offline-optimal gap substantially overstates the gains recovered by representative lightweight causal mechanisms.

📄 PDF Abstract BibTeX arXiv:2608.07911

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework

2026-04-23 · Tharindu Kumarage, Lisa Bauer, Yao Ma, Dan Rosen 외 arxiv

As reasoning capacity and deployment scope grow in tandem, large language models (LLMs) gain the capacity to engage in behaviors that serve their own objectives, a class of risks we term Emergent Strategic Reasoning Risk…

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation

2026-06-13 · Xian Sun, Wei Gao, Yingshuo Wang, Lingdong Kong 외 arxiv

Reasoning models are increasingly used in settings where the final answer is not the only object of review: educational tools may show students intermediate steps, decision-support systems may require human oversight, an…

When One Modality Sabotages the Others: A Diagnostic Lens on Multimodal Reasoning

2025-11-04 · Chenyu Zhang, Minsol Kim, Shohreh Ghorbani, Jingyao Wu 외 arxiv

Despite rapid growth in multimodal large language models (MLLMs), their reasoning traces remain opaque: it is often unclear which modality drives a prediction, how conflicts are resolved, or when one stream dominates. In…

Multimodal Emotion RecognitionMultimodal Reasoning

CausalSim: A Causal Framework for Unbiased Trace-Driven Simulation

2022-01-05 · Abdullah Alomar, Pouya Hamadanian, Arash Nasr-Esfahany, Anish Agarwal 외

We present CausalSim, a causal framework for unbiased trace-driven simulation. Current trace-driven simulators assume that the interventions being simulated (e.g., a new algorithm) would not affect the validity of the tr…

Causal Inference

From Pixels to BFS: High Maze Accuracy Does Not Imply Visual Planning

2026-03-27 · Alberto G. Rodriguez Salgado arxiv

How do multimodal models solve visual spatial tasks -- through genuine planning, or through brute-force search in token space? We introduce \textsc{MazeBench}, a benchmark of 110 procedurally generated maze images across…