paper-with-me

홈 › Papers

When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning

2026-08-03 · Andrew Zhang arxiv

A stable compression score can still select the worse model. In our dense study, a split-half reliable path-quadratic score predicted a 16.1\% gain, while the selected endpoints were 6.0--7.7% worse than two controls. We ask what a compression statistic can justify when deployment cares about the worst supplied group. We treat each statistic as an information interface. Its observation leaves a fiber of compatible endpoint-risk tables, and only orders fixed across that fiber are identified. Cone and fiber identities quantify the remaining uncertainty, while matched observations reverse endpoint order for pooled moments, group-local moments, and reference-path curvature. Sequential composition adds one state variable: the slack from each group risk to the current maximum. This vector determines every unrestricted one-step response, and a margin condition keeps the active group fixed along paths with bounded relative drift. The experiments follow the same ladder. Across three dense LLMs, an early-preserving allocation reduces worst-group perplexity inflation by 12.6--20.9%; target-matched complete-menu selection improves over its references by 2.7--8.0%. Across all 16 routed layers of OLMoE, pooled endpoint refresh lowers held-out worst-group teacher KL by 15.8% over the best static score. A compute-matched hard-max trajectory ends 32.7% worse than pooled, and neither adaptive trajectory improves excess NLL. Local evidence can narrow a menu. Complete endpoints rank that menu, while multistep claims also require control of the evolving active face and future candidates.

📄 PDF Abstract BibTeX arXiv:2608.02940

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Syntactically Look-Ahead Attention Network for Sentence Compression

2020-02-04 · Hidetaka Kamigaito, Manabu Okumura

Sentence compression is the task of compressing a long sentence into a short one by deleting redundant words. In sequence-to-sequence (Seq2Seq) based models, the decoder unidirectionally decides to retain or delete words…

DecoderInformativenessSentenceSentence Compression

Active Context Compression: Autonomous Memory Management in LLM Agents

2026-01-12 · Nikhil Verma arxiv

Large Language Model (LLM) agents struggle with long-horizon software engineering tasks due to "Context Bloat." As interaction history grows, computational costs explode, latency increases, and reasoning capabilities deg…

Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models

2025-11-19 · Singon Kim arxiv

Abstractive compression utilizes smaller langauge models to condense query-relevant context, reducing computational costs in retrieval-augmented generation (RAG). However, retrieved documents often include information th…

Data Augmentation

ACoRN: Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models

2025-04-17 · Singon Kim, Gunho Jung, Seong-Whan Lee

Abstractive compression utilizes smaller langauge models to condense query-relevant context, reducing computational costs in retrieval-augmented generation (RAG). However,retrieved documents often include information tha…

Data AugmentationRAGRetrievalRetrieval-augmented Generation

Eloss in the way: A Sensitive Input Quality Metrics for Intelligent Driving

2023-02-02 · Haobo Yang, Shiyan Zhang, Zhuoyi Yang, Xinyu Zhang

With the increasing complexity of the traffic environment, the importance of safety perception in intelligent driving is growing. Conventional methods in the robust perception of intelligent driving focus on training mod…

3D Object DetectionAnomaly Detection