paper-with-me

홈 › Papers

Beyond Inference-Only Deployment: Comparing Weight-Based Consolidation Against Cascading Compaction

2026-05-23 · Simon Dennis, Kevin Shabahang, Hao Guo, Rivaan Patil arxiv

Major LLM platforms deploy models in an inference-only configuration: the model serves requests but never updates per-user weights. Users must repeatedly re-teach preferences, corrections, and project context, and context-based workarounds consume context-window space and degrade under cascading compaction. We evaluate an alternative: nightly consolidation of interaction knowledge into model weights via reflection, synthesis, and Low-Rank Adaptation (LoRA) fine-tuning on a single consumer GPU. Across ten realistic software development conversations (n = 10, 1,146 test questions across three memory types), three cycles of cascading compaction retain 36.8 +/- 3.0% of knowledge (between an 11.8% no-context floor and a 90.1% full-context ceiling), while consolidation retains 80.4 +/- 1.3% -- a 43.6 pp gain (paired t(9) = 14.8, p < 0.001) that more than doubles what compaction preserves, with the largest gains on procedural corrections (36.3% -> 74.6%) and episodic project facts (31.5% -> 78.2%). As a methodological aside, mean per-token validation cross-entropy is negatively correlated with LLM-judged accuracy (r = -0.51) while median per-token validation cross-entropy tracks accuracy almost exactly (r = +0.99): under evaluators that tolerate surface-form variation, the mean is misleading and a heavy-tail-robust statistic is the faithful signal. Persistent personalization requires moving beyond inference-only deployment toward architectures that consolidate knowledge into weights.

📄 PDF Abstract BibTeX arXiv:2605.24657

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization

2024-11-04 · Eldar Kurtic, Alexandre Marques, Shubhra Pandit, Mark Kurtz 외

Despite the popularity of large language model (LLM) quantization for inference acceleration, significant uncertainty remains regarding the accuracy-performance trade-offs associated with various quantization formats. We…

GPULarge Language ModelQuantization

Lightweight Transformer Models for On-Device Fault Detection: A Benchmark Study on Resource-Constrained Deployment

2026-06-23 · Disha Patel arxiv

On-device fault detection enables real-time diagnostics without cloud dependency, but deploying machine learning models on resource-constrained hardware demands careful tradeoffs between accuracy, latency, and model size…

CoFiDA-M: Concept-Aware Feature Modulation for Cross-Domain Adaptation with Image-Only Inference

2026-05-29 · Nurjahan Sultana, Moi Hoon Yap, Xinqi Fan, Wenqi Lu arxiv

Models for AI-based skin cancer screening suffer a severe performance drop when shifting from expert dermoscopic (source) images to consumer-grade clinical (target) images, hindering real-world deployment. Existing domai…

Domain Adaptation

LagKV: Lag-Relative Information of the KV Cache Tells Which Tokens Are Important

2025-04-07 · Manlai Liang, Jiaming Zhang, Xiong Li, Jinlong Li

The increasing size of the Key-Value (KV) cache during the Large Language Models long-context inference is the main obstacle for its balance between the deployment cost and task accuracy. To reduce the KV cache size in s…

MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices

2025-08-23 · Nishant Gavhane, Arush Mehrotra, Rohit Chawla, Peter Proenca arxiv

The deployment of large-scale Mixture-of-Experts (MoE) models on edge devices presents significant challenges due to memory constraints. While MoE architectures enable efficient utilization of computational resources by …