paper-with-me

홈 › Papers

PRISM: Breaking the O(n) Memory Wall in Long-Context LLM Inference via O(1) Photonic Block Selection

2026-03-23 · Hyoseok Park, Yeonsang Park arxiv

Long-context LLM inference is bottlenecked not by compute but by the O(n) memory bandwidth cost of scanning the KV cache at every decode step -- a wall that no amount of arithmetic scaling can break. Recent photonic accelerators have demonstrated impressive throughput for dense attention computation; however, these approaches inherit the same O(n) memory scaling as electronic attention when applied to long contexts. We observe that the real leverage point is the coarse block-selection step: a memory-bound similarity search that determines which KV blocks to fetch. We identify, for the first time, that this task is structurally matched to the photonic broadcast-and-weight paradigm -- the query fans out to all candidates via passive splitting, signatures are quasi-static (matching electro-optic MRR programming), and only rank order matters (relaxing precision to 4-6 bits). Crucially, the photonic advantage grows with context length: as N increases, the electronic scan cost rises linearly while the photonic evaluation remains O(1). We instantiate this insight in PRISM (Photonic Ranking via Inner-product Similarity with Microring weights), a thin-film lithium niobate (TFLN) similarity engine. Hardware-impaired needle-in-a-haystack evaluation on Qwen2.5-7B confirms 100% accuracy from 4K through 64K tokens at k=32, with 16x traffic reduction at 64K context. PRISM achieves a four-order-of-magnitude energy advantage over GPU baselines at practical context lengths (n >= 4K).

📄 PDF Abstract BibTeX arXiv:2603.21576

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Breaking On-device Training Memory Wall: A Systematic Survey

2023-06-17 · Shitian Li, Chunlin Tian, Kahou Tam, Rui Ma 외

On-device training has become an increasingly popular approach to machine learning, enabling models to be trained directly on mobile and edge devices. However, a major challenge in this area is the limited memory availab…

NavigateSurvey

PRISM: Pareto-Efficient Retrieval over Intent-Aware Structured Memory for Long-Horizon Agents

2026-05-12 · Jingyi Peng, Zhongwei Wan, Weiting Liu, Qiuzhuang Sun arxiv

Long-horizon language agents accumulate conversation history far faster than any fixed context window can hold, making memory management critical to both answer accuracy and serving cost. Existing approaches either expan…

Long-Range Tasks Using Short-Context LLMs: Incremental Reasoning With Structured Memories

2024-12-25 · Dulhan Jayalath, James Bradley Wendt, Nicholas Monath, Sandeep Tata 외

Long-range tasks require reasoning over long inputs. Existing solutions either need large compute budgets, training data, access to model weights, or use complex, task-specific approaches. We present PRISM, which allevia…

Prism-$Δ$: Differential Subspace Steering for Prompt Highlighting in Large Language Models

2026-03-11 · Yuyao Ge, Shenghua Liu, Yiwei Wang, Tianyu Liu 외 arxiv

Prompt highlighting steers a large language model to prioritize user-specified text spans during generation. A key challenge is extracting steering directions that capture the difference between relevant and irrelevant c…

Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving

2025-05-06 · Shan Yu, Jiarong Xing, Yifan Qiao, Mingyuan Ma 외

Serving large language models (LLMs) is expensive, especially for providers hosting many models, making cost reduction essential. The unique workload patterns of serving multiple LLMs (i.e., multi-LLM serving) create new…

GPUScheduling