paper-with-me

Papers Inference Optimization

“Inference Optimization” 태그가 달린 논문 56편 · 필터 해제

Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging

2025-06-29 · Lujun Li, Zhu Qiyuan, Jiacheng Wang, Wei Li 외

Mixture of Experts (MoE) LLMs face significant obstacles due to their massive parameter scale, which imposes memory, storage, and deployment challenges. Although recent expert merging methods promise greater efficiency b…

Inference OptimizationMixture-of-Experts

The Foundation Cracks: A Comprehensive Study on Bugs and Testing Practices in LLM Libraries

2025-06-14 · Weipeng Jiang, XiaoYu Zhang, Xiaofei Xie, Jiongchi Yu 외

Large Language Model (LLM) libraries have emerged as the foundational infrastructure powering today's AI revolution, serving as the backbone for LLM deployment, inference optimization, fine-tuning, and production serving…

Bug fixingInference OptimizationLarge Language Model

Brevity is the soul of sustainability: Characterizing LLM response lengths

2025-06-10 · Soham Poddar, Paramita Koley, Janardan Misra, Sanjay Podder 외

A significant portion of the energy consumed by Large Language Models (LLMs) arises from their inference processes; hence developing energy-efficient methods for inference is crucial. While several techniques exist for i…

DecoderInference OptimizationPrompt Engineering

DSMentor: Enhancing Data Science Agents with Curriculum Learning and Online Knowledge Accumulation

2025-05-20 · He Wang, Alexander Hanbo Li, Yiqun Hu, Sheng Zhang 외

Large language model (LLM) agents have shown promising performance in generating code for solving complex data science problems. Recent studies primarily focus on enhancing in-context learning through improved search, sa…

In-Context LearningInference OptimizationLarge Language Model

Faster MoE LLM Inference for Extremely Large Models

2025-05-06 · Haoqi Yang, Luohe Shi, Qiwei Li, Zuchao Li 외

Sparse Mixture of Experts (MoE) large language models (LLMs) are gradually becoming the mainstream approach for ultra-large-scale models. Existing optimization efforts for MoE models have focused primarily on coarse-grai…

Inference OptimizationMixture-of-Experts

Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints

2025-04-15 · Ruicheng Ao, Gan Luo, David Simchi-Levi, Xinshang Wang

Large Language Models (LLMs) are indispensable in today's applications, but their inference procedure -- generating responses by processing text in segments and using a memory-heavy Key-Value (KV) cache -- demands signif…

GPUInference OptimizationScheduling

SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

2025-04-15 · Junke Wang, Zhi Tian, Xun Wang, Xinyu Zhang 외

This work presents SimpleAR, a vanilla autoregressive visual generation framework without complex architecure modifications. Through careful exploration of training and inference optimization, we demonstrate that: 1) wit…

Inference Optimization

The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation

2025-04-07 · Hao Fang, Runmin Cong, Xiankai Lu, Zhiyang Chen 외

Motion expression video segmentation is designed to segment objects in accordance with the input motion expressions. In contrast to the conventional Referring Video Object Segmentation (RVOS), it places emphasis on motio…

Inference OptimizationReferring Video Object SegmentationSegmentationSemantic Segmentation+3

Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification

2025-02-23 · Arshia Kermani, Ehsan Zeraatkar, Habib Irani

The increasing computational demands of transformer models in time series classification necessitate effective optimization strategies for energy-efficient deployment. Our study presents a systematic investigation of opt…

ClassificationInference OptimizationQuantizationTime Series+1

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization

2025-02-14 · Bowen Pang, Kai Li, Ruifeng She, Feifan Wang

With the development of large language models (LLMs), it has become increasingly important to optimize hardware usage and improve throughput. In this paper, we study the inference optimization of the serving system that …

GSM8KInference OptimizationLanguage ModelingLanguage Modelling+2

DVFS-Aware DNN Inference on GPUs: Latency Modeling and Performance Analysis

2025-02-10 · Yunchu Han, Zhaojun Nan, Sheng Zhou, Zhisheng Niu

The rapid development of deep neural networks (DNNs) is inherently accompanied by the problem of high computational costs. To tackle this challenge, dynamic voltage frequency scaling (DVFS) is emerging as a promising tec…

CPUInference Optimization

Hellinger-Kantorovich Gradient Flows: Global Exponential Decay of Entropy Functionals

2025-01-28 · Alexander Mielke, Jia-Jie Zhu

We investigate a family of gradient flows of positive and probability measures, focusing on the Hellinger-Kantorovich (HK) geometry, which unifies transport mechanism of Otto-Wasserstein, and the birth-death mechanism of…

Inference Optimization

A Survey on Inference Optimization Techniques for Mixture of Experts Models

2024-12-18 · Jiacheng Liu, Peng Tang, Wenfeng Wang, Yuhang Ren 외

The emergence of large-scale Mixture of Experts (MoE) models represents a significant advancement in artificial intelligence, offering enhanced model capacity and computational efficiency through conditional computation.…

Computational EfficiencyDistributed ComputingInference OptimizationKnowledge Distillation+4

FluidML: Fast and Memory Efficient Inference Optimization

2024-11-14 · Jinjie Liu, Hang Qiu

Machine learning models deployed on edge devices have enabled numerous exciting new applications, such as humanoid robots, AR glasses, and autonomous vehicles. However, the computing resources available on these edge dev…

Autonomous VehiclesInference OptimizationManagement

A Temporal Linear Network for Time Series Forecasting

2024-10-28 · Remi Genet, Hugo Inzirillo

Recent research has challenged the necessity of complex deep learning architectures for time series forecasting, demonstrating that simple linear models can often outperform sophisticated approaches. Building upon this i…

Computational EfficiencyInference OptimizationTime SeriesTime Series Forecasting

LLM-Rank: A Graph Theoretical Approach to Pruning Large Language Models

2024-10-17 · David Hoffmann, Kailash Budhathoki, Matthaeus Kleindessner

The evolving capabilities of large language models are accompanied by growing sizes and deployment costs, necessitating effective inference optimisation techniques. We propose a novel pruning method utilising centrality …

Inference OptimizationNetwork Pruning

EdgeRL: Reinforcement Learning-driven Deep Learning Model Inference Optimization at Edge

2024-10-16 · Motahare Mounesan, Xiaojie Zhang, Saptarshi Debroy

Balancing mutually diverging performance metrics, such as, processing latency, outcome accuracy, and end device energy consumption is a challenging undertaking for deep learning model inference in ad-hoc edge environment…

Deep LearningInference OptimizationReinforcement Learning (RL)

CycleBNN: Cyclic Precision Training in Binary Neural Networks

2024-09-28 · Federico Fontana, Romeo Lanzino, Anxhelo Diko, Gian Luca Foresti 외

This paper works on Binary Neural Networks (BNNs), a promising avenue for efficient deep learning, offering significant reductions in computational overhead and memory footprint to full precision networks. However, the c…

Inference Optimization

Revisiting SMoE Language Models by Evaluating Inefficiencies with Task Specific Expert Pruning

2024-09-02 · Soumajyoti Sarkar, Leonard Lausen, Volkan Cevher, Sheng Zha 외

Sparse Mixture of Expert (SMoE) models have emerged as a scalable alternative to dense models in language modeling. These models use conditionally activated feedforward subnetworks in transformer blocks, allowing for a s…

Inference OptimizationLanguage ModelingLanguage Modelling

The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities

2024-08-23 · Venkatesh Balavadhani Parthasarathy, Ahtsham Zafar, Aafaq Khan, Arsalan Shahid

This report examines the fine-tuning of Large Language Models (LLMs), integrating theoretical insights with practical applications. It outlines the historical evolution of LLMs from traditional Natural Language Processin…

Computational EfficiencyInference OptimizationMixture-of-Experts
1–20 / 56 다음 →