paper-with-me

홈 › Papers

How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF

2026-07-22 · Venkata Naga Sai Vishnu Rohit Pulipaka, Anish Katta, Deva Rohit Reddy Peddireddy arxiv

In RLHF pipelines, reward scoring blocks policy updates. Slow scoring bottlenecks the entire loop, since no update runs until every rollout gets a score. And yet most setups just default to PyTorch eager mode or torch.compile, no one checks if that's actually fastest. Scoring itself is small. Rollout generation eats far more of a typical RLHF step. But scoring and generation fight over the same CPU and GPU resources, so a faster scoring engine doesn't shrink step time on its own. It mainly frees up capacity generation can use instead. We built a native C++ inference engine on ONNX Runtime. First step: confirm correctness. Output matched the PyTorch reference to 5.7 x 10^-6 on CPU and 4.2 x 10^-3 on GPU, close enough to trust. Then we tested it against PyTorch eager mode, torch.compile, and FastAPI, on both CPU and GPU. CPU was decisive. Our engine beat every baseline, confidence intervals didn't even overlap. GPU gave a different view: we beat PyTorch and FastAPI, but torch.compile came out ahead. Further testing traced the speedup to ONNX Runtime itself, not C++ as a language. And batching strategy mattered more than either the language or the runtime choice, more than we expected. The results are from repeated, independent runs, since single runs just aren't reliable enough to trust.

📄 PDF Abstract BibTeX arXiv:2607.19712

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fast Graph Representation Learning with PyTorch Geometric

2019-03-06 · Matthias Fey, Jan Eric Lenssen

We introduce PyTorch Geometric, a library for deep learning on irregularly structured input data such as graphs, point clouds and manifolds, built upon PyTorch. In addition to general graph data structures and processing…

GPUGraph ClassificationGraph Representation LearningNode Classification+2

An Empirical Study on Bugs Inside PyTorch: A Replication Study

2023-07-25 · Sharon Chee Yin Ho, Vahid Majdinasab, Mohayeminul Islam, Diego Elias Costa 외

Software systems are increasingly relying on deep learning components, due to their remarkable capability of identifying complex data patterns and powering intelligent behaviour. A core enabler of this change in software…

Deep Learning

Stochastic Gradient Descent without Full Data Shuffle

2022-06-12 · Lijie Xu, Shuang Qiu, Binhang Yuan, Jiawei Jiang 외

Stochastic gradient descent (SGD) is the cornerstone of modern machine learning (ML) systems. Despite its computational efficiency, SGD requires random data access that is inherently inefficient when implemented in syste…

Computational Efficiency

Sockeye 3: Fast Neural Machine Translation with PyTorch

2022-07-12 · Felix Hieber, Michael Denkowski, Tobias Domhan, Barbara Darques Barros 외

Sockeye 3 is the latest version of the Sockeye toolkit for Neural Machine Translation (NMT). Now based on PyTorch, Sockeye 3 provides faster model implementations and more advanced features with a further streamlined cod…

Machine TranslationNMTTranslation

FasterAI: A Lightweight Library for Creating Sparse Neural Networks

2022-07-03 · Nathan Hubens

FasterAI is a PyTorch-based library, aiming to facilitate the utilization of deep neural networks compression techniques such as sparsification, pruning, knowledge distillation, or regularization. The library is built wi…

Knowledge Distillation