paper-with-me

Papers

FlexServe: Deployment of PyTorch Models as Flexible REST Endpoints

2020-02-29 · Edward Verenich, Alvaro Velasquez, M. G. Sarwar Murshed, Faraz Hussain

The integration of artificial intelligence capabilities into modern software systems is increasingly being simplified through the use of cloud-based machine learning services and representational state transfer architecture design. However, insufficient information regarding underlying model provenance and the lack of control over model evolution serve as an impediment to the more widespread adoption of these services in many operational environments which have strict security requirements. Furthermore, tools such as TensorFlow Serving allow models to be deployed as RESTful endpoints, but require error-prone transformations for PyTorch models as these dynamic computational graphs. This is in contrast to the static computational graphs of TensorFlow. To enable rapid deployments of PyTorch models without intermediate transformations we have developed FlexServe, a simple library to deploy multi-model ensembles with flexible batching.

📄 PDF Abstract BibTeX arXiv:2003.01538

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation

2026-03-10 · Yinpeng Wu, Yitong Chen, Lixiang Wang, Jinyu Gu 외 arxiv

Device-side Large Language Models (LLMs) have witnessed explosive growth, offering higher privacy and availability compared to cloud-side LLMs. During LLM inference, both model weights and user data are valuable, and att…

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation

2026-06-22 · Yinpeng Wu, Yitong Chen, Lixiang Wang, Jinyu Gu 외 arxiv

Device-side Large Language Models (LLMs) have grown explosively, offering stronger privacy and higher availability than their cloud-side counterparts. During LLM inference, both the model weights and the user data are va…

Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services

2026-08-19 · José A. Perdiguero López, Miguel A. Durán-Olivencia arxiv

We present Flama, an open-source Python framework for developing and deploying production-ready web APIs, machine learning services, and large-language-model (LLM) applications. Built on the Asynchronous Server Gateway I…

ensemblQueryR: fast, flexible and high-throughput querying of Ensembl LD API endpoints in R

2023-08-13 · Aine Fairbrother-Browne, Sonia García-Ruiz, Regina H Reynolds, Mina Ryten 외

We present ensemblQueryR, a package providing an R interface to the Ensembl REST API that facilitates flexible, fast, user-friendly and R workflow integrable querying of Ensembl REST API linkage disequilibrium (LD) endpo…

Masked LARk: Masked Learning, Aggregation and Reporting worKflow

2021-10-27 · Joseph J. Pfeiffer III, Denis Charles, Davis Gilton, Young Hun Jung 외

Today, many web advertising data flows involve passive cross-site tracking of users. Enabling such a mechanism through the usage of third party tracking cookies (3PC) exposes sensitive user data to a large number of part…

Privacy Preserving