paper-with-me

Papers

Resilient by Design -- Active Inference for Distributed Continuum Intelligence

2025-11-10 · Praveen Kumar Donta, Alfreds Lapkovskis, Enzo Mingozzi, Schahram Dustdar arxiv

Failures are the norm in highly complex and heterogeneous devices spanning the distributed computing continuum (DCC), from resource-constrained IoT and edge nodes to high-performance computing systems. Ensuring reliability and global consistency across these layers remains a major challenge, especially for AI-driven workloads requiring real-time, adaptive coordination. This work-in-progress paper introduces a Probabilistic Active Inference Resilience Agent (PAIR-Agent) to achieve resilience in DCC systems. PAIR-Agent performs three core operations: (i) constructing a causal fault graph from device logs, (ii) identifying faults while managing certainties and uncertainties using Markov blankets and the free energy principle, and (iii) autonomously healing issues through active inference. Through continuous monitoring and adaptive reconfiguration, the agent maintains service continuity and stability under diverse failure conditions. Theoretical validations confirm the reliability and effectiveness of the proposed framework.

📄 PDF Abstract BibTeX arXiv:2511.07202

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Guardians of the Deep Fog: Failure-Resilient DNN Inference from Edge to Cloud

2019-09-03 · Ashkan Yousefpour, Siddartha Devic, Brian Q. Nguyen, Aboudy Kreidieh 외

Partitioning and distributing deep neural networks (DNNs) over physical nodes such as edge, fog, or cloud nodes, could enhance sensor fusion, and reduce bandwidth and inference latency. However, when a DNN is distributed…

Distributed ComputingSensor Fusion

Benchmarking Dynamic SLO Compliance in Distributed Computing Continuum Systems

2025-03-05 · Alfreds Lapkovskis, Boris Sedlak, Sindri Magnússon, Schahram Dustdar 외

Ensuring Service Level Objectives (SLOs) in large-scale architectures, such as Distributed Computing Continuum Systems (DCCS), is challenging due to their heterogeneous nature and varying service requirements across diff…

BenchmarkingCPUDistributed Computing

Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices

2024-06-20 · Li Wang, Liang Li, Lianming Xu, Xian Peng 외

The distributed inference paradigm enables the computation workload to be distributed across multiple devices, facilitating the implementations of deep learning based intelligent services on extremely resource-constraine…

Knowledge DistillationModel Compression

ResiliNet: Failure-Resilient Inference in Distributed Neural Networks

2020-02-18 · Ashkan Yousefpour, Brian Q. Nguyen, Siddartha Devic, Guanhua Wang 외

Federated Learning aims to train distributed deep models without sharing the raw data with the centralized server. Similarly, in distributed inference of neural networks, by partitioning the network and distributing it a…

Federated Learning

CityLearn v2: Energy-flexible, resilient, occupant-centric, and carbon-aware management of grid-interactive communities

2024-05-02 · Kingsley Nweye, Kathryn Kaspar, Giacomo Buscemi, Tiago Fonseca 외

As more distributed energy resources become part of the demand-side infrastructure, it is important to quantify the energy flexibility they provide on a community scale, particularly to understand the impact of geographi…

BenchmarkingManagementreinforcement-learningReinforcement Learning