Tail-Latency-Safe Privacy for Predictive Autoscaling via a Lightweight Encrypted Prediction Head

Publication Date

1-1-2026

Document Type

Conference Proceeding

Publication Title

2026 IEEE 12th International Conference on Network Softwarization Autonomous and Reliable Softwarized Networks in the Age of Distributed Intelligence Netsoft 2026 Proceedings

DOI

10.1109/NetSoft70012.2026.11603491

First Page

219

Last Page

224

Abstract

Predictive autoscaling must meet strict tail-latency service-level objectives (SLOs) while limiting disclosure of tenant telemetry in multi-tenant environments. Full-model private inference is typically too costly for real-time control paths. We present AutoHE-Lite, a serving-time design that keeps the recurrent forecasting backbone in plaintext and places only the final fully connected (FC) prediction head under CKKS approximate homomorphic encryption. This boundary hides the decision-time hidden activation from the provider-side serving node, but does not claim protection against output-based inference, traffic analysis, or side-channel leakage. AutoHE-Lite includes a tail-aware tuner that selects CKKS parameters using explicit gates on 95th-percentile latency (p 95) and fidelity relative to plaintext, and emits a JSON artifact for regression and rollback. Using Google Cluster Trace sequences, the best configuration meets a 50 ms p95 decision budget with negligible drift relative to plaintext and achieves 174.6QPS @p95 through microbatching. The contribution is operational: a deployable, auditable encrypted decision boundary for latency-critical autoscaling.

Funding Sponsor

State of California

Keywords

CKKS, Cloud Control Plane, Homomorphic Encryption, Predictive Autoscaling, Secure Inference, Tail Latency

Department

Computer Science

Share

COinS