| Abstract | Many QoS-driven autoscaling settings rely on a single aggregated tail-latency metric, such as the overall P95 response time, as a compact signal of system behavior. This practice is straightforward when incoming requests are approximately homogeneous, but becomes harder to interpret when a shared backend serves heterogeneous request classes with different latency characteristics and time-varying composition. This paper studies a class-unobservable heterogeneous workload setting in which the decision layer observes only aggregated QoS signals derived from the combined response stream. In such settings, the observed global percentile is computed from a mixture distribution rather than from any one request class in isolation. As a result, the same aggregated tail-latency value may remain compatible with materially different underlying class-level conditions. We use the term composition masking for regimes in which degradation affecting part of a heterogeneous workload is only weakly reflected in the aggregated tail-latency signal. The contribution builds on known aggregation behavior to characterize a class-sensitive interpretability limitation in a class-unobservable QoS setting and to distinguish composition masking from related composition-driven effects. The paper characterizes this limitation across three levels. First, it develops a reduced analytical model that isolates the effect of heterogeneous mixtures under aggregated observability. Second, it uses controlled numerical examples to show how global percentile readings can diverge from class-level behavior in a deliberately reduced two-class setting and in a compact three-class extension. Third, it demonstrates the same qualitative pattern on a cloud-hosted quantum-circuit simulation backend used as a controlled empirical witness. The results show that a global P95 may remain weakly responsive while substantial degradation develops within one part of theworkload, and thatworkload-composition change alone can alter the observed QoS signal even under fixed external request rate in the studied setting. The contribution of the paper is a structured analytical and controlled empirical characterization of a diagnostic limitation of aggregated tail-latency metrics under heterogeneous and composition-varying workloads. |
|---|