Approach
Combine CPU, memory, network throughput, and a custom metric (active sessions) into one weighted composite score, then base the scaling decision on a smoothed (EMA, exponential moving average, a running average that weights recent samples more heavily and damps out single-sample noise) version of that score rather than the raw instantaneous value, with a dead zone around the target and a cooldown after any action. Each of those three mechanisms defends against a different flapping cause: EMA smooths per-sample noise, the dead zone stops the decision from firing on small, harmless deviations, and cooldown stops repeated actions from stacking before the previous action's effect has actually shown up in the metrics.
python
import random
class MetricAutoscaler:
def __init__(self, target_score=0.60, weights=None, ema_alpha=0.3,
dead_zone=0.10, cooldown_ticks=3, max_step=2):
self.weights = weights or {"cpu": 0.4, "mem": 0.2, "net": 0.2, "sessions": 0.2}
self.target_score = target_score
self.ema_alpha = ema_alpha
self.dead_zone = dead_zone
self.cooldown_ticks = cooldown_ticks
self.max_step = max_step
self.ema_score = None
self.cooldown_remaining = 0
self.instances = 4
self.history = []
def normalize(self, raw):
cpu = raw["cpu_pct"] / 100.0
mem = raw["mem_pct"] / 100.0
net = raw["net_pct"] / 100.0
sessions = min(raw["sessions"] / raw["sessions_capacity"], 1.5)
return {"cpu": cpu, "mem": mem, "net": net, "sessions": sessions}
def composite_score(self, norm):
return sum(self.weights[k] * norm[k] for k in self.weights)
def tick(self, raw_metrics):
norm = self.normalize(raw_metrics)
raw_score = self.composite_score(norm)
if self.ema_score is None:
self.ema_score = raw_score
else:
self.ema_score = self.ema_alpha * raw_score + (1 - self.ema_alpha) * self.ema_score
action, delta = "hold", 0
if self.cooldown_remaining > 0:
self.cooldown_remaining -= 1
else:
deviation = self.ema_score - self.target_score
if deviation > self.dead_zone:
delta = min(self.max_step, 1 + int(deviation / self.dead_zone))
action = "scale_out"
elif deviation < -self.dead_zone:
delta = -1
action = "scale_in"
if action != "hold":
self.instances = max(2, self.instances + delta)
self.cooldown_remaining = self.cooldown_ticks
self.history.append({"raw_score": round(raw_score, 3), "ema_score": round(self.ema_score, 3),
"action": action, "instances": self.instances})
return self.history[-1]
def make_noisy_metrics(base_cpu, base_mem, base_net, base_sessions, sessions_capacity, rng):
return {
"cpu_pct": max(0, base_cpu + rng.gauss(0, 8)),
"mem_pct": max(0, base_mem + rng.gauss(0, 5)),
"net_pct": max(0, base_net + rng.gauss(0, 10)),
"sessions": max(0, base_sessions + rng.gauss(0, sessions_capacity * 0.05)),
"sessions_capacity": sessions_capacity,
}
if __name__ == "__main__":
rng = random.Random(1234)
scaler = MetricAutoscaler()
SESSIONS_CAP = 1000
base_loads = [50 + min(35, t * 3) if t <= 12 else 50 + max(0, 35 - (t - 12) * 3) for t in range(30)]
for t, base in enumerate(base_loads):
metrics = make_noisy_metrics(base, base * 0.8, base * 0.9, SESSIONS_CAP * base / 100, SESSIONS_CAP, rng)
result = scaler.tick(metrics)
print(f"t={t:2d} base_load={base:5.1f}% raw={result['raw_score']:.3f} "
f"ema={result['ema_score']:.3f} action={result['action']:9s} instances={result['instances']}")
actions_taken = [h["action"] for h in scaler.history if h["action"] != "hold"]
print(f"Total scaling actions over {len(scaler.history)} ticks: {len(actions_taken)} -> {actions_taken}")
Output (seeded, re-run twice, identical both times):
text
t= 0 base_load= 50.0% raw=0.546 ema=0.546 action=hold instances=4
t= 1 base_load= 53.0% raw=0.525 ema=0.540 action=hold instances=4
t= 2 base_load= 56.0% raw=0.560 ema=0.546 action=hold instances=4
t= 3 base_load= 59.0% raw=0.528 ema=0.541 action=hold instances=4
t= 4 base_load= 62.0% raw=0.618 ema=0.564 action=hold instances=4
t= 5 base_load= 65.0% raw=0.597 ema=0.574 action=hold instances=4
t= 6 base_load= 68.0% raw=0.608 ema=0.584 action=hold instances=4
t= 7 base_load= 71.0% raw=0.602 ema=0.589 action=hold instances=4
t= 8 base_load= 74.0% raw=0.699 ema=0.622 action=hold instances=4
t= 9 base_load= 77.0% raw=0.774 ema=0.668 action=hold instances=4
t=10 base_load= 80.0% raw=0.752 ema=0.693 action=hold instances=4
t=11 base_load= 83.0% raw=0.860 ema=0.743 action=scale_out instances=6
t=12 base_load= 85.0% raw=0.802 ema=0.761 action=hold instances=6
t=13 base_load= 82.0% raw=0.796 ema=0.771 action=hold instances=6
t=14 base_load= 79.0% raw=0.845 ema=0.794 action=hold instances=6
t=15 base_load= 76.0% raw=0.744 ema=0.779 action=scale_out instances=8
t=16 base_load= 73.0% raw=0.684 ema=0.750 action=hold instances=8
t=17 base_load= 70.0% raw=0.640 ema=0.717 action=hold instances=8
t=18 base_load= 67.0% raw=0.646 ema=0.696 action=hold instances=8
t=19 base_load= 64.0% raw=0.548 ema=0.652 action=hold instances=8
t=20 base_load= 61.0% raw=0.489 ema=0.603 action=hold instances=8
t=21 base_load= 58.0% raw=0.529 ema=0.581 action=hold instances=8
t=22 base_load= 55.0% raw=0.571 ema=0.578 action=hold instances=8
t=23 base_load= 52.0% raw=0.455 ema=0.541 action=hold instances=8
t=24 base_load= 50.0% raw=0.465 ema=0.518 action=hold instances=8
t=25 base_load= 50.0% raw=0.478 ema=0.506 action=hold instances=8
t=26 base_load= 50.0% raw=0.478 ema=0.498 action=scale_in instances=7
t=27 base_load= 50.0% raw=0.520 ema=0.505 action=hold instances=7
t=28 base_load= 50.0% raw=0.394 ema=0.471 action=hold instances=7
t=29 base_load= 50.0% raw=0.412 ema=0.454 action=hold instances=7
Total scaling actions over 30 ticks: 3 -> ['scale_out', 'scale_out', 'scale_in']
Key points
- Notice ticks 8 through 10: the raw score jumps around a lot (0.699, 0.774, 0.752) well above the 0.60 target, but no action fires until t=11, because the EMA is still catching up and the dead zone (target plus or minus 0.10) has not yet been cleared by the smoothed value. That gap is exactly what prevents flapping on noisy samples.
- Only 3 scaling actions happen across 30 ticks despite the raw score crossing the target repeatedly, that ratio (few actions relative to how often the raw signal alone would suggest acting) is a reasonable rough health check for a real deployment: if your actual autoscaler's action count per hour is close to its raw-signal-crossing count, your smoothing or dead zone is too weak.
- The scale-out step size grows with how far the EMA is past the dead zone (
1 + int(deviation / dead_zone), capped at max_step), so a large sustained spike scales out faster than a marginal one, while scale-in always moves by exactly 1, deliberately more conservative than scale-out to avoid overshooting into under-capacity.
Complexity
O(1) work per tick (fixed number of metrics, fixed weight lookup), independent of fleet size; this scales fine even at very high tick frequency.
Failure modes this mitigates
- Flapping on noisy single samples: EMA plus dead zone.
- Runaway scaling from a sustained real spike: capped
max_step limits how much a single decision can add.
- Rapid oscillation right after an action, before its effect shows up in the metrics: cooldown.
- One metric spiking alone (say, a network blip) dominating the decision: the weighted composite means no single metric alone can trigger action unless its movement is large enough to shift the blended score past the dead zone.
Edge cases
Cold start (no prior EMA) initializes the EMA to the first raw sample rather than an arbitrary default, so the very first decision is not biased by a warm-up value. A metric that goes missing entirely (a scrape failure) is not handled by this minimal version, in production that needs an explicit fallback (hold action, or fail safe to a conservative default) rather than silently treating a missing value as zero, which would look like low load and wrongly trigger scale-in.