Capacity follows load — and never acts on ignorance.
Four pieces make IG1 elastic: per-instance metrics feed alarms; alarm transitions drive VM autoscaling groups; and Kubernetes worker pools scale through the cluster-autoscaler. One principle runs through all four, and it is the one this page keeps returning to: a system that cannot measure must say so and do nothing — a green light that is green because it is blind is worse than no light at all.
1 · Per-instance metrics
GET /v1/metrics/instances/{server_id} # tier 0 — reading your own load is not privileged
One call returns CPU, memory, disk and NIC figures for one instance, from Nova's diagnostics under your tenant credential — so an instance in another project is refused by Nova itself with its own 404, the same isolation that scopes every proxied compute call. cpu_percent is computed from two samples over a short window (a CPU-time counter is not a utilisation figure), and the raw counters are returned alongside it so your own tooling can derive rates without trusting our rounding. Upstream failures surface verbatim with Nova's status code — never dressed up.
There is deliberately no time-series store behind this endpoint: it samples on demand. If you need history, sample it into your own stack — the endpoint is cheap and tier 0.
2 · Alarms — three states, and the third one is the point
GET /v1/alarms # list yours (tier 0)
POST /v1/alarms # create (tier 1)
GET /v1/alarms/{alarm_id} # one alarm, with its state
DELETE /v1/alarms/{alarm_id} # tier 2 — silencing monitoring is destructive
An alarm is {metric, resource_id, comparison, threshold}, evaluated on an interval by the platform. The metrics you can alarm on are cpu_percent, memory_rss_kb and memory_actual_kb — exactly what the platform can measure as a level. The disk and NIC counters are deliberately excluded: they are cumulative since boot, so a threshold on one would be crossed exactly once in the instance's life and latch forever. That is not a threshold, it is a one-way latch with an alarm's name on it; those metrics become alarmable the day the evaluator differences them into rates.
- Three states: ok, alarm, insufficient_data.
- insufficient_data is honesty, not an error. An alarm whose metric cannot be read has learned nothing about the resource, and the one thing it must never do is say ok. There is no path from an unreadable sample to ok: the absence of a measurement is never evidence of health. If your alarms sit in this state, the metric read is failing — investigate that, not the alarm.
- Events on transitions only. The evaluator runs forever; publishing every pass would drown the topic in steady state that carries no information. An event is emitted when the answer changes — ok → alarm, alarm → ok, either → insufficient_data — and webhook subscribers are delivered to from there.
- Deleting an alarm is tier 2. A leaked operate-tier credential must not be able to silence monitoring.
3 · VM autoscaling groups
Autoscaling groups live on the factory service (https://factory.cloud.ig1.com), which runs the reconcile loop:
POST /v1/asgs # create (tier 1) — admission runs BEFORE the group is stored
GET /v1/asgs # yours, with per-group degraded reasons
GET /v1/asgs/{name} # one group: members, alarms seen, cooldown, drain state
PATCH /v1/asgs/{name} # manual scale: min / max / desired
DELETE /v1/asgs/{name} # drain-first teardown (tier 2)
A group is {name, flavor, image, network, min, max, desired} (bounds 0–60, min ≤ desired ≤ max enforced at create), plus optional user_data, a health check (nova status or a tcp probe with a port), Octavia pool membership (lb_pool_id + lb_pool_member_port, both or neither — half a configuration is refused at the model), and the two alarm hooks: scale_out_alarm / scale_in_alarm, each naming an alarm you own by id or unique name. cooldown_seconds defaults to 300.
| Rule | What it means |
|---|---|
| Transitions, not levels | Scaling acts on an ok → alarm transition only. An alarm that has sat in alarm since before the group existed does not scale it — the group must observe the change. |
| No information moves no capacity | insufficient_data never scales, in either direction — and insufficient_data → alarm does not scale either: the alarm must pass back through ok first, so a monitoring blip cannot manufacture a scaling action out of its own recovery. An alarm state nobody has refreshed recently reads as insufficient_data too: a stale state is a memory, not a measurement. |
| Cooldown freezes observation | During cooldown the engine neither acts nor updates what it has seen, so a transition that lands mid-cooldown fires on the first pass after the window — unless the alarm recovered meanwhile, in which case it correctly never fires. |
| Manual beats alarm | A manual PATCH that lands while the reconciler is mid-pass wins over the pass's alarm decision: a human's explicit number beats a decision computed from a value that no longer exists. |
| Scale-in is drain-first | Removing capacity is: take the member out of the LB pool, wait the drain window, then delete the instance — the window survives factory restarts. If the pool operation fails, nothing is deleted: the group reports it in degraded and the instance keeps serving. Members that are provably dead (Nova ERROR, or past the tcp-probe failure threshold) skip the wait — but still leave the pool first. |
| Admission before actuation | Before any boot — alarm-driven, manual, or a heal — your own Nova quota is checked against the flavor's footprint. A refusal names the numbers and the tier that would grant more; an unreadable quota fails closed for growth. |
| Drift is healed, not punished | Membership is observed, not assumed: Nova is re-listed every pass. A tagged server the group's record does not know is adopted; a member Nova no longer lists is forgotten and the deficit boots a replacement. Deleting a group member by hand just means the group heals. |
| Degraded is a report, not a failure mode | Everything the engine could not do this pass — an unresolvable alarm, a blocked drain, a refused admission, an unreachable Nova — lands in the group's degraded list with the reason and the alternative, visible on GET. A group that cannot observe its scaling signal says so; it never pretends to be healthy. |
Every scaling decision, heal and teardown emits an asg.lifecycle event carrying old/new counts and the trigger (manual | alarm | heal), so the group's history is a topic you can subscribe to, not a log you have to request.
4 · Kubernetes worker autoscaling
PATCH /v1/clusters/{name}/autoscaling # {enabled, min, max} — tier 1, owner only
Worker pools scale through the upstream cluster-autoscaler (Cluster API provider, one instance per workload cluster). Enabling sets the two autoscaler annotations on the pool's MachineDeployment — cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size and …-max-size — and disabling clears them. The annotations are the store: the cluster summary and the autoscaler itself both read them directly, so there is no second record to drift.
- The ceiling is what admission charges. max is what the autoscaler may reach with no further API call, so max — not the current replica count — is checked against your tier's worker-node grant, together with the worst case of your other pools. A breach is a 409 stating all the numbers; an unreachable quota answer is a 503, never a pass.
- Manual scale is refused while autoscaling is enabled. An autoscaled pool has exactly one writer for its replica count; a manual scale racing the autoscaler flaps capacity. The refusal names the way out: disable autoscaling first, or retune min/max here.
- min=0 parks the pool — the autoscaler may drain every worker when nothing is scheduled. Honest caveat, straight from the upstream contract: scale-from-zero additionally needs node-group capacity annotations this platform does not wire yet, so a pool parked at 0 today restarts on the next manual range change. That limitation is recorded, not hidden.
5 · How the pieces compose
The intended loop for a stateless VM tier: a load balancer (created as the one-call composite — networking guide) fronts the group; the group registers and deregisters members in the LB pool as it scales; two alarms on cpu_percent — one gt, one lt — drive scale_out_alarm and scale_in_alarm; and the health check plus drain-first scale-in means membership changes never route traffic to a corpse. For Kubernetes workloads the equivalent loop is pod-level (HPA, inside your cluster) on top of node-level worker autoscaling here.