/ Docs Guides / Versioning MCP ← All guides
Policy

Versioning: one contract, four surfaces, no silent drift.

The served OpenAPI document is the contract. Everything else — the SDKs, the CLI, the MCP tools, the Terraform provider, these docs — is derived from it and gated against it. This page is the rule set that keeps REST ≡ CLI ≡ provider ≡ MCP true wave after wave.

1 · Services: minor-per-wave, independently semver'd

Each service carries its own semver. A program wave (W-numbered workstream, shipped as a numbered phase) that changes a service's contract bumps its minor; fixes and hardening bump the patch. There is no train — a service's version moves only when that service moves.

config/build.yml is the single source for every service version. The playbooks read it, the gitops addon pins are coupled to it, and drift-guards.sh asserts playbook ≡ build.yml ≡ addon pin — so a version that appears anywhere else (this table included) is a copy of that file, never a second opinion. The table below is the build.yml state at this docs release; the file wins if they ever disagree.

→ 0.8.6: get_instance_metrics_history — 137 tools. The live sample answered « now »; the history answers « was it busy at 03:00 » — the same Prometheus the console sparkline reads, with the retention note in the tool text so a truncated window is never reported as idle time. (0.8.6 spent with 0.34.5 — the same bookkeeping truth; 0.8.7 is the committed-tree rebuild.) → 0.8.8: the backup tools — 141 tools. list_volume_backups, create_volume_backup, delete_volume_backup, restore_volume_backup (phase 58, gap register #12): an agent can make the DR copy a snapshot never was, and restore it into a NEW volume. The tool texts say plainly that a snapshot is same-cluster fat-finger protection and a backup is the off-cluster copy. → 0.8.9: restore_database — 142 tools. Phase 58’s second half: an agent can restore a backup-enabled instance into a NEW database (CNPG recovery bootstrap, never in place), with a point-in-time option — the DR story has both its verbs on the agent surface. → 0.8.0: thirteen tools that were in the contract and on no agent surface. The parity gate had asked, since it was written, whether the MCP knew about a DOMAIN — it could not ask whether the MCP could finish a JOB, and the answer for DNS was no on every client at once: get_dns_delegation and import_dns_records were served by the API, carried by all three generated SDKs, and reachable from neither the agent surface, the CLI, nor Terraform. An agent could create a customer’s zone and then had no way to tell them which nameservers to publish — which is the one fact that makes a zone do anything. Also get_dns_zone; list_bucket_objects and presign_bucket_object (so an agent hands out a link instead of pulling object bytes into its context) with get_bucket_versioning / set_bucket_versioning; get_credentials_usage, which is what makes a credential clean-up safe — list_credentials says what exists, this says what anything is still using; get_project_contents, the counterpart to the project delete that is deliberately absent, so an agent can inventory what a deletion would destroy and a human decides with the list in front of them; list_event_topics, because a webhook subscribed to a topic that does not exist is accepted and then silently never fires; list_webhook_deliveries, which separates the three cases that look identical from outside (nothing sent, sent and refused, sent and accepted with the problem downstream); and get_edge_exposure / resync_edge. What stayed OFF the surface is recorded rather than left silent: parity-gate.py grew an OP_EXCEPTIONS table, and rotating a database credential, publishing a status incident, writing an org guardrail and firing a test webhook each carry a written reason there. → 0.8.2: floating-IP tools default to the public pool (public-network, 185.255.84.64/26) instead of whatever Neutron lists first, and the load-balancer publish flow pins its floating IP to the legacy pool — the only one the .75 edge can reach. The tool texts restate the two north-south paths (edge hostname with TLS, zero IPv4; public FIP, direct L4) instead of the retired « internal by design ». (This entry had been sitting in the Secrets broker’s row since it was written — a merge casualty, moved here 2026-08-29.) → 0.8.3: create_database gains phase 52’s durability surface — replicas (1–3, PostgreSQL) and the backup_enabled / backup_retention_days / backup_schedule block (WAL archiving plus a scheduled base backup into the platform bucket, isolated per tenant). An agent can now build the durable instance it could previously only describe; the same 135 tools, and delete_database’s warning now says plainly that a backup-enabled instance leaves a restorable copy behind while anything else leaves nothing. → 0.8.4: resize_database gains the replicas knob (api 0.33.2) — in place, in either direction — and its docstring says why an agent must never recreate an instance to change its replica count: that destroys the data the count exists to protect. → 0.8.5: get_kubernetes_cluster — 136 tools. The detail an agent needs to answer « what is my load balancer’s address » (phase 53): the factory’s load_balancers block, with the docstring saying an empty block means no LB Services exist, not that the read fa → 0.8.11: the floating-IP theft guard fails CLOSED. It read the address before associating it — Neutron silently detaches whatever holds a FIP — but ran that check under if status_code == 200 with no other branch, so a 404, a 500 or an expired token skipped it and associated anyway: the guard absent in exactly the conditions that cause a retry. → 0.9.0: the repairs from a sweep that called all 149 tools against the live cloud. import_dns_records raised on EVERY call — it read the API’s plan (a LIST of per-record entries) as a dict keyed by verdict — and with commit=True it raised after the records were in Designate, so the caller was told the import failed while the zone served the new answers. allocate_floating_ip now refuses a pool the tenant’s router cannot route: allocation and association are two Neutron calls and only the second checks that the address’s network matches the router’s gateway, so the hardcoded default handed out addresses that were billed, held against quota and attachable to nothing — the exact state associate_floating_ip’s own docstring warns about. get_cluster_kubeconfig moves to the factory’s owner-gated endpoint (the raw k8s proxy it used moved to the operator plane, so it 404’d for every customer); the three workload-identity tools ride factory 0.6.0’s new routes instead of endpoints that never existed. run_workflow waits for a VM to be ACTIVE before the step that uses it — its own canonical three-step run failed at step 3 every time and rolled back two resources that were fine. delete_vm reads the instance back instead of printing « deleted » on the accept, which it did five times over five instances that stayed running. snapshot_vm reads the image id out of the Location header. And resync_edge is removed — 148 tools: admin-only, so it answered every customer credential 403 from the day it shipped, and a tool that can only refuse reads as a capability. → 0.9.1: snapshot_vm resolves the new image id by asking Glance for the snapshot whose instance_uuid is this server. 0.9.0 read Nova’s Location header, which the gateway does not forward — the OpenStack passthrough passes only Content-Type — so it correctly reported that it had lost the id rather than printing one. The pair (name, server) is an identity where the name alone is not. → 0.9.2: delete_workload_identity strips the known wi-ig1-kaas-<project>- prefix instead of splitting on the last hyphen, which turned a cluster called web-prod into prod — the wrong cluster, of the same tenant. An id it cannot parse is refused, not half-read. → 0.9.3: delete_load_balancer waits out Octavia's VIP-port release before calling its managed security group litter. One attempt could never hang — and lost the race every time, so every delete left an orphan group behind. Bounded retry, only on the 409 that is the race. → 0.9.4: create_database and resize_database stop refusing sizes above 100 GB before a request is even sent — their published schema still carried the old limit. The platform answers now, and its refusal states the limits and the trade.
SurfaceVersion (config/build.yml)Wave trace
IG1 Cloud API (gateway)0.38.0phase 16 → W4 credentials (26) → W6 data domains (32) → W7 monitoring/status (33) → the parity programme's 0.10.x line: quotas/signups, audit + metrics + alarms, LBaaS + DNS proxy roots, edge exposures, database day-2, object browser → 0.11.0: the key-manager (Barbican) proxy root, the public tier catalogue and the admin tier-tag backfill → 0.11.1: the tier-change RGW quota write fixed (query-string args; tentacle 501d a body) and its 502 made honest (gotcha 149) → 0.11.2–0.11.4: below-usage tier moves — Nova and Neutron forced, Cinder keys held at usage and reported; /v1/quotas carries the object gateway's live ceiling + usage + matches_tier → 0.11.5: every tenant router is uplinked to the external network at onboard (SNAT on; the repair call converges older tenants — gotcha 151, the shared external L2) → 0.11.6: a failed token introspection now says in the LOG which of the three failures it was — TLS trust, the API's own introspection credentials, or a non-JSON answer (the 401 body is unchanged and identical for all three, so an unauthenticated caller still learns nothing — gotcha 150) → 0.11.7: an RGW failure can never answer 2xx again — an unparseable body on a 200 became an HTTP 200 whose payload was an error object, and a tenant with no S3 user yet now lists zero buckets instead (gotcha 160) → 0.12.0: a platform-admin credential is scoped to no tenant, and every endpoint that resolves a tenant from the credential now says so with a 409 instead of a 404 "no such project" (/v1/account) or a 503 that read like an outage (/v1/object-storage/credentials, /v1/databases and the rest of those two families); GET /v1/account gains the admin ?project= override that /v1/quotas already had, and a managed-database resize whose tier cannot be read now fails closed rather than falling back to the namespace cap — which had silently re-granted 100 GB to suspended tenants → 0.12.1: readiness stopped evicting the gateway from its own Service — /health/ready is the kubelet’s readinessProbe and it returned 503 whenever a 1 s probe of Keystone or the k8s apiserver did not answer, so a busy control plane took a healthy pod out of the Endpoints and concentrated the load on whatever replica was left (gotcha 175). Upstream state is still reported in the body; only the verdict changed → 0.12.4: the reference stopped showing null as every response sample. No route here declares a response_model, so FastAPI emitted an empty response schema and Redoc rendered null for all 81 operations — a customer could not learn the shape of a single response without calling it. Declaring models was the wrong fix: response_model makes FastAPI filter the outgoing payload, and on the catch-all proxies (/v1/openstack/{path}, /v1/kubernetes/{path}) the body is whatever the upstream returned, so a wrong model would silently truncate real data. Instead the generated document is enriched with examples captured from the running cloud (scripts/capture-openapi-examples.py) — 48 of 81 operations now carry a real sample, the remaining 33 being writes nobody should execute against production to make documentation. Payloads are untouched → 0.13.0: the fleet digest behind /v1/monitoring/fleet and /alarms was rewritten onto Kolla’s Prometheus and Alertmanager. It used to scrape each node’s Netdata agent on :19999 — a port the 2026-07-31 security review had recorded as loopback-bound on all ten hosts, and which two nodes never ran at all — so the Monitoring page reported the entire fleet unreachable while the fleet was healthy. The node list is now Prometheus’ own scrape targets rather than an environment variable that had already gone stale, node names come from the agents’ uname data instead of bare addresses, and a monitoring source that cannot be READ now answers 502 carrying the upstream’s words instead of an empty fleet that reads as “you have no nodes”. /v1/status degrades its fleet block to unknown rather than failing or claiming zero. → 0.13.1: Alertmanager has its own basic-auth credential, separate from Prometheus’. 0.13.0 sent the Prometheus pair to both and got a 401 from the alarm source, so the Alarms panel reported a refusal above a fleet table that was reading perfectly — the failure was visible precisely because the two sources fail independently and neither is allowed to fake the other’s data. → 0.13.2: the fleet Alarms panel reads BOTH alerting systems. It had been reading only the one that ships no alerting rules, so it published “everything is nominal” while the platform’s own Alertmanager held a critical alert saying its alert deliveries were failing. Sources are a list now, merged, each alarm tagged with the system that raised it; one unreachable source degrades, all of them failing refuses. The always-firing heartbeat alert is filtered, and alerts about a namespace rather than a node no longer render a blank Host. → 0.13.3: removing an org member now revokes the membership record that the invite wrote. DELETE /v1/org/users/{id} deleted the person’s login and left their membership record live, and because the credential ceiling counts every live record a project holds — members included — each removed colleague permanently consumed one of the account’s credential slots. An account on the Discovery tier that invited and removed five people could no longer create an API key, and the refusal named a quota they were not using. → 0.14.1: the gateway proxies the notification integrations surface. The events service has no Ingress, so a route the gateway does not carry does not exist for a browser. → 0.15.0: the container registry becomes SELF-SERVICE. This gateway is now the registry’s token service: it signs short-lived tokens naming exactly the repositories a caller may touch, which is what lets ONE shared registry serve many tenants without a second product. A tenant owns the namespace under their project id and nothing else — the catalogue scope, which would list every customer’s repositories, is never granted; the repository list is filtered server-side instead. The tier decides the verbs: pull at 0, push at 1, delete at 2. The login is the credential you already have, so revoking it revokes registry access with it. → 0.16.0: the gateway is built once and run twice. On the customer plane the cross-tenant endpoints do not EXIST — they are absent, not refused — and the operator plane lives only on the VPN-reachable admin address. Adds /v1/access/policy: your rule for whether IG1 may touch your data, your approvals and revocations, and the full log of what we looked at and why. → 0.17.0: tenant DNS becomes a product. /v1/dns/zones replaces the raw OpenStack passthrough, which could forward but could not qualify, refuse or count: a short name like www was rejected outright, nothing stopped a tenant claiming a zone inside IG1’s own namespace, and no ceiling applied. The gateway now qualifies a short name against its zone in ONE place (so www, www.z.fr and www.z.fr. cannot become three records), refuses a zone inside the platform’s namespace on the label boundary — notig1.com passes, evilcloud.ig1.com does not — enforces the tenant’s zone ceiling, and forwards Designate’s own refusal verbatim. A recordset UPDATES in place (PATCH): delete-then-create would take the name out of resolution between the two calls, and resolvers cache that absence. → 0.17.2: POST /v1/dns/zones answered 500 — the first call a customer makes on the new surface. The route read a field the auth context does not have, so the handler raised before it reached Designate. Fifty-two offline tests were green throughout, because each exercises the zone-name check as a pure function and supplies that flag itself: a pure-function test proves the function and cannot prove its caller. The live gate caught it on its first run, and the offline suite now drives the real application with a real auth context. → 0.18.0: the rate limit stops bounding a TOKEN and starts bounding a TENANT. Keyed by sha256(Authorization), it failed three ways and all of them favoured the caller: a rotated token was a brand-new empty window, so a client refreshing every 50 s was never limited; a second credential was a second budget, so the ceiling scaled with the number of keys the customer chose to create; and every tenant got the same number, so the limit could neither be sold nor protect the platform from the tenant most able to saturate it. The ceiling now comes from the tenant's commercial tier — api_rpm is a dimension of the quota catalogue beside cores and buckets, because a rate limit is sellable capacity — and the RateLimit-* headers report the tenant's real remaining rather than a coarse per-token figure. → 0.18.1: the gateway is ready for TLS on the OpenStack endpoints (review finding F6). Two things, and the first would have hidden the second: every OS_*_URL was derived at IMPORT from a module constant rather than from the configured auth URL, so moving the cloud to TLS would have left compute, network, volume, image, DNS, load-balancer and key-manager POSTing over plain http — and every call would still have succeeded. Siblings follow the effective auth URL now; an explicit override still wins. And the gateway shares one HTTP client, so startup REFUSES two different CA bundles rather than silently honouring one. → 0.18.2: listing your organisation’s users decides ownership BEFORE it builds a record. It used to read every person in the shared directory, assemble a full entry for each — name and address included — and filter on the last line. The answer was always correct; the arrangement meant everyone else’s details passed through the process on every call, and it capped the underlying search at a thousand entries, so a large enough directory would have quietly stopped showing an administrator their own people. (0.17.0 is burned: its image was built and imported 3/3 before version= in main.py had moved off 0.16.0, so the tag would have served the previous version — → 0.18.3: no new code — this release exists because two security fixes had been written and were not running. The pin said 0.18.2, the image tagged 0.18.2 was genuinely that image, and four files had changed since it was built; the gate that compares the running image to the source is what caught it. What starts running here: a read-only credential could revoke production, and the cross-tenant quota routes (a project’s tier, the tier backfill, and the ?project= override on quotas) were mounted on the customer plane as well as the operator plane — the in-handler platform-admin checks were correct and did hold, but on the customer plane those paths now do not exist at all. Also: credential “usage” counted the mint as a call and merged a per-process ring, so the figure depended on which replica answered.) → 0.19.2: the operator console gains a whole-cloud rollup. GET /v1/monitoring/overview answers « is the cloud healthy » before anyone parses a node table: one ok/total pair per subsystem — scrape targets, nodes, OpenStack services, Nova and Neutron agents, the Galera cluster, HAProxy backends and the endpoint probes — plus placement, certificate and instance gauges. Every group carries a numerator AND a denominator, and the denominator decides whether it was read at all: a metric that has vanished answers with an empty vector, and reporting that as 0/0 would paint a green tile over a dead exporter. Ceph is deliberately absent for the same reason — nothing scrapes it here. Alarms also stopped hiding a dead source: one unreachable alertmanager of two used to arrive as « no alarms ». (0.19.0 and 0.19.1 belong to a concurrent session and are spent.) → 0.19.3: the published API examples stopped describing a cloud that does not exist. This reference is generated from the live spec, and its fleet sample carried ten UNREACHABLE nodes blaming a monitoring backend retired in 0.13.0, while the status sample showed nothing reachable underneath a green « ok ». The samples now match what the gateway actually returns, and a test derives the check from the handler itself rather than grepping for the old name — so the next backend change cannot quietly leave this page behind. → 0.21.0: the customer plane stops describing the estate. GET /v1/status carried a fleet block — how many machines this cloud runs, how many are answering, how many alarms are raised — and every row’s probe text: « nova reachable (HTTP 401) », « management k8s reachable », « not configured (ig1-rgw-admin Secret absent) ». That block survived the 2026-08-12 disclosure fix because the line drawn then was « no hostnames, no alarm text » — the most identifying fields — rather than « nothing about the estate ». Fleet size is commercial information, and « 3 of 10 unreachable » times an attack as precisely as a node table does. Rows keep their honest ok/degraded/down; the detail is now the product’s sentence, and the probe’s own words moved to the operator plane. /health/ready lost its component map on the customer plane for the same reason — the Ingress routes path: /, so that map named Keystone and Kubernetes, and reported their reachability, to anyone who dialled the front door. And this reference’s own samples were de-identified: they had been publishing ten physical node hostnames, two tenants’ project names, staff and customer email addresses and the platform’s remaining sellable capacity — captured from admin-gated endpoints whose gates all held. The gate was never the leak; the documentation of the gated endpoint was. Added, not removed: /v1/whoami now carries user_name and user_email, so a caller holding nothing but a token can name the person and the account it belongs to. → 0.21.1: 0.21.0 is burned — it was built by the deploy playbook, which had not been taught to stage config/services.yml into the build context after that file became a fatal startup dependency, so every pod from it died in its lifespan with « service catalogue missing ». No customer impact: the rolling update kept the 0.20.1 pods serving throughout, which is also why the deployment reported the new tag with two Ready replicas the entire time. → 0.21.2: 0.21.1 is burned as well — two sessions took the same next tag off the same crashloop within the hour and both built an image under it, from different trees. Rebuilt from the merged tree, which carries both repairs: the peer’s rebuild and this branch’s fix to the deploy playbook that had never staged the catalogue at all. → 0.21.3: this reference’s samples re-captured from the running cloud through the new de-identification pass — the identity document now shows user_name and user_email, the status sample has no fleet block, and seven customer surfaces that had no sample at all get one (consent, tenant DNS, the registry pair, event topics). → 0.21.4: the identity document gains credential_label. Measured against the live cloud, user_name and user_email both come back EMPTY for the credential type agents actually use — a machine user asserts neither — so the answer to « whose key is this » was an 18-digit id again. The label is the name the customer typed when they minted it. → 0.21.5: the operator plane gains a CAPACITY surface, and cross-tenant listing. GET /v1/capacity answers « can I sell another 32 vCPU » — physical inventory, the SUM of every tenant’s quota, and what is actually running, per resource class and per tenant. The middle number was computed nowhere, so oversubscription was invisible until somebody added quotas up by hand. And Nova and Cinder scope every list to the token’s project, so a platform caller was reading the ADMIN project’s resources: the operator console drew four zeros over a fleet running six machines. all_tenants=1 now travels for platform callers on GET lists only — never for a customer, which is the isolation boundary and is asserted from the customer’s side. 84 paths. → 0.24.2: custom domains on the .75 edge. An exposure may carry a name the CUSTOMER bought — claimed globally (first claim wins across tenants, because a real domain has no tenant slug to disambiguate it), proved by a _ig1-challenge TXT record, and routed only once verified: an unverified claim is a reservation, and serving one would hand traffic to whoever guessed a name first. Certificates are per-domain and SNI-selected from a crt-list. The public front door at 10.57.8.77 is rendered unconditionally, so going public assigns an address rather than introducing a request path nobody has run. Two follow-ups found by deploying it: the store CACHES the rendered configuration and refreshes it only when the exposures change, so upgrading the API upgraded the RENDERER without reaching the edge — the playbook ran green and the controllers installed the previous release’s document. update_store now self-heals a stale render on any write attempt (converging, so a refused request arriving forever still leaves the store byte-identical), and POST /v1/edge/resync is that attempt made deliberate for after a deploy. Neither bumps the generation: the routing table is unchanged, so no exposure is told it is behind. → 0.25.0: phase 48 — the OU tree, and guardrails that inherit down it. /v1/org/ous makes the dotted path the parentage itself (prod.eu-west is a child of prod), so the hierarchy has exactly one representation and there is no parent pointer that can disagree with it. The policy expressions compose the way AWS composes them, for AWS’s reasons: max_tier by MINIMUM, because a leaf that could widen what an ancestor set makes every inherited guardrail advisory; denies accumulate with no allow to cancel one; and ABSENCE CHANGES NOTHING, which is what makes this shippable into a cloud whose accounts hold no policies at all. /v1/org/effective is what makes the rest safe to have — a ceiling a caller cannot see is indistinguishable from a bug, so it reports the tier they would hold WITHOUT the org policy beside the one they do hold and names the policies that produced the difference, at tier 0 and no role: knowing the rules you are subject to is not a privilege. → 0.26.0: a project can be DELETED and an account CLOSED. Phase 47 shipped no DELETE and wrote down why — removing a project destroys everything in it and stays an operator action — which is a defensible default and a poor product: a customer who cannot close what they opened opens a ticket, and the operator then performs the same destruction with less context and no consent trail. GET /v1/projects/{id}/contents is a read-only survey, because « this destroys 3 instances, 2 volumes and 1 bucket » is a consent and « are you sure? » is a formality; DELETE /v1/projects/{id} takes the project’s own name as confirm; POST /v1/account/close takes the account number plus a separate acknowledgement that every project goes with it. Deleting one project is tier 2 + an owning role; closing the account is the Propriétaire alone, because an Administrateur holds billing read-only. The purge re-checks the owner field of every object an admin list returned — an ignored server-side filter returns the whole cloud, and the next line is a delete loop — and refuses to remove the Keystone project unless every service both read cleanly and left nothing behind. → 0.26.2: two defects 0.26.0 shipped, both found by validating against the LIVE cloud rather than the tree — the CAPI-cluster read answered 401 then 403 (no Authorization header, and no cluster.x-k8s.io rule in the ClusterRole), so the step reported unreadable and every project deletion refused; and the owning-role guard, written for creation and reused for deletion, told a customer refused a DELETE that « creating a project is reserved to… ». The new RBAC rule is get+list only. (0.26.1 is spent — built earlier the same session, before a rebase onto a peer's landed release.) → 0.26.3: creating a project silently MOVED the caller into it, and on a timer. A TENANT_MAP-resolved caller holds no credential-store row; POST /v1/projects gives them their first, and resolution prefers the store over the map — so the project just created became the one you were in, but only once the ~30 s per-replica cache expired. Two identical calls a minute apart resolved to different projects. Flipping the precedence would have been a privilege regression (map entries are hard-coded tier 2, so a tier-0 credential would resolve as tier 2), so instead the caller’s current project is written as an entitlement first — genuinely the earlier row. Found by running the phase-49 live gate a SECOND time; the first run was green. → 0.26.4: a purge that GIVES UP now takes its ig1-deleting tag back off. It never did — the tag was add-only — so a refused deletion left the project flagged for good: the selector rendered it disabled as « Suppression en cours » and every retry 409’d once the account was back to one project, so the deletion could neither finish nor be abandoned. The refusal path is the branch where the project SURVIVES INTACT, so marking it as leaving there was not unhelpful but false. → 0.26.5: and the way back for a mark that path never reached. 0.26.4 clears the flag when a purge REFUSES; a gateway killed between the tag and the refusal still leaves it on, and nothing else removes it. POST /v1/tenants/{id}/deletion/cancel (operator plane, admin) drops the mark, converges the tenant network a partial purge removed — clearing the flag alone hands back a project with nothing to launch into — and reports how many credentials can still reach it, because a deletion that got past the purge had already dropped them. → 0.27.0: the identity half of a deletion is suspended, not dropped, until the project is actually gone. The credential rows used to be removed before the Keystone delete, so a failure at that last step destroyed them with no record of who had held them — the cancel verb could only report entitlements: 0 and leave a human to guess who deserved access back. Both hyperscalers avoid that guess by never destroying access first (a closed AWS account keeps its identities for the 90-day post-closure period; a GCP project sits in DELETE_REQUESTED for 30 days with its IAM bindings intact). Order is now suspend → remove the project → drop, a cancel restores exactly the rows that existed and reports restored_credentials, and a suspended credential says so at the door instead of answering « revoked » and sending its holder to mint a new one. → 0.27.1: the public frontends bind the public address itself — EDGE_PUBLIC_BIND_IPS (env IG1_API_EDGE_PUBLIC_BIND_IPS) adds a bind per configured address to ft_ig1_public_443/_80, because the network team attaches the public address to the controllers instead of DNAT-ing to .77 (2026-08-25: 185.255.84.178 on VLAN 104, direct attach); unset renders the pre-direct-attach document byte for byte. (0.26.6 is SPENT: built and imported 3/3 carrying this same change, then 0.27.0 landed mid-flight — pinning it would have regressed the suspend-not-drop deletion under a lower version, so the change re-ships here from the merged tree.) → 0.28.0: closing an account now holds its credentials the same way deleting a single project does. Closure does not loop over the project verb — it runs its own copy of the sequence, which still dropped the credential rows before the Keystone delete. Closure removes one project at a time and stops at the first it cannot empty, and its whole contract is that a failure leaves the customer able to sign in and finish; a drop in that window destroys the access of a project that then survives. Now suspend → remove → drop, and a refusal lifts any hold an earlier attempt left behind. → 0.29.0: closing an account is reversible for 30 days. It used to purge every project and remove the people immediately; it now marks each project, suspends every credential and destroys nothing, with POST /v1/tenants/accounts/{n}/reopen giving it all back and …/finalize doing the destroying once the window expires (early only with a stated reason — an erasure request is one). Why 30 and not AWS’s 90: checked as a legal question first, and no law sets a floor — GDPR prescribes no retention period and Article 5(1)(e) argues for the shorter one, while the 10-year retentions people remember attach to VAT and accounting records, a different dataset under a legal-obligation basis that survives an erasure request. So the window is chosen on cost: during it the resources still exist, still occupy capacity and still bill. Configurable; 0 restores the old immediate behaviour, still tested. → 0.29.1: the ACME lane stops 301-ing its own challenges — http-request redirect evaluates before backend selection in haproxy, so the unconditional redirect fired on the challenge path too and Let's Encrypt met the SNI router instead of the solver on every authorization; the redirect now carries if !acme_challenge. Found by the first live ACME run, minutes after the flip (gotcha 302's fourth face). → 0.29.2: the identity issuer moves to its REAL name — zitadel.cloud.ig1.com, Let's Encrypt on the sign-in page, sign-in green from a phone with no CA installed. The renderer gains exact-match platform SNIs (PUBLIC_PLATFORM_EXACT_SNIS) so the APEX cloud.ig1.com reaches the portal's 301 without opening the -m end label-boundary hole, and the public deny list learns the operator tools' real names (console/grafana/argocd.cloud.ig1.com — records point at the VPN-only admin VIP, the door refuses the SNI). → 0.29.3: the operator tools' real names join the deny DEFAULT, not only the deployment env — guard 33 reads the constant because safe-with-no-config is the property; the byte-pinned fixture follows. (0.29.2 spent: built before the guard spoke.) → 0.30.0: the signup queue and the re-tier button stop disagreeing about how many tenants exist. Both gates spend the same sellable budget and the Locataires page says in as many words that they share one arithmetic, but only re-tier was moved onto tenants.live_tier_counts (the Keystone project tag) when phase 36 introduced it; the queue kept counting its own APPROVED ROWS for twelve days. Measured on the lab: /v1/signups/capacity reported {"discovery": 7} against the two projects /v1/tenants actually held. Not the undercount the docstring promised was safe — an OVERCOUNT of five, because signups.py creates tenants and nothing deletes them (gotcha 276), so every tenant torn down since 2026-08-12 still charged against the budget. The queue was refusing capacity the cluster had, and a real applicant would have paid for it first. → 0.30.1: every name finishes the move — the registry the customers docker-login is registry.cloud.ig1.com (host + token service + the auth realm an internet client is bounced to), and the public deny DEFAULT learns the last two operator names (api-admin/mailpit.cloud.ig1.com — published on the admin listeners by the same route-table edit that refuses them on every customer door). → 0.30.2: onboarding writes down the tier it just spent, and a taken organisation name stops being a dead end. The first two real applicants through the waitlist found both: the new tenant listed as Non étiqueté because create_tenant set quotas FROM the tier and never tagged the project with it — so live_tier_counts, the census both admission gates spend against, counted it at DEFAULT_TIER regardless of what was sold (0.30.0 fixed the reader while the writer stayed silent) — and the second applicant, a colleague who typed the same organisation name, met a raw Keystone 409 rendered into a toast as braces and quotes, with no route forward: nothing about retrying changes the name. The clash is now detected BEFORE provisioning and answered with an instruction — ask that account's owner to invite them — while tenant_name lets an operator accept a genuinely different organisation that shares a name. No auto-generated ig1-2: inventing a name hides a collision the operator must see. POST /v1/signup tells a colleague the same thing at submission rather than queueing a request that can never be approved — but only when their email DOMAIN is already inside that tenant, because a company name is guessable and a bare \u201cthat org exists\u201d would hand any anonymous caller the customer list. → 0.32.0: a customer can move a domain bought anywhere onto IG1 without a support ticket — GET /v1/dns/zones/{id}/delegation compares the nameservers the zone publishes (read from its own apex NS RRset, never a constant, so the second one appears by itself the day the pool serves it) against the ones the internet names for it, and POST /v1/dns/zones/{id}/import turns a pasted zone file or registrar table into recordsets, dry-run by default. A delegation pointing at dead nameservers SERVFAILs, and is reported as unreachable rather than as « no records » — which would have told a customer who got it half right that they had not started. → 0.32.1: the operator finally HEARS about a new application. The signup queue has mailed the operator since it was built (“how do I approve if I never hear about it?”) — into Mailpit, a CATCHER, so not one of those mails has ever been delivered and the queue was silently poll-only again. Now sent through Google Workspace’s SMTP relay, IP-allowlisted on this cluster’s egress so there is no credential to store; _send_mail_blocking gained TLS at the same time, because it could only ever speak plaintext and pointing it at a real relay produced a dead connection and a notification that silently never arrived — the same failure as the catcher, one layer down. IG1_API_PORTAL_BASE also stopped naming the pre-flip nip.io address. → 0.32.3: an existing tenant router is never re-gatewayed by configuration (phase 50 — ensure_tenant_network sets a gateway only on routers that have none, so flipping TENANT_EXTERNAL_NETWORK to public-network cannot migrate a tenant and strand its floating IPs; migration is the explicit runbook), and new tenants uplink to the public pool by default. → 0.32.4: the public door's ingress backend is a health-checked SET, not one address — IG1_API_EDGE_INGRESS_BACKEND takes a comma list and the rendered haproxy gets one server per entry, so losing any single k3s mgmt node no longer takes portal/api.cloud.ig1.com with it (the 2026-08-28 outage and its root cause: OPERATIONS gotcha 355). → 0.32.5: the operator rollup learns the bus. GET /v1/monitoring/overview gains two RabbitMQ tiles, made possible by the per-queue metrics phase 51 turned on (gotcha 359): an ok/total pair that calls a queue healthy when it is consumed OR shallow — and NAMES any queue filling past the fresh-trickle window with zero readers, the exact signature of the notifications_designate.error incident — and a gauge reading the deepest queue in the cluster, the number gotcha 357's 383k would have moved. As everywhere on this rollup, a denominator that vanishes reports the tile unread rather than painting 0/0 over a dead exporter. → 0.33.0: managed Postgres gains real durability (phase 52 — gap register #2). POST /v1/databases takes replicas (1–3 — synchronous replication with operator failover) and a backup block (daily base backups plus continuous WAL archiving to the platform's object store — point-in-time recovery): the same CNPG barman-cloud pattern the platform's own identity cluster has run since 2026-08-08, exposed per tenant with an archive path no other tenant can write into. The storage budget now counts what is actually consumed — a database's footprint is size × replicas, on create AND on resize — and the budget gate runs on create for the first time (it was resize-only, so a first instance with replicas could overshoot before it existed). Deleting an instance kills its schedule first and leaves the archive; that is the point of DR. Also: the gateway block on /v1/edge/exposures finally says the customer-domain TLS sentence out loud (your own domain gets a publicly-trusted certificate automatically). → 0.33.1: the durability the release above sells now actually starts. The first backup-enabled fixture sat at « Setting up primary » for over an hour with no pod object at all: the tenant namespace’s ResourceQuota covers cpu/memory, so the apiserver refuses any pod whose containers declare no resources — and the barman-cloud plugin’s INJECTED init container arrives with resources: {} that neither the tenant nor the Cluster spec can fill. Every tenant namespace now converges a LimitRange (ig1-container-defaults, 10m/32Mi requests) beside its quota — the admission contract that lets an injected sidecar in — converged on drift like the network policies, and pinned by the RBAC grant, the offline suite and the live phase-52 gate alike. → 0.33.2: a database’s replica count changes IN PLACE. PATCH /v1/databases/{engine}/{name} takes replicas beside size_gb — CNPG bootstraps or retires an instance live, in either direction, so nothing ever has to replace the cluster (and destroy the data the count exists to protect) for a replica change. The budget gate sees the final shape (size × replicas), a no-op count is refused like a no-op size, and Kafka answers 422 — its single dual-role pool has no second shape to scale to. → 0.33.4: the backups the releases above sell can now actually LEAVE the cluster. Every Backup CR on the first live fixture failed with the same rpc error: Unavailable / EOF while the database itself read perfectly healthy: the tenant namespace’s NetworkPolicies are default-deny on egress, and the barman sidecar’s first act — barman-cloud-check-wal-archive against the RGW endpoint — is a dial no rule allowed. The database needs no such route; only its durability does, which is why nothing else ever met it. Every tenant namespace now converges allow-backup-egress beside the other policies, and the endpoint is ONE constant both the policy and the ObjectStore render read — a rule that allows a different address than the store uses is the bug that shape exists to make impossible. (0.33.3 is spent: allocated by a build that was killed before it ran, when the same command’s git checkout was caught discarding this very fix — the second uncommitted-work loss of the day, after portal 0.39.3.) → 0.33.5: SPENT — built and imported with the old LimitRange constants inside. The sizing edit sat uncommitted in the shared checkout, a git stash pop never applied it (and said nothing about not applying it), and the build archived the clean tree. Third uncommitted-work loss of the day, after portal 0.39.3 and the netpol fix — the rule from gotcha 364 is now enforced here: the payload commits to a worktree BEFORE any version cut. → 0.33.6: the LimitRange the fixture was rescued with no longer OOMKills the sidecar it exists to admit. 100m/64Mi passes the wal-archive check and dies mid-backup (exitCode 137; the Backup CR’s error reading from server: EOF is the plugin connection dying with the process) — the gzip base-backup stream needs more. The defaults are 50m/128Mi requests, 500m/512Mi limits now, and the suite reads the constants instead of retyping them, so the next sizing is a one-line change. → 0.33.7: the backup block reports on-demand backups too. last_successful_at read only the ScheduledBackup’s own status, so a customer who took a base backup on demand saw null forever — found by the first green run of the phase-52 live gate, which completed its immediate Backup and read null back. The detail now names the newest COMPLETED Backup CR for the cluster, any origin. → 0.33.8: and it actually finds them. A Backup CR carries NO labels in CNPG 0.29 (--show-labels is <none>), so the selector the previous release listed by matched nothing and the answer stayed null with the backup completed beside it. The cluster link is now filtered client-side from spec.cluster.name. → 0.33.9: and the completed time is read from the field that exists. CNPG 0.29’s Backup status carries startedAt/stoppedAt — no completionTimestamp — so the completed list was always empty and the detail still answered null with the backup finished beside it. Third read of the same block, third live-only miss, each one found by the phase-52 gate rather than by a customer. → 0.33.10: the PITR window is read from where it actually lives. With the plugin architecture the Cluster’s own firstRecoverabilityPoint NEVER populates — the identity cluster itself has 14 days of completed dailies and no such field. The plugin writes the window under status.serverRecoveryWindow on the ObjectStore; the detail reads it there, and the phase-52 gate now asserts the API’s product field instead of the operator internals that never change. → 0.34.0: the load-balancer quota is a real row (phase 53). The tier table gains octavia.load_balancer (discovery 0, standard 2, extension 5), every tier write PUTs it to Octavia’s own quota API — the service enforces it, and a discovery tenant’s refusal lands on the Service’s events instead of a silent pend — and /v1/quotas reads it back as a load_balancers block. The factory’s 0.5.2 turns the tenant OCCM’s controller on in the same wave, so a type=LoadBalancer Service now gets an address or a readable refusal, never a pend. → 0.34.1: the audit trail survives a redeploy (phase 54, gap register #7). GET /v1/audit used to read a per-pod ring of 2,000 rows — two replicas, half the trail each, and a rollout reset the answer to zero. It now answers from the ring only when the ring provably covers the requested window; otherwise it reads the durable api.audit topic through the events service’s new history endpoint, and every response says which source answered (source: ring|durable). The tenant filter runs at BOTH layers — neither trusts the other’s — and a dead events service degrades honestly: ring rows plus a degraded string, never a silent empty window. → 0.34.2: the managed-database storage ceiling moves 100 → 250 GB (phase 55, gap register #3), with the derivation measured, not asserted (docs/capacity-measured.md, 2026-08-30: 250 × 3 replicas = 750 GiB, 26 % of the 2,864.3 GiB MAX_AVAIL, inside the stored-bytes stop at 44 %). The binding change is the namespace: the tenant ResourceQuota’s storage ceiling is now rendered from the tier’s storage footprint (discovery 100, standard 400, extension 1,700) and drift-converged on every database write — before, a flat 100 Gi bound every tier below its grant, so a standard tenant with 400 GB of storage could not put more than 100 GB into databases. The budget gate and the quota now share one derivation, and a tier-lookup outage fails closed at the pre-55 default. The quota’s PVC count moved 4 → 8 in the same change (a 3-replica database plus its restore drill is six). → 0.34.3: the proxy becomes the billing ledger’s publisher (phase 56). Every successful forwarded write that changes a resource’s billable life — a power action, a server or volume delete, a volume or floating-IP create — emits its lifecycle event fire-and-forget: a refused action never bills, a dead bus never changes what a write already answered, and nothing secret travels. The proxy is the only path a tenant can act through, which is what makes its publish authoritative; anything that bypasses it is the billing reconciler’s. (0.34.3 is spent on a bookkeeping truth: its image predated the commit of its own version constant, so the source-vs-pin guard could not prove content — 0.34.4 is the same bytes built from the committed tree.) → 0.34.5: tenant metrics history (phase 57, gap register #8). GET /v1/metrics/instances/{id}/history answers « what did it look like an hour ago » from kolla’s Prometheus — the libvirt exporter had been scraped all along; what was missing was the JOIN (openstack_nova_server_status{id,tenant_id} → instance_libvirt → the series), so ownership is settled before any range query and a foreign server 404s. The offline suite caught a real bug on the way: since=0 was silently defaulted to the last hour. (0.34.5 and 0.8.6 are spent on the bookkeeping truth again — the constants commit landed after their builds; 0.34.6 and 0.8.7 are the committed-tree rebuilds, and the rule now has two receipts.) → 0.34.7: the tier table gains the backup quotas (phase 58): backup_gigabytes = gigabytes (one full copy of the estate), backups = 2 × volumes (one copy plus rotation). A quota table is inside the api’s build context, so the quota change is an api release — the guard caught it before the tree claimed otherwise. → 0.35.0: POST /v1/databases/{engine}/{name}/restore — phase 58’s second half. Restore a backup-enabled instance into a NEW cluster (CNPG recovery bootstrap from the source’s own archive, never in place), with to_point_in_time mapping to the recovery target. The source without an archive 422s with the reason; the recovered cluster archives under its own serverName. → 0.35.1: a retag, and the reason is worth the line. nat_gateways became a real quota class — config/quota-tiers.yml grants it and both readers expose it, so the tier lookup that phase 59 needed stopped 409ing where it should 404. The 0.35.0 IMAGE already carried that code: it was built from a working tree whose changes had not been committed, so for two days the running bytes were correct and the tree was behind them. validate-image-content could not report it — it compares the image to the tree and the tree was the wrong half. What caught it was the guard that reads git: a version number must not describe two different byte sets, so the tag is spent even though the bytes never move. → 0.35.2: no signup could be approved. Onboarding wrote the tier’s Octavia quota at /v2/octavia/quotas/<id> — the amphora sub-tree, which has no quotas member — so Octavia answered the PUT with 405 and the tenant rolled back; the same wrong path on the READ 404’d, and every tenant’s load-balancer ceiling rendered « unavailable ». The route is /v2/lbaas/quotas/<id>, now built in one place for all three call sites (onboard, re-tier, quota read). → 0.35.3: the Capacity page’s DISK card finally has a commitment. It had read « commitment not published » under a note that was true — openstack_cinder_limits* was zero metric names in this Prometheus — and whose cause was one string: the metrics exporter resolves the Keystone service type volumev3, kolla registers cinder as block-storage, so the collector failed to enable on every scrape for months while logging it once a minute. The class now sums both Cinder ceilings, gigabytes + backup_gigabytes — each an enforced quota, both landing in the Ceph the denominator measures. It runs well above physical, and that is the point: storage is sold thin, so the ratio is the commercial position rather than an outage. → 0.35.4: the DISK commitment now carries the whole promise. 0.35.3 summed Cinder's two ceilings while the admission gate summed block + object — two different populations answering the same question. All three land on the same OSDs, so each gigabyte is counted once and all of them are counted, with committed_parts beside the total because one number cannot say what it covers. The object half is read from RGW's admin API (1 list + N user calls on one credential) because RGW publishes no per-tenant quota series and, unlike Cinder, no catalog registration would make it a scrape. An account with no ceiling is counted and named, never summed as zero. → 0.35.5: the per-tenant table gains storage — block from Prometheus, object from the same RGW read the fleet total makes, merged by project id. A tenant RGW has never seen renders unread, never zero. → 0.35.7: an alarm row carries its detail — the label set, the fingerprint, the generating query, the last refresh. The console could render « / is below 15% free » and had no way to ask WHICH filesystem: mountpoint and device live only in the labels, and stopped here. (0.35.6 is burned — built and imported before anyone noticed main.py still announced 0.35.5; a changed tag is never reused.) → 0.36.0: two fixes of the same shape, found by calling every MCP tool against the live cloud. rotate_s3_credentials answered 500 on every call: RGW’s add-key subresource replies with the user’s key ARRAY and the code read it as the user OBJECT — and the AttributeError came after the PUT had minted a real keypair, so each call left an unusable key on the tenant’s RGW user and threw the secret away. And IRSA token exchange resolves an issuer through the FACTORY when IRSA_CLUSTER_MAP cannot: that map ships empty in every environment, so the map-only lookup answered 403 « not mapped to any tenant » for every cluster this platform has ever provisioned. A cluster’s issuer is its own control-plane endpoint and its namespace names its owner, so the fleet IS the registry — nothing to register, nothing to keep in step. An unreachable factory is a 503, never a 403: « I could not ask » and « the answer is no » are different facts, and returning the second for the first turns an outage into a permissions bug somebody spends an hour on. → 0.36.1: the same code as 0.36.0, which is burned. That image built clean and crash-looped in the cluster on a MISSING config/services.yml — lab-build.sh archived the session worktree’s symlink instead of the file it points at, and the catalogue assertion did its job: the pod never became Ready and the old ReplicaSet kept serving. Fixed in the build script with tar --dereference. → 0.36.2: the OpenStack proxy strips locations and direct_url from Glance image bodies. Those name where an image’s bytes live, and they are about to start appearing — show_multiple_locations is going on so nova can take an RBD clone instead of copying a whole 20 GiB disk. Nova reads them internally as a service user; a tenant’s only door is this proxy. Shipped BEFORE the flag, so the read never opens. → 0.36.3: 0.36.2’s location strip was a no-op. It matched "/image/" against the path the route hands over — image/v2/images/<id>, no leading slash — so nothing was ever stripped, and the live tenant read carried full rbd:// URLs once the Glance flag went on. The unit tests passed because they fed the full request path, not the handler’s. Selects on the first path segment now. → 0.36.4: image locations are no longer writable through the gateway. A location names the storage object behind an image, so pointing one at another project’s object and booting it is the surface Glance’s own docs say policy cannot fully close. Refused on both routes (POST .../locations and the JSON-patch to /locations), at every tier. Uploads, imports and snapshots are unaffected. → 0.36.5: the quota tier table, re-measured, with the top tier lowered to 1,500 GiB (block 1,200 unchanged, object 500 → 300). config/quota-tiers.yml is copied INTO this image rather than mounted, so a tier edit reaches the cloud only when the image ships. extension had not been raised — the stored-bytes stop it was measured against fell to meet it, 1,870.6 → 1,680.6 GiB, as ordinary tenant growth took 316 GiB off Ceph’s MAX_AVAIL (which is what is LEFT, not what exists). Signup admission was never affected: it reads Ceph live through Prometheus, not this table. → 0.36.6: the per-database storage ceiling, re-derived — and the second ceiling it should always have been derived against. Two limits now, because the SIZE was never the problem: one instance may still be 250 GB, and size × replicas may reach 300 GB. A 3-replica 250 is 750 GB of volume, which the cluster cannot deliver on either of the two ceilings it is measured against, so it is refused at create with the arithmetic and the trade in the message rather than half-provisioned. Trade replicas against size: 250 × 1, or 100 × 3. → 0.37.0: “the organisation’s users” now means ONE organisation. GET /v1/org/users answered a platform-admin credential with every tenant’s people in one flat list, and no row said whose member anyone was; the console rendered that under the heading “Organization members” behind a banner apologising for it. The three readings are now asked for by name — no argument is the caller’s own account, ?project_id= is one named account, ?scope=all is every account and is platform-only — and every row carries project_id and project_name. Inviting also stopped failing on the ordinary case: every tenant shares one Zitadel directory, so an address that already had a login came back as Zitadel’s own 409 "User already exists" wrapped in a 502. An existing person is now ATTACHED to the account — unowned or already yours attaches, another account’s refuses with the remedy and names no tenant, and re-inviting a current member answers 200 with their role updated instead of duplicating them. And a platform admin must now name project_id: their invite used to bind the new member to auth.tenant_id, which for a platform credential is the literal string admin — a service scope name, never a Keystone project, so the member could sign in and reach nothing. → 0.38.0: and the cloud owner can invite his own colleagues. 0.37.0 answered a platform credential's “my organisation” with the UNOWNED set — live, one colleague and five purged signup fixtures — and its invite offered an account picker of sixteen customers and no way to say IG1. Leftovers are not colleagues: membership of IG1 is the PLATFORM ROLE, which is what resolve_tenant reads first. GET /v1/org/users now answers a platform caller with IG1's team and scope=platform; role=admin on the invite adds a colleague to IG1 itself, writing no member row (one would bind a colleague to a project id, which is the “admin”-as-a-project bug in a new coat). Refused to anyone but a platform admin (403 on who is asking, not a 422 on the role name), refused alongside project_id, and refused for someone who already belongs to a customer account — promoting a customer's user to IG1 staff would hand them every other customer's cloud. The orphans stay reachable under scope=all, in the bucket that says what they are.
Factory (KaaS)0.6.5phase 22 → day-2 (upgrade, protection, nodes, autoscaling, ASGs; phase 41) → 0.3.0: leader election, so the service runs TWO replicas while its reconcile loops stay singleton — a node drain no longer takes the KaaS control surface with it. → 0.4.0: the cluster lifecycle memo moves onto the cluster record itself. Held per instance it was wrong in both directions — a watcher could be told a cluster was ready twice, and after a node recycle the second “ready” could never arrive. Control-plane port allocation now confirms its write, so two instances cannot hand the same port to two clusters. → 0.4.1: the operator console’s origin joins the CORS allow-list. The console moved to its own address on 2026-08-22 and this service was never told, so every call to it from the operator console died in preflight — a refusal that carries no HTTP status, which the page renders as an empty cluster list rather than as an error. Second occurrence of the same miss; a gate now derives the origin list from the deployed console configs. → 0.4.2: creating a cluster works again after the platform moved its OpenStack APIs to TLS. This service was still dialling the old plaintext address, and an endpoint that only speaks TLS answers a plaintext request by hanging up — which every client reports as a dropped connection rather than as a wrong protocol, so a healthy cloud looked like a dead one. Three separate places were affected and all three surfaced identically, as a new cluster that stays in provisioning forever: the credential the cluster builder authenticates with, the trust anchor a new worker needs before it can register itself, and the file that credential is read out of — read by position rather than by name, so it broke the moment a second entry had to be added to it. → 0.4.3: a cluster whose setup was interrupted can finish. If the service was restarted — or simply timed out — while installing a new cluster’s networking, the half-finished install left a marker that made every later attempt refuse to run, and the retry that was supposed to recover was the one command that never could. Such a cluster would have stayed in provisioning indefinitely with nothing obviously wrong. Interrupted installs are now cleared before the next attempt; completed and cleanly failed ones are left untouched, so a healthy cluster never has its networking removed underneath it. → 0.5.1: a retag rather than a change. 0.5.0’s image was built before a peer’s 0.4.3 landed, so the rebase carried their wedged-helm fix into the tree and NOT into the image already sitting on the nodes — an older base regressing a newer fix under a higher tag, which is exactly what « IMPORTED 3/3 » cannot show. Rebuilt, retagged, and verified INSIDE the image to carry both this line’s cluster_namespace and their _PENDING_STATES. → 0.5.2: the tenant OCCM runs the load-balancer controller (phase 53 — register gap #4). Every tenant type=LoadBalancer Service used to pend forever with NO error to read — the worst failure shape there is — because the bootstrap rendered the OCCM in node-lifecycle-only mode. Octavia itself shipped in phase 37; this is the last mile. Existing clusters converge by themselves: the bootstrap-addons annotation is now a STACK MARKER and the reconcile re-runs the idempotent helm upgrade on any cluster carrying the old one — no node rollout. The floating IPs come from the phase-50 public pool; an internal LB annotates itself and never claims one. → 0.5.5: the render now matches the platform it lands on. The first pilot upgrade proved the controller reaches Octavia and that the refusals are READABLE on the Service — and then ate three live-only misses: amphora is disabled here (the cloud runs the OVN provider), the provider key is lb-provider not provider, and OVN accepts exactly one algorithm (SOURCE_IP_PORT; ROUND_ROBIN is a 501). Each got pinned by the offline gate only after the live event named it. (0.5.3 and 0.5.4 are spent — both carry wrong-key images.) Proven on the pilot: a type=LoadBalancer Service got 10.168.210.162 and answers 200. The floating network stays the phase-50 public pool by default — right for post-50 tenants; the pilot uplinks to the 2034 shared L2 and got its override by hand. → 0.5.6: the cluster detail carries the tenant’s LoadBalancer Services, read live through the cluster’s own admin kubeconfig — best-effort by contract, so an unreachable tenant apiserver degrades to an empty block, never a failed read of the cluster itself. The CLI, the MCP and the console read the same document. → 0.5.7: …and 0.5.6’s read never actually worked live. httpx 0.28 silently drops cert=(crt, key) when verify is a path string, so every call arrived as system:anonymous and the best-effort blanket read the 403 as « no LBs » — over a cloud with a live one. The TLS context is now built explicitly (gotcha 368), with a regression test pinning the transport shape. 0.5.6 is spent. → 0.5.7: the LB read-back works live. httpx 0.28 silently drops cert=(crt, key) when verify is a path string, so every tenant-kubeconfig call arrived as system:anonymous; the TLS context is now built explicitly (gotcha 368). → 0.5.8: IRSA / workload-identity wiring for new tenant clusters (register #16). The Kamaji control plane now carries --service-account-issuer and --service-account-jwks-uri from the cluster’s own endpoint, so in-cluster workloads can authenticate to external services with projected service-account tokens. The httpx response is consumed inside its async with block so lazy streaming can no longer return an empty list. → 0.5.9: a retag on the same story as api 0.35.1. The 0.5.8 image was built and rolled on 2026-09-01 from a working tree whose version= bump had not been committed, so the pods reported a release the repository did not contain — the running bytes were right and the tree was the half that was behind. A version number may not describe two different byte sets, so the tag is spent even where nothing executes differently. → 0.6.0: phase 60.4’s missing half. 60.4 wired the OIDC issuer into every new tenant apiserver and shipped the API’s token-exchange route, but nothing could ASK or CHANGE that state — so the MCP’s list/create/delete_workload_identity tools called endpoints that existed nowhere and 404’d from the day they were published. GET /v1/workload-identity, plus GET/PUT/DELETE on /v1/clusters/{name}/workload-identity. The state is derived from the KamajiControlPlane’s --service-account-issuer flag rather than stored beside it, so it can never report « enabled » for an apiserver that signs nothing; enabling is idempotent, because the patch rolls the tenant control plane and doing it twice must be free. → 0.6.1: the tenant namespace gets its secret-read RoleBinding — phase 47’s other unfinished half. Clusters moved to ig1-kaas-<project> and the factory’s per-namespace secret Roles did not follow, so the kubeconfig endpoint answered 502 for every cluster created since. It binds a STATIC factory-kaas-secrets ClusterRole (get on secrets, no list, no watch) and holds bind on that one name only. → 0.6.2: 0.6.1 is burned — its four workload-identity routes all answered 500 on a real cluster, because CAPI spells the control-plane endpoint as an object and Kamaji spells it as a string, and the new helper assumed one shape. Both are read now, by one function, pinned by tests. → 0.6.3: the IRSA issuer reaches a tenant apiserver for the first time. 60.4 patched spec.extraArgs.apiServer; the field is spec.apiServer.extraArgs, and the CRD prunes the unknown one silently with a 200 — so that stage had been a no-op since it shipped and every cluster reported no issuer while the code reported success. Enabling now reads the flag back and refuses to claim a patch that did not land. → 0.6.4: the factory receives the tenant-credential map it has always claimed to read. TENANT_CREDS was described as the API’s map from the day it was written and never wired, so create_asg answered no Keystone application credential for EVERY project while the credential sat in that very secret. Same secretKeyRef the api addon uses; no admin fallback. → 0.6.5: new Kubernetes clusters build again. Creating one had been failing since 2026-08-27 for every customer: the cluster definition did not name which external network to egress through, leaving the cluster builder to find it by looking for the one network marked external. A second such network was added that day for the public front door, so the search started matching two, and the builder stopped before creating any worker — a cluster that sits in provisioning and never finishes. The network is now named explicitly. Existing clusters are unaffected and keep serving; a cluster created while this was broken cannot be repaired in place and must be recreated.
Billing0.8.1phase 21 → W7 cost/budgets (33) → budget-breach events + invoices → 0.4.3: GET / reports the version the service is actually running. It had answered 0.4.1 out of a 0.4.2 image since 2026-08-16, because the number was written down twice and the bump moved one copy. → 0.5.0: budgets stop being clobbered across replicas. A budget's durable write used to replace the whole set with one instance's view, so a second instance that had not seen your newest budget silently deleted it — a spending guard disappearing with nothing to show for it. Writes name the rows they change now, and every budget read re-checks the durable copy first. → 0.5.1: the same sibling-URL defect the gateway had, copied here with the comment that admitted it. Every OpenStack service URL was derived at import from a module constant, so moving the auth URL moved Keystone and left compute, network and volume behind — billing would have read usage over plain http while believing it had been secured, and every call would still have succeeded. → 0.6.0: one rating table. Rates had lived in two places — rating.Rates’ field defaults and config.Settings.RATES — each carrying an instruction to keep them MIRRORED, and the mirror had already drifted onto a page customers read: the console’s demonstration dataset still carried the pre-2026-08-13 compute rates, understating compute by 40% against what the engine charges. The rates are NOT restated here, deliberately — they live in one place and a guard holds every published quote equal to a charged rate, which is the same rule this release exists to enforce. The fix was a DELETION rather than a synchroniser — config.RATES is now empty and the env-override channel only, because validating an empty dict already yields the default table and a partial one overrides just the keys it names. → 0.6.1: this service enforces the billing-access rule itself. « Cost is visible only to a member the account owner granted it to » was written in August as one line in the API gateway’s proxy — and those proxy routes are a front for THIS service, which answers on its own address. So the rule covered one of two doors: a read-only credential the gateway refused read the full cost breakdown and the tenant’s budget list straight off the billing host. The flag now travels with the identity this service already resolves, and guards the cost breakdown, the forecast, budgets, invoices, usage and the AWS-compatible cost-and-usage report — the last of which the gateway never fronted, so it had been gated on NEITHER door. Both doors now refuse in the same words. → 0.6.3: the tenant edge bills at ALB parity (2026-08-26 — « we should charge the same »): edge_exposure_hour_eur = 0.02 per exposure-hour, collected from the api’s /v1/edge/exposures under the caller’s bearer, rated as an edge line at the exact default rate, and landed in UNTAGGED by the tag join (exposures carry no Neutron tags). An unreadable exposure list answers 503 — a zero bill is never written for a source that failed. → 0.8.0: the meter learns power state and deletion (phase 56, gap registers #9/#10). A durable lifecycle ledger in the state Secret folds every proxied power action and create/delete; a 5-minute reconciler corrects it from Nova’s own state for anything that bypassed the proxy; and the collector splits every instance window per state — SHUTOFF/SUSPENDED/SHELVED* time bills NO vCPU and shows on the line as storage-only, a volume or floating IP deleted mid-period bills its exact measured window, and every line the ledger cannot cover is marked estimated rather than dressed as measured. The consumer restarts itself on a Kafka outage; billing’s other duties never depend on it. (0.7.0 is unshipped — allocated twice by one mangled command, never built; the minor is 0.8.0. 0.8.0 itself is spent the same day: the state store’s merge kept only the two keys it knew and silently deleted the lifecycle section on every read+write cycle — measured live, the ledger held only the newest event — and 0.8.1 ships the pass-through.)
Events0.5.7phase 23 → schema adds (26/27) → W7 webhooks (33) → publisher authorization + api.audit / billing.budget.breached topics → 0.3.0: two consumer groups — a per-pod reader group so every replica sees every event, and a shared webhook group so each event is still delivered exactly once. That split is what made a second replica correct rather than silently partial. → 0.3.2: the same version-reporting repair as Billing 0.4.3 — the root endpoint had been answering 0.3.0 out of a 0.3.1 image. → 0.4.0: webhook subscriptions are PERSISTED. They lived in one pod's memory while the service ran two replicas, so a subscription created on one replica answered 404 from the other — about half of all calls — and every restart silently deleted every customer's webhooks. They now live in a Kubernetes Secret shared by both replicas, written under optimistic concurrency so two simultaneous creates cannot discard each other. A store that cannot be read answers 503 with the reason rather than an empty list, and a create that cannot be persisted is not reported as created. The delivery LOG remains per-instance and the endpoint now says so. → 0.4.1: the bus becomes the platform's NOTIFICATION FABRIC. Alertmanager now posts to an in-cluster receiver on this service (not a gateway route — nothing outside the cluster calls it); it had been posting to an empty URL and failing 3274 of 3274 notifications, including the two rules whose job is to report that alerting is broken. /v1/integrations makes a destination a RESOURCE you create yourself (Slack, Teams, PagerDuty, Opsgenie, signed webhook, SMTP) instead of a vendor URL in an Ansible vault: platform scope for IG1 staff, tenant scope for your own events. Two topics a live producer had been publishing into a map that did not carry them are registered at last — every tenant.alarms transition had been 404ing silently. → 0.4.3: the Watchdog heartbeat moved into the shared store and the Watchdog route dropped from a 24-hour repeat to 5 minutes — a dead-man's switch whose beat is slower than its own staleness threshold reports the alert path dead on a healthy platform. → 0.5.0: the alert de-duplication ring is shared between the two instances instead of held by each. A repeat notification landing on the other instance used to be republished, so the same alert reached your tools twice; and the ring had no eviction, growing for as long as the service ran. → 0.5.1: the operator-access topic reaches the running service. It was declared in the schema by the plane-split release without a rebuild, so the event registry advertised ten topics and served nine. → 0.5.2: the durable audit trail becomes queryable (phase 54). a new server-side history endpoint reads the api.audit topic itself — offsets_for_times for the window start, a bounded forward fetch, cursor pagination that owes a cursor whenever the window is still open (a page under limit is NOT the end: the sparse-tenant rule, OPERATIONS gotcha 369), and the ring’s own read rule applied server-side — a customer sees only their project’s rows, an untenanted row stays admin-only. Every other topic is a 422 that NAMES the one it serves. (0.5.2 is spent: its endpoint awaited partitions_for_topic, a plain function in aiokafka 0.14, and left position() — a coroutine — un-awaited; the first live read 500’d on every call while the offline fake carried the same wrong signatures. 0.5.3 is spent the same way: the sync metadata cache never carries the topic at all on this client — measured live — so partition discovery moved to the async metadata refresh; 0.5.4 ships the measured shape.) → 0.5.5: three lifecycle topics register (phase 56): compute.instance.lifecycle, storage.volume.lifecycle, network.floatingip.lifecycle — the billing ledger’s input. 7-day operational class, tenant-stamped from the payload: the topics are transport, billing’s own durable ledger is the memory. (0.5.5 spent with 0.34.3, same reason; 0.5.6 is the committed tree.) → 0.5.7: webhook delivery gets its own trust store, which is the difference between a feature and a decoration. Every leg this service had was internal, so one client verified against the mounted bundle — the lab CA plus the ISRG roots, three certificates. A customer’s endpoint is signed by whichever CA they use, so that bundle rejected it: verified live against two unrelated public hosts, every delivery failed CERTIFICATE_VERIFY_FAILED. The API also refuses to register a plain-http URL, by design — so the two rules together meant no webhook this platform accepted could ever be delivered, while create_webhook reported success and handed the customer a signing secret to store. The delivery context trusts the public roots AND the lab bundle: a tenant receiver inside the lab is a legitimate destination too, and picking one trust domain would break the other.
MCP server0.9.4phase 18 → per-customer caller-Bearer (27, MCP SDK 2.0.0) → the W9/W11 hardening + parity tool surface → 0.4.0: Barbican secrets (metadata/write/delete — payload read stays off the agent surface) + the tier catalogue → 0.4.1: the router tools follow the uplinked tenant router (gotcha 151) → 0.4.2: internet_facing on the load-balancer composite (publish through the .75 edge; delete cleans up). The 0.5.x line was withdrawn upstream — those tags are burned and never re-used. → 0.4.3: manage_security_group — the composite opens the listener port on a <name>-lb group and the delete takes it away. → 0.5.0: the DNS tools stop being a passthrough. They ride /v1/dns/zones now, so a SHORT name is qualified by the gateway instead of refused, and the namespace and quota refusals come back in the gateway’s own words. Three tools that were missing: create_dns_zone, delete_dns_zone (which the destructive-tool list already named without it existing) and update_dns_record — an agent could publish a record but not change one, and delete+create is an outage the PATCH does not have. (0.5.0 is burned for the same reason as api 0.17.0.) → 0.6.0: whoami. One hundred and fourteen tools could spend a customer’s capacity and not one could answer whose capacity it was — the parity matrix even recorded the tool as deliberately absent, on the grounds that « identity is implicit in the caller’s bearer ». Implicit to the server, which resolves it on every call; opaque to the agent, which holds a token and a hundred tools that act on « the caller’s project » without being able to ask which project that is. It answers with the person, the role, the account (name, alias and number), the project and the tier — and says so in capitals when the credential is not confined to one tenant. → 0.6.1: it names the CREDENTIAL when there is no person to name — a machine user asserts no display name, and an id is not an answer. → 0.6.2: one credential may hold several projects, and the ONE request funnel carries the choice. X-IG1-Project is stamped in a single place because 115 tools would otherwise be 115 chances to forget it — and a tool that forgets does not fail, it acts in the default tenant. whoami gained the other projects the credential holds, printed only when there is genuinely a choice: an agent shown one project name will otherwise assume it is the only one. → 0.6.3: claim_edge_domain, verify_edge_domain, release_edge_domain — 118 tools. Every answer repeats the claim→TXT→verify→A order, because an agent that reports a domain live before the verify call is describing an outage that does not exist. → 0.7.0: list_projects, create_project — 120 tools. The parity matrix had recorded project enumeration as deliberately absent — « an agent must never enumerate projects » — which was right while one identity meant one project, when enumerating could only mean reading the PLATFORM’s. Once a credential holds several, « which of MY projects am I in » is the question that makes the selector usable, and an agent that cannot ask it acts on a subject it cannot name. The platform-wide list stays admin-only and off the tool surface, where the old prohibition still applies. → 0.7.1: the same 120 tools, rebuilt. 0.7.0’s image had been built before the three edge-domain tools of 0.6.3, which a peer landed while it was building — so the tag shipped a registry the catalogue had already moved past, and the stated tool count disagreed with what main.py registered. Retagged onto the rebased tree rather than left to describe a surface that no longer existed. → 0.7.2: list_org_units, get_effective_org_policy — 122 tools. The org grew a tree in api 0.25.0, and an agent acting under an inherited ceiling has to be able to READ the ceiling: effective answers with the tier it would hold without the org policy beside the tier it holds, so the refusal it is about to meet is legible before it meets it. → 0.7.3: the default issuer follows the identity migration to zitadel.cloud.ig1.com — agents authenticate against the public issuer like every other client. → 0.7.4: the trust-store DEFAULT follows it too. ZITADEL_CA_BUNDLE still named the bare lab CA (/etc/ig1/ca.crt), which stopped verifying that issuer the moment it was re-signed by a public root — and because the W9 write-ahead audit mints its publish token there, all 28 destructive tools refused to act while read tools, /health and both replicas stayed green. The addon was moved onto the shared ig1-trust-bundle (lab CA + ISRG roots, one file) and a gate now requires that env, so this default no longer decides anything; it is corrected because a constant contradicting every deployment that reads it is the next reader’s wrong turn. OPERATIONS gotcha 315.
Secrets broker0.2.02026-08-13 — the only secret-read permission in the cluster, so the API holds none → 0.2.0: the number catches up with the code. Rotation landed in two commits (7ba4e178, then 0cdf0c8e fixing a rotation that could destroy a credential nothing regenerates) and the version never moved, so the broker ran rotation while self-reporting the release that predates it. validate-image-content.sh passed on it throughout — correctly, because the bytes matched the tree. What was wrong was the number, and hashing cannot see that. → 0.8.10: the seven tools phases 59-60 wrote and never shipped. dfec4e2d added list_nat_gateways / create_nat_gateway / delete_nat_gateway and the four workload-identity tools to services/mcp, and no image was cut — so the API served /v1/nat-gateways and /v1/workload-identity/token-exchange while the running broker held 142 tools against a tree that defined 149. The feature worked, which is exactly why nobody saw it: an API that answers reads as shipped, and only the CLIENTS were absent. The guard that caught it is the one that reads git rather than the cluster — « a service’s source has moved past the image it is pinned to ». 149 tools.
Portal0.41.1phases 29–31 console → A-series professionalization → parity surfaces → 0.14.0: internet-facing load balancers (the Public column, the wizard toggle, and the hint that the members must serve TLS themselves) → 0.15.0: the wizard opens the listener port by default (a managed <name>-lb group, shown in the table and the drawer) and says what unchecking it costs. → 0.15.1: the console's asset URLs are content-fingerprinted, so a new console release reaches a browser that already has the old one — previously every asset was cached immutable for a year under an unchanging URL, and a returning user kept the console they first loaded until they hard-reloaded. → 0.15.2: the console stopped serving the AppleDouble metadata sidecars the macOS build tar had been packing into the image. → 0.15.3: object storage and managed databases are no longer requested by a platform-admin session, which has no tenant and was being shown two refusals as if they were failures. → 0.15.4: a data source that fails TRANSIENTLY (429/5xx/network) is retried before it counts as failed, so a 200 ms hiccup upstream no longer paints the red « Partial sync » banner. A refusal (401/403/404/409) is still asked exactly once, and a real outage still raises the banner. → 0.16.0: the console wears the « IG1 Cloud » lockup, light and dark. The artwork ships under NEW filenames (logo-cloud-black.svg / logo-cloud-white.svg) because the logos are the two assets the fingerprinter deliberately skips and are served immutable for a year — overwriting the old names would have been correct on the node and invisible in every browser that had already loaded the console. The light-mode filter: brightness(0) is gone with it: it was normalising a monochrome mark and would have crushed the new lime accent to black. → 0.16.1: the console's own gate now follows the gateway's. /v1/monitoring/fleet and /alarms became platform-admin only on 2026-08-12; the console kept offering Monitoring to every customer, which could only answer 403 — and it drew that refusal as an outage, over node counts taken from the demonstration fixture. Monitoring is now a reserved surface (customers are pointed at Statut, which is the tenant-appropriate view of the same facts), a refusal renders as a refusal carrying the gateway's own words, and no view can print fixture data as a live reading. → 0.17.2: the console moves onto pkg/ui — one shared token layer for the whole estate, with a light mode that was designed rather than inverted. --lime-ink moved #6d8210 → #5f720e: the old value measured 4.07:1 on the light page, under the AA floor, and small text had been using it for as long as the light theme existed. (0.17.0 and 0.17.1 are burned — both were built from a tree predating the 0.16.1 Monitoring gate.) → 0.18.0: the Supervision view follows the gateway onto Prometheus — the page named Netdata as its source in the lede, the sub-header and the empty state, and the empty state pointed at an environment variable that no longer exists. The host column now shows the real node name (mox-pa6-ctrl-1) over its address, which the previous source could not supply. → 0.19.1: the console gains Integrations — configure where your notifications go yourself, per destination type, with a test button that takes the real delivery path. Destination secrets are write-only: the API returns a fingerprint, never the value, and there is no reveal route. → 0.21.0: the VPN page sells the service instead of announcing a refusal. Tunnels are still set up by IG1 — that has not changed — but a customer who needs one now learns the capability exists and how to ask for it, rather than reading “not available” and going elsewhere. → 0.22.0: « Cancel » closes the drawer. It did not, in 73 of the console's 75 drawers — the footer button carried the same data-close as the header × and a singular querySelector could only ever reach the first one. In the same release the console stops serving French under an English menu: the whole Container registry page, half the New-destination drawer and the Copy button on every code block were untranslated in all five languages. The translation layer is now a shared package with a gate that DERIVES its string list from the rendered console rather than from a hand-written list, so a new string cannot ship untranslated. And the Integrations counters stop saying « 0 armed destinations » over their own « read failed » — a zero after a failed read is a lie, and on that page it is the text of the outage the page exists to prevent. → 0.25.0: the top bar fits the window. Between 768 and 1023 px it carried its full desktop contents PLUS the burger — 1001 px of content measured in an 800 px viewport — and everything past the theme toggle, including the account chip, was pushed off-screen with no way back. A second, never-reported band did the same on the desktop bar between 1024 and 1107 px. Only the text-bearing controls shrink now, and they truncate rather than paint over their neighbour; tap targets keep a 38 px floor. Nothing dropped from the bar becomes unreachable — documentation returns to the account drawer, where it had been missing on phones entirely. In the same release, « ? » stops standing in for the customer's name: the id_token carries none of name / preferred_username / email, which OIDC permits, so the console now reads them from /userinfo — where it had always been entitled to. → 0.25.1: a platform admin’s every console refresh stopped raising « Synchronisation partielle » on a healthy cloud. The console still fetched the raw Kubernetes proxy to list tenant control planes; api 0.16.0’s plane split had REMOVED that path from the customer plane — which is the plane this console talks to — so it could only answer 404, and a 404 is not retried. The source is gone rather than re-pointed at the admin VIP, because this page loads without a VPN. No cluster row is lost: the factory listing already took precedence by name and is served on both planes. → 0.26.1: the DNS page. Zones and records the customer creates and changes himself, under Réseau. Editing a record is a PATCH, never a delete-then-create; the name and the type are LOCKED while editing, with the reason written down, because Designate cannot rename a recordset or change its type in place and offering those fields would offer a refusal. Found while wiring it: both DNS drawers passed onMount: where drawer.open reads opts.mount, so the hook was dropped in silence and neither submit button had ever been wired — Cancel still worked, because drawer.open binds the close handlers itself, which is exactly what made a dead form look alive. (0.26.0 is burned: its image was built from a tree predating 0.25.1 above, so pinning it would have regressed that fix — a newer tag carrying older bytes.) → 0.26.2: the console's language stopped depending on what a test happened to render. The rendering gate could only judge what it rendered, so a tab nobody opened, an invoice state no demo fixture carried, an error branch and a page it never loaded all stayed French under an English console. A second gate reads the source instead — it lexes every asset for literals, template fragments and markup text and fails on French the dictionary does not serve and a reviewed allowlist does not excuse, each allowlist entry carrying a written reason. 111 strings came in with it, including 11 the tenant DNS wave had just shipped French, caught on rebase before they reached the lab. → 0.27.0: the console is built once and deployed TWICE. plane=customer serves the customer front door; plane=operator serves the admin-console VIP and points at the operator plane of the gateway. On the customer build the operator views are ABSENT — not hidden: no nav entry, no served route, no view function. That is ergonomics and defence in depth, not a security boundary — the same bundle ships to both, and the boundary is the dedicated Envoy listener, the SNI reject on every customer frontend, and the gateway’s own plane split. It exists because a platform admin on the customer console asked for /v1/monitoring/fleet, got the 404 of a router that is absent from that plane, and read « Flotte indisponible » — an outage message for a structural absence. Landing with it, all reported from the live console: the sidebar counted the demonstration fixture (« Webhooks 3 » above a page saying there were none), and counts are now a number or nothing, never an unearned zero; a created agent credential could not be found in the list because the drawer showed the Zitadel client_id while the list is keyed on the store id, and GET /v1/credentials did not return the former at all; three « create » buttons for one action collapse to one, on a page that now READS the credentials it used to ignore; the agent token is called IG1_TOKEN, as the CLI and these docs have always called it; the notifications bell reads GET /v1/events instead of announcing that the stream is unwired, which it was not; DNS — shipped in 0.26.1 — becomes reachable at all, having been absent from the production view set and therefore filtered out of every customer’s navigation; and the render path stops repainting four to six times per boot, including an unbounded fetch→render→fetch loop that froze the tab on an expired session. → 0.27.1: a second pass after eight adversarial verifiers went at the new i18n gate. It was reading five of the nine scripts the console loads — the absentee being data.js, the dataset the Status, Supervision and Incidents pages render — so its file list is now derived from the pages rather than written by hand. The status and incident prose had been marked data-noi18n to quiet a gate, which suppressed English the dictionary already held. And three DNS sentences had been “fixed” by enrolling their fragments as dictionary keys: green gate, unchanged browser, because the walker matches whole text nodes. They are single translated blocks now, proven by running the shipped engine against the shipped dictionary. → 0.27.2: the second wave of the console audit, all correctness. The credentials list stops calling MINTS “calls” — it read the mint/revoke trail, so a credential appeared used the moment it was created, and it merged a per-replica ring so the number also depended on which pod answered; it now reads api.audit and agent.actions, deduped by envelope id, and the column says it counts WRITES (a read-only credential doing its job correctly shows zero for ever, which is the true answer). The cross-tenant quota routes leave the customer front door for the operator plane. The shown-once client secret and PAT no longer outlive their drawer in the DOM. data-i18n-vars escapes the VALUES it splices into innerHTML — escaping at the call site protected the attribute and only looked like it protected the value. Thirteen radio-card drawers gain a focus ring that was being painted fully transparent, including the one that picks whether a credential can destroy. And the org-users table names its PLATFORM scope before the first row, where an operator had been reading every customer’s staff on a page identical to a customer’s own. → 0.27.3: accessible names on the four copy buttons of the credential-created drawer. A screen reader announced all four as « Copier », so the one that yields the PAT — the only one an agent runtime needs, and the only one that is never recoverable once the drawer closes — was indistinguishable from the other three. → 0.28.0: the accessibility wave. Six drawer tab strips were a bare row of buttons — no tablist, no selected state, no link to the panel each governed, and arrow keys did nothing; fixed at the single mount point all six pass through. There was no skip link at all, so the first Tab landed in the topbar and every keyboard user walked some twenty navigation stops to reach the content, on every view. And moving between views moved nothing and said nothing: focus stayed on the item you clicked, so the next Tab resumed from the middle of the menu, and no screen reader was told the page had changed. Ten controls had no accessible name, fourteen form and error paths delivered validation only as a disappearing toast, and three operator pages showed every tenant’s data with no scope indication — one of them saying « Périmètre : votre projet » while listing the whole estate. → 0.28.3: the console's own Content-Security-Policy learns the operator gateway. connect-src listed only the .64 origins, so every api call the OPERATOR console made was blocked by the browser before it was sent — a bare TypeError: Failed to fetch, no HTTP status, and zero lines in the gateway's log. The page rendered that as « Fleet unavailable », « Tenant read refused » and « Capacity unread » over a healthy cloud. One image serves both consoles, so the list must be the UNION of every origin either calls. → 0.28.6: the operations console gets a DOOR. One Zitadel application serves BOTH consoles, so a CUSTOMER account signs in to the operations console perfectly well — and was shown a console bearing their own name, carrying the operations navigation (Tenants, Capacity, Fleet monitoring), each page refusing them one click later. The refusals were correct; the surface was not theirs to read. A non-platform account now meets a single page that says so and points at its own console. This is not the security boundary — the gateway is, it held throughout, and a test now drives a customer identity against EVERY operator route and asserts that none answers 200. Also undoes a 0.28.5 regression: the OPERATOR dashboard rendered on the CUSTOMER plane for an admin, which then read « Capacity unread » because /v1/capacity does not exist there by construction. The plane decides which data exists; the scope decides what the caller may read. → 0.28.7: the exposure view carries the customer’s own domain, in TWO columns and not one — « reserved » and « served » read as each other, and a customer who confuses them publishes an A record early and meets a 421 from a gateway working as designed. The detail panel prints the exact TXT record. In the same release the gateway card stops naming the WRONG blocker: public certificates never waited on the DNS delegation, they wait on this platform being reachable at all, because a public CA proves control of a name by connecting to it. → 0.28.8: 0.28.7 is burned — two sessions built that tag within the hour from different trees and the later import won, so all three nodes carried one session's console door and NOT the other's domain-serving surface, under a pin naming the latter. Rebuilt from the merged tree and verified to carry BOTH on all three nodes. → 0.28.9: the Capacity page stops drawing Placement’s local-ephemeral inventory as the platform’s storage — « 0 TiB / 122 TiB », a resource nothing on this cloud consumes, rendered as entirely free — and reads Ceph instead: the same pair the quota stop already admits against, with four Ceph groups and a fill gauge on the fleet rollup. → 0.29.0: the project selector, connected to phase 47 at last. Every other client got it on 2026-08-23 — ig1 --project, the three SDKs, the provider’s project_id, the MCP funnel — and the console, the surface a non-technical customer actually uses, could neither name nor change its project. It did something worse than omit it: the control was written VISIBLE in the markup and hidden by script after boot, so every reload flashed a scope selector that then disappeared. It is now built from whoami.projects, sends X-IG1-Project on every call including the OpenStack proxies, and empties the collections before re-hydrating so no project’s rows are ever shown under another’s name. « All projects » is gone in production: a customer token resolves to ONE Keystone project, and there is no consolidated cross-project read to back the promise. → 0.29.1: the console stops serving French inside the zones its own translator is told to skip. Dates were the bulk of it — fmtDay and fmtDayTime asked for fr-FR in hard, so « 5 août 09:42 » and « 31 déc. 2026 » went out in all five languages, 37 of them across ten views — alongside « Jamais », « Courriel (SMTP) », « réseau privé », every security-group rule description, seven code-block comments and four alert destinations. None of it was ever red: it all lives in SKIP_SEL (pre, .mono, [data-noi18n]), the walker does not enter those zones, and the coverage gate READS SKIP_SEL from the engine rather than restating it — so the renderer and the gate that judges it shared one blind spot, and a blind spot shared by both halves cannot be seen from inside. The gate now renders the whole console a second time in the TARGET language and refuses French in any excluded zone, plus a second section for prose stitched to a locale-formatted value (localising the month broke « août 2026 · à ce jour », which used to match the dictionary whole). Both were seeded by being watched red. And demoText() draws the line the allowlist could not: the demonstration prose we write is translated, a customer's own text never is — which retired 21 allowlist entries, one of whose stated reasons was simply false. Also: <meta name="description"> follows the language on both pages (applyI18n starts at document.body, so nothing in <head> was ever walked). → 0.30.0: the console can now CREATE a project. 0.29.0 gave it the switcher; a customer who wanted a second project still had to reach for ig1 projects create, which is not a thing a portal customer has. « Créer un projet » sits at the bottom of the project menu, offered only to who the API will accept — tier 2 and an owning role, two different questions. The tier ceiling is shown rather than discovered by refusal (GET /v1/projects serves tier/limit/held for exactly that reason), and at the ceiling the row stays visible and disabled with its reason. Projects of the account this token cannot address are rendered inert rather than hidden, because « 2 of 15 » above one line reads as a bug. The name is previewed live, since the gateway slugifies it and prefixes ig1-. After the 201 the project is not auto-selected: whoami is served by two replicas each caching the credential store, so a bounded poll waits for it to become selectable instead of dropping the customer into the revoked-project fallback. → 0.31.0: and then it did not work — the only trace being ?name=… in the address bar, a native GET form submission that reloaded the page before any request was issued. Seventeen drawer forms guarded themselves with onsubmit="return false", an ATTRIBUTE handler, which this console’s script-src 'self' has been blocking all along; only a one-field form submits on Enter, so no older drawer ever exposed it. Guarded now at drawer.open, where the eighteenth drawer inherits it, and Enter performs the primary action rather than nothing. → 0.33.0: the console can now DELETE a project and CLOSE the account. « Gérer les projets » sits below « Créer un projet » and deliberately not as a per-row action inside the switcher — a switcher is used quickly and often, and an irreversible destruction one pixel from « look at this » is a design that is eventually right about somebody’s afternoon. The delete drawer shows what the project CONTAINS before the confirm field becomes usable, and stays inert if that survey failed: not knowing what you are destroying is not a reason to proceed. (0.32.0 is spent — a peer had built it on the nodes without landing it, and the allocator refused rather than clobber their release.) → 0.34.0: signing out of one console signed the other one out. end_session does not close a page’s session — it terminates the browser’s session at the identity provider, and every token minted in it is revoked, whichever application asked; the gateway introspects every token, so the other console 401s on its next call. A surface now signs out by REVOKING ITS OWN TOKENS, and « Se déconnecter partout » keeps the old behaviour for a shared machine, where a surface sign-out would otherwise leave a session anyone can walk back into. Signing in asks which account, so the two consoles can be two different people at once. Each console also has its own OIDC application now, which is what keeps their stored tokens apart — it is not what keeps their sessions apart. → 0.35.0: a project you had just created only appeared after a page reload. The selector’s rows were built from the account list alone — which is always present here, since you open the selector to reach « Créer un projet » — and the poll that follows a creation re-read only the OTHER list, so it announced « Projet prêt » over a menu that did not carry the project. One derivation now serves both, as the union of the two lists, and the new project shows immediately as a disabled row saying it is being prepared — which it is, for as long as the gateway replica answering has not refreshed its credential cache. → 0.35.1: the console finally has ASCENDING breakpoints. The page is centred and its width ceiling rises with the screen (1440 → 1680 → 1920 px at 1800 / 2600 px of window) — a 2560 display had 845 px of dead band on one side, 37 % of the content area. Table headers pin under the topbar while a 50-row page scrolls (≥ 1440 px); the drawer follows the route in both directions, so the back button and pasted links close a drawer whose resource left the route (only in-app navigation did before); and the spend curve’s stroke and terminal dot no longer distort when the chart stretches to a wide card. → 0.35.2: the issuer migration release — config.js moves to the public issuer and the public service bases (a browser on 5G calls api.cloud.ig1.com, not a VIP it cannot route), and nginx answers the APEX with a 301 to the canonical host: one URL for bookmarks, sessions and OIDC redirect URIs. → 0.35.3: the CSP connect-src learns the public origins — the guard that exists because a browser refuses BEFORE sending caught the union missing api/factory/docs/mcp/zitadel.cloud.ig1.com. (0.35.2 spent, same reason.) → 0.35.4: default_server, said explicitly — nginx's implicit default is the FIRST block for the listen address, so the apex 301 block silently became the default and every Host got redirected to the portal, consoles included; found by the flip play's own portal probe minutes after 0.35.3 rolled. (Spent.) → 0.36.0: the signup queue becomes the LISTE D’ATTENTE on Locataires. It shipped on 2026-08-13 as a « Demandes de compte » tab under Sécurité, beside API keys and agent credentials — the wrong neighbourhood: Sécurité describes what an EXISTING tenant holds, and who becomes a tenant is not that. An operator looking for where to accept the people who signed up opens Locataires, whose own lede already promises « the same arithmetic as accepting a signup », and found nothing. MOVED, not copied — two surfaces for one admin gesture are two states that drift — and loaded independently of the tenants table, so a 403 on one cannot blank the other. The tier offered on acceptance is derived from the sellable catalogue rather than a list written into the page. → 0.36.1: the operator console calls its gateway by its REAL name — apiBase api-admin.cloud.ig1.com (VPN-only A record), and the CSP union follows. → 0.36.2: the public signup page renders the 409 instead of a number. It had branches for 422/503/429 and a numeric fallback, so \u201cyour organisation already has an account — ask its owner to invite you\u201d reached a prospect as erreur 409: a code in place of the only thing left to do. → 0.38.0: the DNS page gains « Connecter votre domaine » — the nameservers to publish, copyable, read off the zone itself; three registrar-agnostic steps (no provider is named, in any of the five languages); a delegation check that says where the domain points TODAY; and an import drawer that previews the plan before writing. The page's own lede stopped claiming zones « resolve publicly » full stop: they resolve once the domain is delegated, and that sentence was the only explanation a customer had for why a brand-new zone answered nobody. → 0.39.0: the « Connecter votre domaine » panel stops concluding before it has read. It said « this zone's nameservers are not published yet » whenever its list was empty — and the list is empty just as often because the console has not FETCHED the zone's records, which is the state a customer photographed, next to a table where the NS records were plainly there. It now loads them and says so. Two additions from the same screenshot: replace-don't-append (two authoritative nameserver sets for one domain answer differently depending on the resolver), and a warning when the zone still holds no records at all — delegating an empty zone takes the domain offline the moment it takes effect. The « only one nameserver » line retired itself the same day, because it was always conditional on the list read off the zone rather than on a constant. → 0.39.1: the Supervision view renders the two RabbitMQ tiles the gateway gained in 0.32.5 — « Files RabbitMQ » (consumed-or-shallow over every queue, naming any unconsumed backlog) and « Messages en file (max) », the deepest queue in the cluster. Labels ship in all five languages, and the demonstration fixture carries the same two tiles, so a demo render can never disagree with the live shape. → 0.39.2: the database wizard sells phase 52’s durability. Two PostgreSQL-only fields join the drawer — Répliques (1–3, with the size × replicas charging note) and Sauvegarde (daily 03:00 base + continuous WAL, 7-day retention) — and the detail’s overview tab shows what was actually bought: the replica count, the schedule, the oldest point-in-time-recovery point and the last successful backup. Kafka hides the section rather than offering a 422. The demonstration fixture finally agrees with its own event log — checkout-primary had claimed daily backups for months while the overview showed none. → 0.39.3: SPENT — built, imported and never pinned. Its build ran from a shared checkout whose uncommitted resize-drawer edits had just been discarded by a git checkout -- . issued during the api 0.33.2 land, so the tag carried 0.39.2’s bytes under a number claiming the drawer. → 0.39.4: the resize drawer scales replicas too (api 0.33.2), for real this time. PostgreSQL rows gain a Répliques select beside the size field — in place, in either direction — and the size field stops defaulting ten gigs above current: it now defaults to the CURRENT size and is sent only when it actually changes, so a « replicas only » gesture can never grow the volume by accident. → 0.39.5: the cluster drawer shows the Load balancers (phase 53). It fetches the factory’s detail (0.5.6) and lists every type=LoadBalancer Service with its address — an em-dash while the OCCM is still ensuring, because an empty address is a fact, not a failure. The read failure renders as a retry hint, never as an empty list that would read as « no LBs ». → 0.39.6: the same bytes with the loop variable renamed — the loadbalancers smoke’s i18n harvester greps the source for lb.-prefixed keys and swept the drawer’s own lb loop variable into its mute-key assertion. The harvester cannot tell t('lb.x') from lb.x, so the variable is svc now. (0.39.5 is spent — built before the rename.) → 0.39.7: the integrations demo’s topic set follows the bus registry — the three phase-56 lifecycle topics (the portal smoke pins the demo to the registry, and it caught the new topics missing within the hour). → 0.39.8: the instance drawer gains its history — a SVG sparkline per series under the live gauges (CPU busy, memory %), loaded once from the new history endpoint, with the coverage note so a curve that ends mid-window is not read as « the instance was off ». → 0.39.9: the billing period label stops being asked of the dictionary. It is already complete in the active language when it leaves costPeriodLabel — the month from Intl, the tail from tr() — so it renders under data-noi18n now. The coverage gate had been demanding the whole composed phrase, and the catalogue carried exactly one: « août # · à ce jour ». On 1 September nothing matched and the offline lane went red over a console that was correct in all five languages. The exclusion is not a blind spot: the gate’s excluded-zone pass re-renders in the TARGET language and now PROVES the label comes out as « September 2026 · to date ». → 0.39.10: the console could print a remedy it made impossible to follow. A name collision is refused with « accept again with an explicit tenant name » and approveSignup sent only the tier, so two colleagues from one company left an application that could never be accepted. The field now travels, and an « accept under another name » drawer asks for it — opened ahead of the click when the tenant table already knows the collision, and on the gateway’s verbatim refusal otherwise. → 0.39.11: a tier downgrade looked like it had done nothing. The write was correct end to end — extension → standard re-posed Nova, Cinder and Octavia, and the metrics already carried the new figures — but the Capacity view cached its reading for the whole session and only the tenant table was invalidated, so the operator re-read the document from before the change, indefinitely, with nothing failing. All three writes that move committed capacity (re-tier, signup acceptance, tier backfill) now invalidate it, and the toast says the page follows at the next reading — the exporter is scraped every 60 s, and a latency nobody mentions reads as a bug. → 0.39.12: the signup page said « the signup service is unreachable » to one applicant while two of his colleagues signed up without trouble the same morning. Everything measured healthy — route, CORS preflight from both origins, CSP, certificate and its SAN list, pods with zero restarts, queue at 26 of 200, zero 5xx at the gateway and nothing at the public front door — and a valid submission answers 202 in 387 ms from a real browser. His request never arrived. The defect was that one sentence was the only rendering of four unrelated causes: no network, the API host blocked at the applicant's end, CSP, and any exception thrown while handling the response, which the end-of-chain catch swallowed too. It now separates transport from rendering, probes its own origin to say which network is at fault, names the unreachable host, and prints the browser's own error so the next screenshot is diagnosable. → 0.39.13: the Disk card shows what its commitment is made of — block and backups, objects, the count of object accounts with no ceiling, and « partial total » when a half could not be read. → 0.39.14: the Disk card’s foot no longer overflows into its neighbour (.s-foot was a single-row flex with no wrap), and the per-tenant table carries Block and Object columns. → 0.39.15: an alarm row opens. It rendered four columns and listened to nothing, so a disk alarm named a threshold and never the filesystem; the panel now shows the measured value, the full label set, how long it has been firing, which alertmanager raised it, and a link to the query that fired. → 0.39.16: the launch wizard can ask for SSH from the internet, allocating a public floating IP and attaching it to the new instance. → 0.40.0: the four things a customer test run tripped over on 2026-09-09. The S3 credentials card hands out a ~/.aws/credentials profile instead of a sentence telling you to write your own export — a tenant access key is tenant$<project>-…, and an unquoted export truncates it at the $ before the AWS CLI ever sees it, which RGW answers with an error the CLI then renders as a Python TypeError (gotcha 391). A public IP can be attached and detached from the instance itself rather than only from Network › Floating IPs. Addresses, instance id, hostname and key-pair name are click-to-copy. And an attached address appears immediately: the console refreshed the floating-IP list but never the server list, so the new address showed up only after something else re-polled Nova — which a reboot did, making it look like the interface needed one. → 0.41.0: the member list is split in two, because it was answering two questions under one title. Security › Organization users is now ONE account’s members for everybody; the platform-wide list moved to Tenants › Members, where each row names the account it belongs to and a filter narrows to one. The banner that used to warn “this list spans EVERY tenant” is gone with what it was warning about. Inviting from the platform view now asks which account the person joins, and the drawer says plainly what happens when they already have an IG1 Cloud login. → 0.41.1: Security › Organization users is “The IG1 team” for a platform admin, and its button invites a COLLEAGUE — no account picker, because IG1 is not a customer. Seating someone in a customer account is a different gesture and stays on Tenants › Members, where an account is already selected. The drawer states the one platform role that exists and why a customer's member cannot be promoted into it.
Docs (this hub)0.12.56phase 35 (W8) → 0.4.16: the DNS status caught up with reality — cloud.ig1.com was delegated on 2026-08-17 and the hub still said no zone had been, so the networking guide now carries the delegation, the hardening, and the reason it resolves for nobody; the certificate line stopped promising that delegation would bring public certificates (it blocked that path instead), and the Terraform example stopped naming a project id that had been deleted. → 0.4.17: the version table follows factory and events to 0.3.0, the release that made them safe to run as two replicas. → 0.4.18: the Portal row follows the console to 0.15.1. → 0.4.19: same AppleDouble sweep as portal 0.15.2, and the Portal row follows to 0.15.2. → 0.4.20: the Portal row follows to 0.15.3. → 0.4.21: routine sweep. → 0.4.22: the API reference stopped rendering null as every response sample — the gateway now ships examples captured from the running cloud (api 0.12.4), so the reference finally shows what a call returns. The audit behind it also found the hub itself sound: no broken links, no cited CLI command that does not exist, and every cited /v1 path resolving. → 0.4.24: the Factory row follows to 0.3.3 and the Portal row to 0.15.4 — the release that stopped a transient upstream read from reaching a customer as an outage. → 0.4.25: the API row follows to 0.12.5 — every idempotent read of the management cluster now retries a transient failure, and the alarm and status stores stopped answering a read they could not complete with an empty collection. → 0.5.1: the topbar wears the « IG1 Cloud » lockup instead of the lime "IG" chip and the wordmark beside it — the artwork carries both, so the text was saying the name twice. The WHITE variant is hard-coded here, which is the light-on-dark rule applied rather than skipped: this hub has one theme, so a prefers-color-scheme swap would put the black lockup on a black page for anyone whose OS is in light mode. The Portal row follows to 0.16.0. (0.5.0 was built before this table was corrected and is burned — the hub is SERVED bytes, so a page edit is a release.) → 0.5.2: the Portal row follows to 0.16.1, and two pages stopped telling a customer that fleet monitoring was theirs to read — the API gateway card and the Terraform provider's out-of-v1 callout both listed monitoring among the read-only domains a tenant can call. It has been platform-admin only since 2026-08-12. → 0.6.2: the hub adopts pkg/ui and gains light/dark, having been single-theme since phase 24; docs.css is fingerprinted, which is what lets a returning reader actually receive the new tokens instead of a year-immutable cached copy. The Portal row follows to 0.17.2. (0.6.0 and 0.6.1 are burned — both predated the 0.5.2 monitoring correction.) → 0.6.3: the Billing row follows to 0.4.3 and the Events row to 0.3.2 — the release in which both services stopped reporting a version they were not running. This hub is SERVED bytes, so a row edit is a release. → 0.6.4: the API row follows to 0.13.0 and the Portal row to 0.18.0 — the release that moved fleet monitoring off a port it could never reach. → 0.6.5: the API row follows to 0.13.1. → 0.6.6: the Events row follows to 0.4.0 — the release that stopped losing customers' webhooks. → 0.6.7: the API row follows to 0.13.2. → 0.6.8: the API row follows to 0.13.3 — the release in which removing a member stops costing the account a credential slot forever. → 0.6.9: the version table follows Events to 0.4.2, the API to 0.14.1 and the Portal to 0.19.1 — the release that gave the platform a notification path it owns. → 0.6.11: the API row follows to 0.15.0 and the Portal row to 0.20.0 — the release that made the container registry self-service. (0.6.10 was BUILT and imported 3/3 before this branch rebased onto a mox that had corrected the Events row to 0.4.2 — a tag whose image exists with different bytes is spent, so it was retired unused rather than rebuilt over.) → 0.6.12: the Billing row follows to 0.5.0 and the Portal row to 0.21.0. → 0.6.13: this hub fits a phone. docs.css contained the string @media zero times, and measured at 375 px NINE of the eleven guides scrolled sideways — mcp.html laid out 844 px of content in a 375 px viewport. Reference tables now scroll inside their own container instead of taking the page with them, long monospace identifiers may break, and the topbar wraps. The Portal row follows to 0.22.0. → 0.6.14: the Events row follows to 0.5.0 and the Factory row to 0.4.0 — the replica-state audit. → 0.6.15: the version table follows the API to 0.16.0. → 0.6.16: the API card, the reference banner and the version row all follow the 0.16.0 contract (77 paths). → 0.6.17: the Portal row follows to 0.23.0 — the release that stopped reporting a device’s missing certificate as a platform outage. (0.6.15 and 0.6.16 are burned: peer sessions took both while this branch was unrebased, and an image had already been built and imported under each — a tag whose image exists with different bytes is spent.) → 0.6.18: the Events row follows to 0.5.1. → 0.6.21: the Portal row follows to 0.25.1 — the release that stopped a platform admin’s every console refresh from reporting a partial sync over a source api 0.16.0 had moved off the customer plane. (0.6.19 and 0.6.20 are spent. 0.6.19 was burned while pinned: this hub briefly ran a pod labelled 0.6.19 that served 0.6.18’s bytes — Running, Ready, and wrong — because the tag already existed from a concurrent session and the image was never rebuilt here. 0.6.20 is that session’s own release and was left alone.) → 0.6.22: the tenant-DNS wave — the API row follows to 0.17.1 (81 paths) and the Portal row to 0.26.1, the provider page gains its DNS section (ig1_dns_zone, ig1_dns_record, and why name/type/zone_id force replacement), and the MCP catalogue gains create_dns_zone, delete_dns_zone and update_dns_record. Every count claim follows with them: 22 provider resources, 114 MCP tools. (0.6.21 is spent too, by the same race one turn later: it was picked as « absent from all three nodes » — true when checked, false minutes later, because a concurrent session was mid-build on that exact tag. Absence from the nodes is not a lock; between the check and the build there is a window. 0.6.22 was BUILT and grepped inside before this row was pushed, which is the ordering that ends the loop.) → 0.6.23: the API row follows to 0.17.2. → 0.6.24: routine sweep with the DNS wave's open item — a new tenant zone resolves but is recorded PENDING indefinitely, which is the pool behind Designate rather than the API in front of it. → 0.6.26: the API row follows to 0.18.0. → 0.6.27: the API row follows to 0.18.1 and Billing to 0.5.1. → 0.6.28: the API row follows to 0.18.2. → 0.6.25: that note corrected to what was measured — four failures in five consecutive live runs, the passing one resolving end to end, so it is a race the pool usually loses rather than a path that is broken. Understating a surface is still misstating it. → 0.6.30: the table follows portal 0.27.0 — the console’s split into a customer plane and an operator plane, and the seven defects reported from the live portal that land with it. → 0.6.32: the table follows portal 0.27.2 — the second wave of the console audit. → 0.6.33: the table follows portal 0.27.3. → 0.6.35: the API row follows to 0.18.3 — the release that put two already-written security fixes into the running cloud. (0.6.34 is burned: it was built and imported 3/3 from a tree that did not yet carry the 0.6.33 row, because a peer landed it mid-build. Shipping it would have regressed that row under a higher tag — pin forward, pod Running and Ready, page silently backwards.) → 0.6.42: the API row follows to 0.19.2 and the Portal row to 0.28.1 — the release that gave operators a whole-cloud rollup instead of two tables. (0.6.39 is burned: built and imported 3/3, then a peer landed 0.6.40 and 0.6.41 while this branch was mid-flight, so that image predates rows the hub now carries.) → 0.6.43: the API row follows to 0.19.3 — the release that corrected this reference's own samples. → 0.6.46: the Factory row follows to 0.4.1 — the release that let the operator console talk to the factory at all, and the API row to 0.20.1. (0.6.45 is burned: it was BUILT and imported 3/3 from a tree that did not yet carry the 0.6.44 row a peer landed mid-flight, so shipping it would have regressed that row under a higher tag — pin forward, pod Running and Ready, page silently backwards. 0.6.44 is that peer's own release and was left alone.) → 0.6.47: the API row follows to 0.21.0 and the MCP row to 0.6.0, and the MCP catalogue gains whoami — 115 tools, with every count claim on the hub following, and server.json gaining an identity domain so an agent directory meets it first. (0.6.45 is burned by this session: picked while origin/mox said 0.6.44, then built and imported 3/3 before the rebase that revealed a peer had already taken 0.6.45 and 0.6.46 — an image with different bytes exists on the nodes under that tag, and ctr images ls cannot tell the two apart.) → 0.6.49: the API row follows to 0.21.2. (0.6.47 and 0.6.48 are spent — each was built and imported naming an api tag that burned underneath it.) → 0.6.52: the API row follows to 0.21.4 and the MCP row to 0.6.1. (0.6.47, 0.6.48 and 0.6.49 are spent — each was built and imported naming an api tag that burned or moved underneath it.) → 0.6.53: the Factory row follows to 0.4.2 — the release that restored cluster creation after the platform’s OpenStack APIs moved to TLS. → 0.6.55: the API row follows to 0.21.5 (84 paths) and the Portal row to 0.28.2 — the release that gave the operator console a capacity page and a cross-tenant view of the estate. (0.6.54 is burned: BUILT and imported 3/3 from a tree that did not yet carry the 0.6.53 row a peer landed mid-flight, so shipping it would have regressed that row under a higher tag.) → 0.6.56: the Factory row follows to 0.4.3 — the release that lets a cluster whose setup was interrupted finish. (0.6.54 was built independently by TWO sessions the same afternoon — this one and the one that then took 0.6.55 — which is what a burned tag looks like from both ends: each picked it while the other was mid-build, and ctr images ls cannot tell the two sets of bytes apart. Nothing pins it.) → 0.6.59: the Portal row follows to 0.28.3 — the release that let the operator console talk to its gateway AT ALL: the console's own CSP connect-src had never learned api.10.57.8.69.nip.io, so the browser blocked every call before sending it. (0.6.58 is burned: BUILT and imported 3/3 from a tree that did not yet carry the 0.6.57 row a peer landed mid-flight.) → 0.6.64: the API row follows to 0.22.2 — the release in which the platform’s Ceph capacity stop stopped reading a hand-written constant. Both halves of that stop (stored bytes and MAX_AVAIL) now come from the Ceph mgr exporter through Prometheus, and an unreadable source narrows the budget rather than widening it. → 0.6.65: the Portal row follows to 0.28.6. (0.6.64 belongs to a concurrent session.) → 0.6.66: networking §6.1 — bring your own domain, the four steps in order, and a callout that says plainly a browser will warn until this platform has public ingress. The API row follows to 0.23.2, the MCP row to 0.6.3 (118 tools), the Portal row to 0.28.7, and the provider inventory to 23 resources with ig1_edge_domain_verification. → 0.7.2: the Portal row follows to 0.28.8. (0.6.67–0.7.1 belong to concurrent sessions and were left alone.) → 0.7.3: the same bytes on a free tag — a peer took 0.7.2 while this build was running. → 0.7.4: the API row follows to 0.24.3, the release in which a renderer upgrade finally reached the edge — the store had gone on serving a config cached from the previous release, so update_store now self-heals a stale render on any write attempt and POST /v1/edge/resync is that attempt made deliberate after a deploy. This release also put back the API row’s own narrative, truncated sixteen minutes earlier by the merge described at 0.8.2. → 0.7.5: the API row follows to 0.24.4 and the Portal row to 0.28.9 — the release in which the operator console stopped drawing Placement’s local-ephemeral inventory as the platform’s storage (« 0 TiB / 122 TiB », a resource nobody consumes rendered as entirely free) and started reading Ceph, alongside four Ceph groups and a fill gauge on the fleet rollup. → 0.8.0: mcp.html declares its example.com allowance. claim_edge_domain’s catalogue entry uses app.example.com to show the shape of a domain a CUSTOMER owns, which tripped this hub’s placeholder gate — a failure that predated the branch that fixed it and had been failing make ci for every session. RFC 2606 reserves the name for documentation and provider.html already used the gate’s own exemption for its DNS examples: the same marker, with the reasoning written out rather than a bare pragma. A minor, because the whole 0.7.x patch range had gone to concurrent sessions. → 0.8.1: the table follows api to 0.24.4 and returns the Portal row to 0.28.9 — 0.8.0 had been cut from a tree whose table read 0.28.8 while build.yml pinned 0.28.9, and this projection is generated from that file, never maintained beside it. → 0.8.2: the MCP row gets its 0.6.3 bullet back. The conflict resolution in 0b1fc4b6 kept the peer’s COMMENT lines and took our VALUE lines — a recipe that yields a plausible file rather than an error — so a 2-line version bump silently dropped 1,022 characters of narrative from two rows. The API row was restored the same day; this one was not, and the three edge-domain tools it names have been on the served surface since 0.6.3. → 0.8.5: the two version histories stop skipping releases — the MCP row gains 0.6.2, 0.7.0 and 0.7.1, and this row gains 0.7.3 through 0.8.1. Eight releases had moved a version CELL without leaving a line about what shipped in them. Different cause from 0.8.2 and the same shape: the cell is a PROJECTION of build.yml and follows a pin on its own, while the sentence beside it is written by hand — so the two drift apart silently and the table reads as though those releases carried nothing. (0.8.3 and 0.8.4 are spent, not skipped. 0.8.3 was a peer’s build for the same 0.6.3 restoration this hub shipped as 0.8.2. 0.8.4 was this text with one interval wrong — it put an hour between the truncation and the API row’s repair where the commits are sixteen minutes apart — caught before its pin moved, so it was rebuilt rather than corrected under a tag that had already shipped.) → 0.9.0: the table follows api to 0.25.0 and the MCP row to 0.7.2 — phase 48’s OU tree and the guardrails that inherit down it. → 0.9.3: the pair is gated. A row’s CELL is generated from build.yml and follows a pin by itself; the sentence beside it is written by hand and was checked by nobody, which is how eight releases and one truncation went unnoticed. So audit_version_narratives() in gen-doc-tables.py, with a twin in the offline lane so a push cannot carry it either, now fails when a row’s narrative never mentions its own cell. It found SIX rows red the day it was written — api, factory, billing, MCP, portal and this one — and those are the entries added above. Deliberately not a contiguity check: this trace is curated (build.yml has pinned 98 docs versions and 55 are narrated), so demanding every version would be a gate that fails without looking. (0.9.1 belongs to a concurrent session. 0.9.2 is SPENT: it was built and imported 3/3 carrying this same text with the Billing line quoting a rate the engine does not charge — caught by that very guard before the pin moved, so it was rebuilt rather than corrected under a shipped tag.) → 0.9.4: carries the Portal row’s 0.29.0 line — the console’s project selector. A hub edit is a release because the hub is SERVED bytes: leaving the tree ahead of the image is exactly the served-vs-declared gap 0.4.9 and 0.4.13 were burned for. → 0.9.5: mcp.html stops naming the console’s rungs in FRENCH inside English prose — “reserved to the Propriétaire/Administrateur rungs”, where the console itself renders Owner and Administrator, so the sentence named a control the reader could not find. Landing with it, this page finally DECLARES the example.com allowance that its own mcp 0.8.0 entry describes: a changelog cannot record that a page declared an allowance without writing the word, and the placeholder gate had been red on origin/mox since that entry was written. And the Portal row follows to 0.29.1 — the release in which the console stopped serving French inside the zones its own translator is told to skip. → 0.9.6: carries the Portal row’s 0.30.0 line — project creation from the console. A hub edit is a release because the hub is SERVED bytes. → 0.9.7: and the 0.31.0 line beside it — the CSP had been blocking the guard that kept those forms from navigating. → 0.9.8: carries the API row’s 0.26.0 line — project deletion and account closure. → 0.9.10: the Portal row above carries 0.34.0 and the sign-out change behind it. (0.9.8 and 0.9.9 are both spent — a concurrent session holds one and this work’s pre-rebase cut holds the other, and both are imported on the nodes.) → 0.9.11: carries the API row’s 0.26.2 line — the two defects phase 49 shipped and live validation caught. → 0.9.12: and 0.26.3 beside it — the third, found by running the destructive gate twice. → 0.9.13: the Portal row above carries 0.35.0 and the project-selector fix behind it. → 0.9.14: the Portal row above carries 0.35.1 — the responsive release: centred pages with ascending width tiers, pinned table headers, the drawer following the route both ways. → 0.9.15: carries the API row’s 0.26.4 line — a refused purge now takes its ig1-deleting tag back off. → 0.9.16: the API row above carries 0.26.5 and the deletion-cancel verb. → 0.9.17: the API row above carries 0.27.0. → 0.9.19: the API card, the reference meta and the version table follow to 0.27.1 — the direct-attach release (185.255.84.178). Two tags spent on the way, both to the same mid-flight landing: a SECOND 0.9.17 was built and imported by this branch before the pin above landed (one tag, two byte-sets — ctr images ls cannot tell them apart), and 0.9.18, this branch's correction on top of its own 0.9.17, never saw its pin land at all. → 0.9.20: the API row above carries 0.28.0. → 0.9.21: the API row above carries 0.29.0. → 0.9.22: the API row above carries 0.29.1 — the ACME-lane redirect fix, found by the first live authorization after the flip. → 0.9.23: the API card and reference meta follow to 0.29.1, and the flip gotcha renumbers to 301 — a peer landed their own 300 mid-rebase, the numbering collision the worktree-scan rule exists for. (0.9.22 is spent: built and imported 3/3 with the 0.29.0 card strings still inside.) → 0.9.24: the flip gotcha settles at 302 — TWO peers took 300 and 301 during this branch's rebase cycles, which is what the numbering guard exists to catch at push. (0.9.23 spent, one number stale inside.) → 0.9.25: the issuer-migration sweep — server.json and every page that names the identity host follow to zitadel.cloud.ig1.com, and the version table carries api 0.29.2 / portal 0.35.2 / mcp 0.7.3. → 0.9.26: the table follows api 0.29.3 and portal 0.35.3 — the two respins the CORS/CSP and deny-default guards forced before anything shipped. → 0.9.27: the MCP card, the API card and the reference meta all follow the respins (0.7.3 / 0.29.3) — the hand-carried strings the hub's own tests refuse to let drift. (0.9.26 spent.) → 0.9.28: the MCP lede too (0.7.3) — the one hand-carried string 0.9.27 missed. (Spent.) → 0.9.29: the table follows portal 0.35.4. (0.9.28 spent with it.) → 0.9.30: the API card, the reference meta and the spec banner follow api 0.30.0, and the api + portal rows narrate the waitlist move and the tenant-census fix behind it. → 0.9.31: the table follows api 0.30.1 / portal 0.36.1 — the naming completion release. → 0.9.32: the API card and meta follow to 0.30.1. (0.9.31 spent.) → 0.9.33: the API card, the reference meta and the spec banner follow api 0.30.2, and the api + portal rows narrate the onboarding tier tag, the taken-organisation refusal and the signup page's 409. → 0.9.34: the guides speak PUBLIC — 46 customer-facing endpoint examples across ten pages move to cloud.ig1.com, the auth guide says plainly that the public endpoints are Let's Encrypt-trusted (the internal-CA ritual is lab-only now), and server.json hands agents the public MCP endpoint with an honest TLS note. The CLI's baked defaults follow (api/factory/billing/mcp → public) — a customer on the internet uses it with zero configuration. (0.9.33 is BURNED WHILE PINNED — this release's bytes were rebuilt under the landed tag when a version-cut refusal was grepped into looking like success; the 0.6.19 scar, re-earned. Pods running the old bytes stay honest only because the pin moves here.) → 0.9.36: the API card and the reference page follow api 0.32.0 — 101 paths, two of them new: the DNS delegation check and the bulk record import a customer uses to move a domain here from wherever they bought it → 0.9.37: retagged over 0.9.36, whose bytes predated the cloud.ig1.com client-surface sweep landing beside it — an image built on a stale base ships the work it is missing as absent, under a tag that claims to include it. → 0.9.38: the API card, the reference meta and the spec banner follow api 0.32.1, and the api row narrates the operator-notification relay. → 0.9.39: the Billing row follows to 0.6.1 — the release in which the billing service enforces the billing-access rule itself, instead of relying on the API gateway that fronts it. A hub edit is a release because the hub is SERVED bytes. → 0.9.40: the hub follows portal 0.39.0 and the second nameserver — ns2.cloud.ig1.com, a real address on a second host, so the registrar forms that demand two entries are satisfiable. → 0.9.41: the MCP card, its lede and server.json follow mcp 0.7.4 — the release that corrects the trust-store default that had all 28 destructive tools refusing to act behind a green /health (OPERATIONS gotcha 315). Three copies of one version number, each with its own offline gate; they are what caught this sweep before it shipped half-done. → 0.10.0: the hub stops describing a cloud that was retired on 2026-08-25. The flip made cloud.ig1.com resolve worldwide and put Let’s Encrypt on every customer surface, and the client-surface sweep at 0.9.34 moved the endpoint STRINGS — but the prose explaining them stayed pre-flip on seven pages, which is the harder half and the half a reader believes. The quickstart opened by telling a customer to pull ig1-internal-ca-secret out of the cluster with kubectl: a first step that is both unnecessary (the endpoints are publicly trusted) and impossible (no customer has cluster access), and every curl / CLI / Terraform example under it carried the --cacert that step produced. The networking guide’s §7 still read « delegated, hardened, serving, and resolving for nobody » and offered dig api.cloud.ig1.com @8.8.8.8 → REFUSED as proof — that command now returns the address, from every public resolver. Its §6.1 callout told customers a browser would warn on their own domain because « this platform has no address reachable from the internet at all », while the HTTP-01 issuer beside it has been on production Let’s Encrypt since flip day. The certificate line is now split by NAME rather than by date, which is the distinction that actually survives: the generated *.10.57.8.75.nip.io hostname keeps the internal CA permanently — it resolves into private space and no public CA can validate a name it cannot connect to — and a customer domain on the same exposure is publicly trusted. Three shipped surfaces the hub had never mentioned land with it: DNS is on all four surfaces (the migration guide still called console, CLI and Terraform « on the roadmap »), /v1/dns/zones/{zone_id}/delegation and /import make moving a domain here a flow rather than a support ticket, and webhook subscriptions have PERSISTED since 2026-08-21 while two rows still warned that a redeploy forgets them — only the delivery log is per-replica now. The SDK row splits: Python ships token_from_env(), Go and TypeScript still mint the token the way the CLI does. A minor, because a reader who trusted the old prose planned the wrong cutover. → 0.11.0: the CLI page stops asking customers to compile the CLI. Its Install section had said « Build from source (Go 1.26+) » since phase 28 — a fine instruction for a contributor and an absurd one for a customer, since it demands a Go toolchain to obtain one static binary that has no runtime dependencies at all. The page now serves linux/amd64, linux/arm64, darwin/amd64, darwin/arm64 and windows/amd64 from /downloads with SHA256SUMS, and says plainly that they are not notarized yet. The binaries are BUILT BY THIS IMAGE from source in a first stage rather than committed or staged into the context, so « the image was rebuilt » and « the download is current » are one statement; the build context moves to the repo root to reach cli/ and sdk/go. They are versioned for the first time too — phases.cli.version now exists, because the Makefile had been deriving the stamp from git describe --tags against a repo carrying a single spent tag, so every build stamped a bare commit sha and one made outside a checkout said dev. Also: the MCP catalogue grows to 135 tools (0.8.0) and the provider gains its sixth data source, ig1_dns_zone, which is the first surface anywhere that can answer whether a customer’s domain is actually delegated here. → 0.11.1: a hedge this hub wrote one release ago was already false when it deployed. 0.10.0 said « neither is an MCP tool » of get_dns_delegation and import_dns_records, verified against the served server.json (122 tools, 7 in dns) at the moment it was written. A peer landed both tools, both CLI verbs and the delegation_status attribute on the ig1_dns_zone data source that same afternoon — 135 tools — so 0.11.0 shipped carrying the sentence, now wrong, in three places. Corrected here, in the DNS intro, the surface table and the migration guide’s Route 53 row: an agent CAN drive a domain move end to end and verify it landed. The lesson is gotcha 331’s second half and it is not about DNS: a claim that something does not exist yet is a claim about OTHER PEOPLE’S in-flight work, so it must be re-read against origin/mox immediately before the build that ships it — the discipline the version number already gets, applied to the sentence. → 0.12.0: the Terraform provider stops being a build instruction too. Its Install section had said « build from source, with a dev_overrides entry » — and dev_overrides is a DEVELOPMENT mechanism: it skips terraform init entirely, warns on every plan, and makes the version constraint and the lock file do nothing at all. Recommending it to customers was worse than recommending a build, because it looks like it works. The hub is now a provider network mirror at /providers: six packages (linux, darwin, windows × amd64, arm64), SHA256SUMS, and the protocol documents that make terraform init resolve ~> 0.1, verify the h1: hashes and write a real .terraform.lock.hcl — which is what the pending registry publish is a substitute for, not a workaround around it. Three things a real terraform init taught that no document states: the mirror URL must be https, the address is normalised to lower case (a mirror written MoxForge 404s and the error reads as a missing listing), and a mirror-installed provider locks only the CALLER’S platform — so the page shows the terraform providers lock line that fixes a colleague’s broken install before it happens. Also corrected: .goreleaser.yaml built the archive binary as terraform-provider-ig1 with no _v<version> suffix — packages a registry accepts and the CLI then cannot load, invisible because make release is staged and never run. → 0.12.1: the provider page’s terraform providers lock line gains -net-mirror. Without it the command consults the ORIGIN registry — that is deliberate Terraform behaviour, not a bug — and fails with « registry.terraform.io does not have a provider named registry.terraform.io/moxforge/ig1 », which is true and is the whole reason the mirror exists. 0.12.0 shipped the command without the flag: the remediation for a single-platform lock file could not itself run. → 0.12.2: the provider page showed terraform init output nobody gets. It printed « Installed MoxForge/ig1 v0.1.0 (unauthenticated) » and explained that word at length; the real line is « Installed moxforge/ig1 v0.1.0 (verified checksum) » — lower-cased, because Terraform normalises a source address, and VERIFIED, because the mirror publishes the h1: hashes and init checks them. What is actually absent beside a registry install is the publisher signature, so the callout says that instead, and names the TLS certificate as the other half of the chain. Written from a real init against the live hub rather than from the protocol documents, which is how the first two versions of this page were wrong in three different ways. → 0.12.4: the version table carries api 0.32.3, billing 0.6.3 and mcp 0.8.2 (phase 50 + the ALB-parity billing line), and the networking guide gains the public-pool prices and the raise-on-request note. → 0.12.5: carries the API row’s 0.32.4 line — the public ingress backend becomes a health-checked set (the gotcha-355 fix), and the version table follows api 0.32.4. → 0.12.6: the api card and the reference meta follow to 0.32.4 as well — the hub is SERVED bytes, so a card edit is a release. (0.12.5 spent: built before the card strings moved.) → 0.12.7: the api and portal rows narrate 0.32.5 / 0.39.1 — the Supervision view's RabbitMQ tiles, on per-queue metrics since phase 51 (gotcha 359) — and the version table follows all three releases. → 0.12.8: the whole phase-52 story, in one release. The version table follows api 0.33.10 / portal 0.39.4 / mcp 0.8.4; the api row narrates the durability line and its four live-only repairs (LimitRange, backup egress, the backup-state reads, the ObjectStore window); the migration guide closes its RDS row — durability is a create-time choice now, the dump-cron advice is gone; the provider page documents replicas and the backup block and stops claiming custom-domain certificates are lab-CA; and the SSOT’s RPO table gains the managed-Postgres row. OPERATIONS gotchas 362–366 are the week’s diary: admission, egress, the checkout losses, the freeze mechanics, and the field that never populates. → 0.12.9: the phase-53 week, served. The MCP page documents all 136 tools incl. get_kubernetes_cluster; the server card and README say 136; the version table follows api 0.34.0 / portal 0.39.6 / mcp 0.8.5 / factory 0.5.7 — the tenant LoadBalancer line from the silent pend to the working Service, including the two spent releases (0.5.6: the cert that never presented; 0.39.5: the loop variable the smoke swept). Gotchas 367 (three documented values, each learned live) and 368 (httpx 0.28’s cert drop) are in OPERATIONS. → 0.12.12: the phase-54 story, served. (0.12.10 is spent: built while the hub’s own offline gates still had three stale strings to catch — two version tags and a narrative that cited the events history endpoint as a customer /v1 path when the service has no ingress; the gates went red, the strings were fixed, the tag was already imported. The build-before-gate ordering scar, re-earned. 0.12.11 joined it the same hour: imported with the events cell reading 0.5.3 while the pin had already moved to 0.5.4 — the regen is part of the payload, not a step that can trail the build.) The version table follows api 0.34.1 / events 0.5.3; the events row narrates the durable audit history endpoint (and the spent 0.5.2, whose aiokafka call shape was guessed rather than introspected); the api row narrates the source: ring|durable switch; the migration guide’s CloudTrail row loses its « read path on the roadmap » gap; the CLI page documents the source field; the gap register’s #7 row closes. Gotcha 369 is in OPERATIONS: the cursor that hid a sparse tenant’s rows, the sed & that ate two gate runs, and the fake that fed too fast. → 0.12.13: the phase-55 story, served. The version table follows api 0.34.2; the api row narrates the 250 GB managed-database ceiling and the tier-derived namespace quota; the migration guide’s RDS references move to 250; the gap register’s #3 row closes. No gotcha this time — the capacity answer shipped in docs/capacity-measured.md before the number moved, which is the ritual the register row asked for. → 0.12.14: the phase-56 story, served. The version table follows api 0.34.3 / events 0.5.5 / billing 0.8.0 / portal 0.39.7; the migration guide’s EC2 row and day-5 dry-run stop saying stopped is not free; the gap register’s #9 and #10 rows close together (the meter follows power state and deletion now); billing’s row narrates the durable lifecycle ledger; the retention table in the phase-23 suite records why the lifecycle topics are 7-day operational. → 0.12.15: the billing cell and row follow to 0.8.1 — the merge fix the first live proof caught. (0.12.14 is spent on the same cell: the 0.8.1 cut landed after its build for a real product fix; the regen is part of the payload, said here for the third time because the third time was not the charm either.) → 0.12.16: the api and events cells follow the retags (0.34.4 / 0.5.6 — the committed-tree rebuilds), and the portal-admin row is whole again. → 0.12.17: the SDK auth shim is three languages, served. The migration guide’s SDK row and the gap register’s #14 close: sdk/go/auth.go and sdk/typescript/src/auth.ts mirror the Python shim exactly (same grant, same scope asserted against the CLI’s constant, same cache, same refusal), all three ride the PRESERVE list, and phase 45 gained a Go and a TypeScript suite — each mints live against the lab and reads /v1/whoami back 200. The TS shim is deliberately dependency-free (ambient Node declarations — a @types/node dependency would die on every regeneration). → 0.12.18: the CLI gains auto-pagination and waiters (cli 0.2.0, gap register #21). Every list now follows limit/marker to the end of the collection — a tenant’s 1,001st object used to be silently invisible, and the register called auto-pagination « the one that silently returns wrong answers today » for exactly that reason. ig1 wait (vm/volume/lb/cluster) plus --wait on the vm/volume create/delete verbs; exit 4 timeout, exit 5 error-state. The CLI page documents both, and the exit-code table gained its two new rows. The binaries are rebuilt by THIS image, so the download is current with this release. → 0.12.19: the phase-57 story, served. The version table follows api 0.34.5 / portal 0.39.8 / mcp 0.8.6 / cli 0.2.1; the MCP catalogue gains get_instance_metrics_history (137 tools — page, card, server.json, README and the index chip all moved together); the gap register’s #8 row closes; the CLI page gains the metrics verbs. → 0.12.20: the api and mcp cells follow the retags (0.34.6 / 0.8.7). → 0.12.21: the phase-58 story, served. The version table follows mcp 0.8.8 / cli 0.2.2; the MCP catalogue gains the four backup tools (141 tools — page, card, server.json, README and the index chip all moved together); the CLI page gains the backup verbs; the gap register’s #12 row closes. The cinder-restore lesson is gotcha 371: /restores is a Nova shape, Cinder restores are an action on the backup, and the body the URL names the backup for takes only name/volume_id. → 0.12.22: the api cell follows 0.34.7 (the backup quotas in the tier table). → 0.12.23: flavor_id resizes in place and the downloads are signed. The provider gains 0.2.0 — ig1_server.flavor_id is an in-place resize (register #15a: resize → VERIFY_RESIZE → auto-confirm, proven live end to end) instead of the destroy-and-recreate every practitioner feared; the CLI gains 0.2.2; and every published artefact now carries a minisign signature (register #22) built into the image build itself — the private key is a BuildKit secret, never a layer; the public key is served at /ig1-minisign.pub and both download pages carry the one verify command. → 0.12.24: the provider page’s prose catches its own header — the sample init output and the zip’s inner binary name still said v0.1.0 against the 0.2.0 downloads, and a hand-verifying reader would have concluded the package was wrong. → 0.12.25: and the constraint itself — version = "~> 0.1" only matches < 0.2, so following the page literally answered « no matching versions » against a 0.2.0 mirror. It reads ~> 0.2 now, in the block and the sample output. → 0.12.26: the restore ship, served. The version table follows api 0.35.0 / mcp 0.8.9 / cli 0.2.3; the MCP catalogue gains restore_database (142 tools); the CLI page gains the restore verb; make validate-58 now proves both halves of the DR story live — the volume copy off the block layer AND the database recovery into a new cluster. → 0.12.27: the hub catches up to phases 59-60. It had been serving 142 tools and mentioning NAT gateways and workload identity zero times while both endpoints were live on api 0.35.0 — the same un-rebuilt-image gap as mcp 0.8.10 beside it. The version table follows api 0.35.0 / mcp 0.8.10; the MCP catalogue gains the seven NAT and workload-identity tools (149 tools); the customer-facing pages document both surfaces for the first time. → 0.12.28: the table follows portal 0.39.9 and this row’s own cell. A hub that publishes the version table is behind the moment any pin moves, which is why a version cut anywhere lands a docs cut beside it. → 0.12.29: the API row follows to 0.35.2 and the Portal row to 0.39.10 — the release in which the operator queue became usable again. No signup could be approved at all: onboarding wrote the tier’s Octavia quota at /v2/octavia/quotas, which is the amphora sub-tree, so the PUT was a 405 and the whole tenant rolled back; and the one applicant that got past it met a refusal naming a remedy the console could not send. → 0.12.30: the API row follows to 0.35.3 and the Portal row to 0.39.11 — the DISK commitment this cloud could never publish, and the stale reading behind it. → 0.12.31: the Portal row follows to 0.39.12 — the signup message that covered four unrelated causes and proposed a remedy for none of them. → 0.12.32: the API row follows to 0.35.4 and the Portal row to 0.39.13. → 0.12.33: the API row follows to 0.35.5 and the Portal row to 0.39.14. → 0.12.35: the hub follows api 0.35.7 and portal 0.39.15 — the alarm detail. This row exists because the table is the one place a reader can check what is actually served against what was built — and 0.12.34, built seven minutes before this page was finished, served the previous contract number under the new tag. → 0.12.36: the hub follows mcp 0.8.11 and portal 0.39.16. → 0.12.37: the hub follows mcp 0.9.0 — resync_edge is gone and the count is 148, restated on the landing page, the MCP page and server.json because a stated count that nothing derives is three chances to disagree. Also api 0.36.0, factory 0.6.0 and events 0.5.7, each narrating its own row here. → 0.12.38: the corrected bytes of 0.12.37, which was built before three version strings inside these pages had followed the pins beside them — the MCP lede, the landing page’s API card and api.html’s contract line. A changed tag is never re-used, so the fix ships under a new one. → 0.12.39: follows api 0.36.1. 0.12.38 was built against 0.36.0, which burned an hour later — so the hub is rebuilt rather than left naming a tag that never served. → 0.12.40: follows mcp 0.9.1. → 0.12.41: follows mcp 0.9.1 and factory 0.6.1. → 0.12.42: follows factory 0.6.2. → 0.12.43: follows factory 0.6.3. → 0.12.44: follows mcp 0.9.2. → 0.12.45: follows mcp 0.9.3. → 0.12.46: follows api 0.36.2 and factory 0.6.4. → 0.12.47: follows api 0.36.3. → 0.12.48: follows api 0.36.4. → 0.12.49: follows factory 0.6.5, and publishes the API row that 0.12.48 was built one edit too early to carry — this hub had been showing the gateway at 0.36.3 while it answered 0.36.4. → 0.12.50: follows portal 0.40.0, and the AWS-migration page gains the two things a customer test run cost on 2026-09-09 — that a tenant access key must be quoted or read from a profile because it contains a $, and that when the AWS CLI prints argument of type 'NoneType'… it is crashing while formatting RGW's error rather than reporting one, so the SDKs will tell you what actually happened. The stale AWS_CA_BUNDLE line went with it: the published endpoint is public TLS now and needs no internal CA. → 0.12.52: this table catches up the two API releases it had missed — 0.36.5 and 0.36.6; it had kept showing 0.36.4 — the ig1 CLI 0.2.4 download stops refusing database sizes above 100 GB, the Terraform provider mirror serves 0.2.1, and the MCP server’s database tools (0.9.4) stop refusing those sizes too. The miss was ours: the check that forces this hub to be rebuilt when its sources move was reading the wrong files. → 0.12.53: the Terraform provider mirror serves 0.2.2, whose NAT-gateway and workload-identity resources now carry their descriptions in the provider itself (the provider row below says what changed). → 0.12.54: this page, for API 0.37.0, console 0.41.0 and CLI 0.3.0 — “the organisation’s users” now means one organisation, and an invitation can attach someone who already has an IG1 Cloud login. The CLI follows: ig1 org user list reads one account (--project, or --all-tenants for the platform view) and ig1 org user invite gains --project, without which a platform admin would read a refusal it could not follow. → 0.12.55: 0.12.54 was built before these pages were finished, so its number described a tree it predates — rebuilt under its own tag rather than re-cut under the spent one, which is the failure nothing catches (gotcha 426 found the build path that could not run at all). → 0.12.56: API 0.38.0, console 0.41.1 and CLI 0.3.1 — the cloud owner's own page invites COLLEAGUES into IG1, not people into customers' accounts.
ig1 CLIsemver per build · schema v1phase 28 (W10) — ig1 version --client prints the pin agents record
terraform-provider-ig1~> 0.2phase 34 (W11) + the elasticity wave's ig1_cluster / ig1_asg / ig1_edge_exposure + the Barbican wave's ig1_secret — the tenant DNS wave's ig1_dns_zone / ig1_dns_record — the custom-domain wave's ig1_edge_domain_verification — the NAT gateway wave's ig1_nat_gateway — the IRSA wave's ig1_workload_identity — the 25-resource v1 inventory → 0.2.1: ig1_database’s size_gb stops telling you the limit is 1–100 — it had not been since the API moved to 250 — and says the platform bounds both size_gb and size_gb × replicas, with the live values in the API reference. If your committed lock file pins 0.2.0, run terraform init -upgrade once: the mirror serves the current release only. → 0.2.2: ig1_nat_gateway and ig1_workload_identity now describe themselves in the provider exactly as their registry pages already did — a NAT gateway is for tenant outbound internet access, workload identity is for a tenant Kubernetes cluster. Those words had only ever been typed into the generated pages, where the next regeneration would have erased them; they live in the provider’s schema now, so your editor and terraform providers schema show them too. A lock file on 0.2.0 or 0.2.1 needs terraform init -upgrade once.

2 · OpenAPI is the contract

3 · The SDK regen rule

When a service's spec changes, the SDKs are regenerated first — before any consumer is built against the new surface (the standing rule from the W10 CLI and W11 provider workstreams):

make sdk-build    # scripts/build-sdks.sh — pinned openapitools/openapi-generator-cli:v7.24.0

4 · Deprecation policy

5 · The parity gate — how CI enforces REST ≡ CLI ≡ provider ≡ MCP

  1. One source. The services' routers are the only definition of the surface; the spec is served, not transcribed.
  2. Generated consumers. CLI and provider sit on the regenerated sdk/go; the MCP server calls only served paths (gateway, factory, billing) with the caller's bearer.
  3. The docs-don't-rot gate. node services/docs/tests/docs.smoke.js (zero-dependency, offline) parses every page on this hub and asserts: every guide is linked from the index; every ig1 … command in a code block exists in the CLI's Cobra tree (parsed from cli/cmd); every /v1/… path exists in the services' routers (parsed from services/); every MCP tool in the catalog page exists in services/mcp with the right tier; every provider resource exists in the W11 spec inventory; and no placeholder hosts or tokens appear anywhere.
  4. Per-wave validators. scripts/validate-phase-*.sh prove each wave's contract live against the lab; the offline CI suite (scripts/ci/run-ci.sh, wired as make ci) runs the build/vet/drift guards on every change.
If a doc and a service disagree, the service is right — and the gate turns red. That is the whole trick: examples on this hub are checked against the source tree, so a renamed command or path breaks the build instead of rotting a page.