Bug description
When a Tenant namespace is deleted while its Tenant is being reconciled (routine with short-lived CI namespaces), the ResourceQuota sync fails with
update of outer ResourceQuota items has failed: resourcequotas "capsule-<tenant>-0" not found
and the whole Tenant is requeued, on both the capsule/tenants and the capsule/tenant-resourcequotas controllers, although nothing is left to do for that quota.
Root cause, in internal/controllers/tenant/resourcequotas.go at 4c30567c:
syncResourceQuotas lists the Tenant's ResourceQuotas through the API reader (L176).
- For each resource name it calls
resourceQuotasUpdate → persistResourceQuotaState, which reads every listed quota back (L469) and updates it.
- If the namespace was deleted in between, the
Get returns NotFound, errgroup.Wait() surfaces it (L487) and Reconcile returns the joined error → requeue with backoff.
The create/update path already skips namespaces that are gone or terminating since #2104 (L391: IsNotFound || NamespaceTerminatingCause → return nil), but the read-back/update path in persistResourceQuotaState has no such guard. Same family as #2101.
How to reproduce
- A Tenant with
spec.resourceQuotas.items and a few dozen namespaces.
- Continuously create and delete namespaces of that Tenant (a CI pipeline does it for us: 377 namespace terminations in 7 days, 134 distinct names).
- Watch the manager logs:
Reconciler error … update of outer ResourceQuota items has failed: resourcequotas "capsule-<tenant>-0" not found, from both controllers.
Nothing is actually broken: the surviving quotas are updated (the errgroup runs every goroutine to completion) and the next reconcile is clean because the List no longer returns the vanished quota. But every hit is a full requeue of the Tenant, and ours has 97 namespaces.
Expected behavior
A ResourceQuota that disappeared between the List and the read-back is skipped, exactly like a namespace that is gone or terminating is skipped in syncResourceQuotas. No error, no requeue.
Logs
{"level":"error","ts":"2026-09-19T05:07:11.013Z","msg":"Reconciler error","controller":"capsule/tenant-resourcequotas","namespace":"","name":"opencell","reconcileID":"c95d91a4-6e68-4a44-9a49-5319d90db13d","error":"sync ResourceQuotas: resource quota 0: update of outer ResourceQuota items has failed: resourcequotas \"capsule-opencell-0\" not found"}
{"level":"error","ts":"2026-09-19T00:43:45.185Z","msg":"Reconciler error","controller":"capsule/tenants","controllerGroup":"capsule.clastix.io","controllerKind":"Tenant","Tenant":{"name":"opencell"},"namespace":"","name":"opencell","error":"cannot sync resourcequota items: resource quota 0: update of outer ResourceQuota items has failed: resourcequotas \"capsule-opencell-0\" not found"}
Measured over 7 days on one cluster (one Tenant with 97 namespaces churned by CI):
- 121 occurrences of this error (4 to 26 per day), against 377 namespace terminations.
- 73 % of the errors fall inside a namespace termination window, while such windows cover 16 % of the time (×4.4 over chance); the rest is consistent with the 60 s sampling of the termination start.
- 31 % of the terminations (115/377) produced at least one such error.
- Tightest example: namespace
opencell-selenium-regression-16x fully gone at 00:44:47Z, error logged at 00:44:46Z — the quota vanished while the sync was in flight.
Additional context
- Capsule version: v0.14.6 (first observed on v0.14.2;
persistResourceQuotaState is unchanged on main at 4c30567c)
- Helm Chart version: 0.14.6
- Kubernetes version: v1.35.0
A one-line fix (IgnoreNotFound on the read-back and on the update) with a unit test that fails without it follows as a PR.
Bug description
When a Tenant namespace is deleted while its Tenant is being reconciled (routine with short-lived CI namespaces), the ResourceQuota sync fails with
and the whole Tenant is requeued, on both the
capsule/tenantsand thecapsule/tenant-resourcequotascontrollers, although nothing is left to do for that quota.Root cause, in
internal/controllers/tenant/resourcequotas.goat4c30567c:syncResourceQuotaslists the Tenant's ResourceQuotas through the API reader (L176).resourceQuotasUpdate→persistResourceQuotaState, which reads every listed quota back (L469) and updates it.Getreturns NotFound,errgroup.Wait()surfaces it (L487) andReconcilereturns the joined error → requeue with backoff.The create/update path already skips namespaces that are gone or terminating since #2104 (L391:
IsNotFound || NamespaceTerminatingCause→return nil), but the read-back/update path inpersistResourceQuotaStatehas no such guard. Same family as #2101.How to reproduce
spec.resourceQuotas.itemsand a few dozen namespaces.Reconciler error … update of outer ResourceQuota items has failed: resourcequotas "capsule-<tenant>-0" not found, from both controllers.Nothing is actually broken: the surviving quotas are updated (the errgroup runs every goroutine to completion) and the next reconcile is clean because the List no longer returns the vanished quota. But every hit is a full requeue of the Tenant, and ours has 97 namespaces.
Expected behavior
A ResourceQuota that disappeared between the List and the read-back is skipped, exactly like a namespace that is gone or terminating is skipped in
syncResourceQuotas. No error, no requeue.Logs
Measured over 7 days on one cluster (one Tenant with 97 namespaces churned by CI):
opencell-selenium-regression-16xfully gone at 00:44:47Z, error logged at 00:44:46Z — the quota vanished while the sync was in flight.Additional context
persistResourceQuotaStateis unchanged onmainat4c30567c)A one-line fix (
IgnoreNotFoundon the read-back and on the update) with a unit test that fails without it follows as a PR.