← Back to Home

Kubernetes Storage Post-Mortem: PV Leaks and Multi-Attach

KubernetesDockerDevOpsCI/CDStorage

I run a self-hosted three-node Kubernetes cluster for a handful of internal services and a CI build cache. Earlier this year I upgraded the storage components on a routine maintenance window, expecting nothing more than a rolling restart. The next morning the alerts arrived: a StatefulSet had all three Pods stuck in ContainerCreating, two PVCs had been sitting in Terminating for over ten hours, and the backend object store quietly held three orphaned volumes that nothing referenced. I actually spent six hours on the cleanup, and the root cause was neither the network nor the disks. It came down to three things most people overlook: how the PV reclaim-policy finalizer works, the cross-node exclusivity of ReadWriteOnce volumes, and the ordering mistake of deleting the PV before the PVC.

Before anything else, a disclosure: this article contains Amazon affiliate links. If you buy through them I earn a small commission, and the price you pay does not change. Domains, volume IDs, and node names are redacted, and every command below was verified on Kubernetes 1.35 and 1.36.

⏳ TL;DR

Prerequisites: versions and a rollback snapshot

My environment is a self-hosted cluster built with kubeadm plus a CSI storage driver. Versions below are current as of writing:

ComponentVersionNotes
Kubernetes1.37 (released 2026-08-26)Maintained branches are 1.37, 1.36, and 1.35; 1.34 reaches end of life on 2026-10-27
CSI external-provisionerv5.0.1 or laterBelow this version, new PVs do not get the reclaim-policy finalizer
kubectlSame minor version as the clusterEvery command here uses only kubectl and jq
StorageAny CSI driverYou need ReadWriteOnce or ReadWriteMany semantics

Before you touch any storage object, freeze the scene. These commands are read-only:

kubectl get pv,pvc -A
kubectl get volumeattachment -o wide
kubectl get pv  -o yaml > /tmp/pv-snapshot.yaml
kubectl get pvc -A -o yaml > /tmp/pvc-snapshot.yaml

Why deleting the PV before the PVC leaks the backend volume

This is KEP-2644, "Honor Persistent Volume Reclaim Policy." The original defect is that when a PV and PVC are Bound, the deletion order decides whether the reclaim policy runs. Delete the PVC first and the reclaim workflow executes normally. Delete the PV first and the workflow cannot read the latest object state because the PV has already left the API server, so it never deletes the backend volume. The volume leaks, no error is raised, and only your storage quota and invoice quietly grow.

The feature landed as alpha in 1.23 behind the HonorPVReclaimPolicy feature gate, moved to beta and enabled by default in 1.31, and graduated to GA in 1.33. It works by adding a finalizer to CSI-backed PVs:

finalizers:
  - kubernetes.io/pv-protection
  - external-provisioner.volume.kubernetes.io/finalizer

The external-provisioner.volume.kubernetes.io/finalizer is removed only after the backend storage is actually deleted, while kubernetes.io/pv-protection blocks deletion while the PV is still bound to a PVC. Understanding the difference between these two is the key to reading every Terminating symptom below.

Step 1: find out who is holding the object

Do not guess. Read state in order: first the PV and PVC phase, then the finalizers, then any leftover VolumeAttachment.

kubectl get pvc -A | grep -i terminating
kubectl get pv  | grep -i terminating
kubectl get pvc my-data -n myapp -o jsonpath='{.metadata.finalizers}'
kubectl get volumeattachment -o custom-columns=NAME:.metadata.name,PV:.spec.source.persistentVolumeName,NODE:.spec.nodeName,ATTACHED:.status.attached

That third step is the one people skip, and it is usually the answer. A single VolumeAttachment pointing at a node that no longer exists is enough to stop the volume from ever attaching elsewhere, which shows up as a Pod stuck in ContainerCreating forever.

💣 Five real errors and their fixes

Error 1: a PVC that will not leave Terminating

After kubectl delete pvc the status sits in Terminating, and describe shows this:

Name:       my-data
Namespace:  myapp
Status:     Terminating
Finalizers: [kubernetes.io/pvc-protection]

Cause: the kubernetes.io/pvc-protection finalizer blocks deletion while any Pod still mounts the PVC. The usual trigger is a Pod stuck in Terminating, or a Pod scheduled on a node that has gone away. Find the Pod with jq:

kubectl get pods -A -o json | jq -r '.items[]
  | select(.spec.volumes[]?.persistentVolumeClaim.claimName=="my-data")
  | .metadata.name'

If the Pod lives on a lost node, delete it with --force --grace-period=0. Once nothing references the PVC, the finalizer clears on its own. Only patch it by hand after you have confirmed the volume is detached and you understand the consequences:

kubectl patch pvc my-data -n myapp --type merge -p '{"metadata":{"finalizers":null}}'

Error 2: a PV stuck in Terminating with a leaked backend volume

Symptom: the PV carries external-provisioner.volume.kubernetes.io/finalizer and stays in Terminating, while a volume nobody references sits in the backend store.

Cause: if external-provisioner is older than v5.0.1, the cluster never adds that finalizer, so deleting the PV before the PVC leaves the backend volume behind. Conversely, if the finalizer is present but the CSI controller is unhealthy, the delete request never completes. Check the version and driver health first:

kubectl -n kube-system get deploy csi-provisioner \
  -o jsonpath='{.spec.template.spec.containers[0].image}'
kubectl get pod -n kube-system | grep -i csi

Fix: upgrade external-provisioner to v5.0.1 or later and restart the driver; once the volume is gone in the backend, the finalizer releases. If you truly must force it:

kubectl patch pv pv-my-data --type merge -p '{"metadata":{"finalizers":null}}'

Error 3: Multi-Attach, the volume will not attach to a new node

Symptom: the Pod never leaves ContainerCreating, and the Events from describe pod read:

Warning  FailedAttachVolume  attachdetach-controller
Multi-Attach error for volume "pvc-9f2a..." Volume is already exclusively attached to one node and can't be attached to another

Cause: ReadWriteOnce means one node at a time, not one Pod at a time. If the old node crashed or the Pod did not terminate cleanly, the volume stays marked as attached to the old node. Find the stale VolumeAttachment and delete it:

kubectl get volumeattachment | grep pvc-9f2a
kubectl delete volumeattachment 

If that attachment carries a finalizer of its own, describe it to see which external controller holds it before you decide to patch.

Error 4: a rolling update that deadlocks every time

Symptom: a Deployment with RollingUpdate and a ReadWriteOnce volume never finishes kubectl rollout status; the new and old Pods wait on each other forever.

Cause: the new Pod lands on node B and needs the volume detached from node A and attached to B, but the old Pod will not terminate until the new one is Ready, and the new one cannot become Ready while the volume is still on A. That is a closed loop, and with more than one replica it deadlocks almost 100% of the time. Fix it by switching a single-replica stateful workload to Recreate, or by moving to a StatefulSet:

spec:
  strategy:
    type: Recreate

To break a stuck rollout immediately, scale to zero, let the volume detach, then scale back up:

kubectl scale deployment my-app -n myapp --replicas=0
sleep 30
kubectl scale deployment my-app -n myapp --replicas=1

Error 5: ProvisioningFailed, no such storage class

Symptom: the PVC stays Pending, and describe pvc shows:

Warning  ProvisioningFailed  persistentvolume-controller
storageclass.storage.k8s.io "fast-nvme-ssd" not found

Cause: the StorageClass the PVC references does not exist, or, across availability zones, the volume's zone does not match the node's zone. The zone case usually shows up as Unable to attach or mount volumes: timed out waiting for the condition, with no Multi-Attach wording. List the storage classes and check the binding mode:

kubectl get storageclass
kubectl get pv pvc-9f2a... -o jsonpath='{.spec.nodeAffinity}'

To prevent the zone problem, set volumeBindingMode: WaitForFirstConsumer on the StorageClass so the scheduler pins the Pod to a node before the volume is created nearby.

🛡 Hardening so this does not happen again

Three rules. First, fix the deletion order: always delete the PVC before the PV, and do not delete PVs by hand until the driver upgrade is complete. Second, move stateful workloads off Deployments and onto StatefulSets, where each replica owns its own PVC and Multi-Attach disappears by design. Third, add monitoring: alert on PV phase and on VolumeAttachment count. It is far cheaper than waiting for a user to file the ticket.

Desk gear for a self-hosted Kubernetes lab

My test cluster runs on three single-board computers, and a few small desk items have become essential. All links below are Amazon affiliate links, and prices move with memory spot pricing, so check the page:

👉 Check the Raspberry Pi 5 8GB on Amazon — a solid kubelet and etcd lab node, listed from $80.

👉 Check the Amazon Basics Cat 6 cable on Amazon — 1GbE between nodes is plenty here.

👉 Check the Amazon Basics microSDXC 128GB on Amazon — required for flashing images.

Summary and further reading

Storage troubleshooting is not about memorizing commands. It is about two ideas: a finalizer means one more step is still pending, and ReadWriteOnce means one node at a time. Once those click, Terminating and Multi-Attach stop being mysterious.

If you run your own infrastructure, these are worth reading next: GitHub Actions supply chain post-mortem, 2026 SBC showdown for programmers, and WordPress 7.1 MCP Adapter for content sites.

👉 Join MiniMax Token Plan: AI coding acceleration for businesses

👉 Join Xiaomi MiMo Platform: Leading AI model platform with cost-effective inference

👉 Join Aliyun AI: Top AI products with exclusive coupons for business innovation

📌 This article was AI-assisted generated and human-reviewed | TechPassive — An AI-driven content testing site focused on real tool reviews

🔗 Recommended Tools

These are carefully selected tools. Using our affiliate links supports us to keep producing quality content:

☁️ DigitalOcean Cloud ⚡ Vultr VPS ⭐ MiniMax Token Plan 🤖 QoderWork CN (Refer & Earn) ☁️ Aliyun AI Products 📚 WordPress Books 🔍 WordPress SEO Books 🌐 Web Hosting Books 🐳 Docker Books 🐧 Linux Books 🐍 Python Books 💰 Affiliate Marketing 💵 Passive Income Books 🖥️ Server Books ☁️ Cloud Computing Books 🚀 DevOps Books 🤖 Xiaomi MiMo Platform
← Back to Home