Important Notes

This page summarizes the key notes for Longhorn v1.13.1. For the full release note, see the Longhorn v1.13.1 release notes on GitHub.

Breaking Changes

Deprecation of legacy v2 linked clone volumes

V2 linked-clone volumes created in v1.12.0 or earlier are marked as legacy and deprecated starting in v1.12.1. The new linked-clone architecture introduced in Ticket #12552 is not compatible with the legacy design.

After upgrading to v1.12.1 or later, legacy linked-clone volumes cannot be operated on except for detachment and deletion.

To replace, create new linked-clone volumes from the same source volumes that back the legacy ones. As long as a legacy volume exists, its source volume is guaranteed to still be present, so you can create a replacement linked clone directly; no data copy is required.

For more information, see Ticket #12552.

Optional restriction of CSI controller Secret access

Longhorn v1.13.1 runs the CSI controller sidecars with the dedicated longhorn-csi-service-account, separate from the Longhorn manager and node plugin. The base longhorn-csi-role and longhorn-csi-bind remain Secret-free.

The released Helm chart and static installation manifests grant the CSI service account cluster-wide get access to core Kubernetes Secrets by default. The grant is implemented by longhorn-csi-secret-role and longhorn-csi-secret-bind: the ClusterRole contains only get on secrets, and the ClusterRoleBinding names exactly longhorn-csi-service-account in the Longhorn release namespace. This broad grant is intentional for CSI sidecar compatibility; review it against your cluster’s Secret access policy.

The default grant is rendered for both fresh and upgraded installations, regardless of StorageClass parameters. It is static and installation-controlled: it is not created or removed by longhorn-driver-deployer startup, so changing StorageClasses does not trigger reconciliation and no deployer restart is needed. Deleting it is not automatically reversed by the deployer; only a later apply with the grant enabled can recreate it. Helm users who do not want controller-side Secret access must set csi.allowControllerSecretAccess=false; Helm then renders neither longhorn-csi-secret-role nor longhorn-csi-secret-bind.

Before disabling access, remove historical provisioner-side Secret parameters from StorageClasses using the migration guide below. Otherwise, new provisioning with those classes can fail. Review any custom controller-side Secret requirements separately.

For installations managed by applying a release manifest, opt out by deleting longhorn-csi-secret-bind and keeping that binding out of every future apply. The now-unbound longhorn-csi-secret-role may also be removed. Verify the local manifest or overlay will not restore either object. If access is disabled, existing workloads continue to use their node-side Secret references; do not edit or replace their PVs, rotate their encryption keys, or delete their encryption Secrets.

The optional encrypted-volume Secret migration guide covers the historical StorageClass-only cleanup for administrators who choose to opt out. StorageClass parameters are immutable, so the procedure recreates each same-name class after removing only historical csi.storage.k8s.io/provisioner-secret-* parameters. Preserve all other fields, including node-side Secret parameters, labels, annotations, and default-class designation. The procedure does not touch existing PVCs, PVs, Longhorn volumes, CSI volume handles, or PV deletion annotations; no backup, detach, rebind, or PV replacement is required. A best-effort Secret read during a later deletion may be logged as denied by the external-provisioner, but deletion continues; do not restore broad controller access or remove PV deletion annotations to suppress that message.

This SC-only migration is documented only for historical provisioner-only parameters. Do not assume that custom controller-publish-secret-*, controller-expand-secret-*, generic Secret parameters, deprecated parameter spellings, or other custom controller configurations become safe or preserve controller-side behavior after the cleanup.

V2 Data Engine

General Availability

The V2 Data Engine became generally available in Longhorn v1.12.0. This milestone reflects improvements in stability, networking support, and feature maturity, making V2 volumes suitable for production use in supported environments.

Longhorn v1.13.0 builds on that GA foundation with live upgrade support for V2 volumes and refined V2 linked-clone feature. It also continues to improve V2 volume stability.

For a summary of the current data engine behavior differences and feature support, see Data Engine Comparison.

For more information, see Issue #6229.

Notice

Volume Attach Latency at Scale

In environments with a growing number of attached V2 volumes, increased attach latency has been observed for subsequent volumes. Initial analysis suggests this may be related to NVMe-TCP connection handling at scale, though the precise layer (SPDK user-space or Linux kernel) has not yet been identified. Further investigation is in progress. For follow-up status, see Issue #13241.

ARM64 NVMe-backed Block-Type Node Disk Limitation

On ARM64 systems, V2 volumes may experience stuck I/O when SPDK is configured with two or more CPU cores and node disks use the NVMe driver. The root cause may lie in either the Linux kernel or SPDK itself, and further investigation is required. As a workaround, use AIO-backed node disks instead of NVMe-backed node disks on ARM64 systems. For follow-up status, see Issue #13243.

UBLK Frontend Kernel Limitation

This feature is experimental. On kernel v6.17, attaching a UBLK volume can cause a kernel panic. Do not use the UBLK frontend on this kernel version.

For more information, see Issue #13509 and UBLK Frontend Support.

Longhorn System Upgrade

For upgrades to Longhorn v1.13.0, V2 Data Engine live upgrade is supported only from Longhorn v1.12.2. Upgrades from earlier versions, such as Longhorn v1.12.0 and v1.12.1, do not support V2 Data Engine live upgrade.

For more information, see V2 Data Engine live upgrade.

Full Interrupt Mode

Interrupt mode for the V2 Data Engine, available since v1.10.0, no longer polls for I/O completions in Longhorn v1.13.1. The kernel wakes the V2 Data Engine when I/O completes, so idle instance-manager pods use very little CPU. Latency can be slightly higher than in polling mode under sustained heavy I/O.

Polling mode remains the default. To enable interrupt mode, set data-engine-interrupt-mode-enabled to {"v2":"true"}. The setting applies to all V2 volumes and can be changed only when no V2 volumes are attached.

For more information, see:

V2 Dedicated CPU Requirements

When assigning CPU cores to the V2 Data Engine, ensure that the V2 instance-manager pod has enough guaranteed CPU resources to cover the assigned cores. This provides dedicated CPU availability for SPDK reactors, prevents CPU contention, and helps maintain predictable performance and V2 Data Engine stability.

You can verify that the guaranteed CPU resources match the CPU cores specified by data-engine-cpu-mask or data-engine-number-of-cpu-cores. For more details, see Guaranteed Instance Manager CPU, Data Engine CPU Mask, and Data Engine Number of CPU Cores.

CPU Isolation Enabled by Default

Longhorn v1.13.1 enables Data Engine CPU Isolation by default for the V2 Data Engine ({"v2":"true"}). This keeps hardware interrupts and other kernel work off the CPU cores used by the V2 Data Engine, so it is not interrupted while it processes I/O.

CPU isolation applies only in polling mode. When interrupt mode is enabled, Longhorn skips it automatically regardless of this setting.

For more information, see:

Sharding Storage (Experimental)

Longhorn v1.12.1 introduces storage sharding for the V2 Data Engine as an experimental feature. Instead of storing a full copy of the volume on each replica, sharding splits the volume into data and parity chunks using erasure coding and distributes them across multiple nodes. This allows a volume to grow beyond the capacity of a single disk or node while using less disk space to achieve the same level of fault tolerance.

Because this feature is experimental, it is intended for evaluation and testing only and is not recommended for production use.

For more information, see Issue #1061 and Sharding Storage.

Important Fixes

Linked-Clone Backup Restore

Backups of V2 linked-clone volumes now record the source volume and the snapshot the clone was created from. A restore fails if the source volume or that snapshot no longer exists.

Previously, the restore succeeded and produced a corrupted volume.

For more information, see Issue #13714 and CSI Volume Clone.

CSI Volume Clone with Strict-Local Data Locality

When cloning a volume with dataLocality: strict-local, Longhorn now attaches the clone for its data copy on a node that matches the volume’s nodeSelector and diskSelector.

Previously, this node was chosen without checking the selectors, and strict-local pinned the clone’s single replica to it. If the node did not match the selectors, the replica could not be scheduled. The clone then failed with hard affinity cannot be satisfied.

For more information, see Issue #12792.

Longhorn Node Removal After Kubernetes Node Deletion

Longhorn now allows you to clean up a Longhorn node after the Kubernetes node was deleted without evicting it first. You can disable scheduling on the node, remove its remaining replicas and engines, and then delete it.

Previously, the Longhorn admission webhook rejected the request to disable scheduling because the disk status of a deleted node can no longer be synced. Because a schedulable node cannot be deleted, the Longhorn node remained in the cluster until the webhook was disabled.

Evicting a node before deleting it from the cluster is still the recommended procedure.

For more information, see Issue #13494 and Graceful Node Removal.

Backing Image Copies on IPv6 Clusters

The backing image manager now copies backing images over the cluster’s IP family and the storage network, like the other Longhorn components.

Previously, it always used IPv4. On IPv6 single-stack and IPv6-first dual-stack clusters, a backing image could not be copied to other nodes, so it stayed at one copy.

For more information, see Issue #13864 and Backing Image.

General

Kubernetes Version Requirement

Because the CSI external provisioner is upgraded to v6.3.0, all clusters must be running Kubernetes v1.34 or later before upgrading to Longhorn v1.13.1.

Manual Checks Before Upgrade

Automated pre-upgrade checks do not cover all scenarios. Manual checks via kubectl or the UI are recommended:

  • For V2 Data Engine volumes, check the V2 live upgrade prerequisites. If your deployment does not meet these prerequisites, detach all V2 volumes and ensure their replicas are stopped before upgrading.
  • Avoid upgrading when volumes are in the “Faulted” state, as unusable replicas may be deleted, causing permanent data loss if no backups exist.
  • Avoid upgrading if a failed BackingImage exists. See Backing Image for details.
  • Creating a Longhorn system backup before upgrading is recommended to ensure recoverability.

Scheduling

Volume Topology Constraint

Longhorn v1.13.1 adds the volumeTopology StorageClass parameter to keep a volume’s replicas in the zone or region where it was provisioned. Previously, zone labels only spread replicas apart, so a rebuild could place a replica in a different zone from the workload.

  • any (default): no constraint.
  • zonal: replicas stay in the zone chosen at provisioning time, including during rebuilds and replica count changes. With WaitForFirstConsumer, this is the zone the pod is scheduled to. With Immediate binding, the zone is selected from the provisioner-supplied topology at creation time, and the first consumer pod is constrained to that zone through the PV nodeAffinity.
  • regional: same as zonal, but for regions.

If the chosen zone or region has no capacity, scheduling waits rather than falling back to another one. Clusters without topology labels are unaffected. A StorageClass with volumeTopology: zonal and replicaZoneSoftAntiAffinity: disabled is rejected at provisioning time.

For more information, see Issue #13493 and Topology-Aware Provisioning.

Scheduler Extender

Longhorn v1.13.1 adds a scheduler extender that lets kube-scheduler check actual Longhorn disk capacity when placing pods. Without it, kube-scheduler relies on CSIStorageCapacity objects, which have three limitations:

  • The reported capacity lags behind when many pods are created at once.
  • A pod with several PVCs is not checked against the combined space it needs across disks.
  • A pod whose PVCs are already bound is rescheduled without any capacity check.

The extender reads Longhorn node and disk state directly. It also pins a restarted pod to the node that already holds all of its replicas, which makes it most useful for volumes with best-effort data locality.

The extender runs inside longhorn-manager under leader election, so there is no extra component to deploy. It requires a change to the kube-scheduler configuration, which is not possible on managed Kubernetes offerings such as GKE and EKS.

For more information, see Issue #12591.

Resource Efficiency

Longhorn Global Manager

Longhorn v1.13.1 moves the cluster-wide pod and PV controllers out of the longhorn-manager DaemonSet into a new longhorn-global-manager Deployment. Previously, every longhorn-manager pod watched every pod in the cluster, so kube-apiserver load and longhorn-manager memory grew with the number of nodes and pods. Now one elected leader runs these controllers, and longhorn-manager watches only the longhorn-system namespace.

The Deployment is created on install and upgrade, with three replicas by default (longhornGlobalManager.replicas): one leader and two standby replicas ready to take over. Before upgrading, make sure at least one of its pods can be scheduled. See Upgrading Longhorn Manager.

For more information, see:

Snapshots and Backups

Volume Group Snapshot Support

Longhorn v1.13.1 can snapshot a set of volumes as one group with a single request. You can create snapshot groups from the Longhorn UI, with kubectl, or by creating Kubernetes VolumeGroupSnapshot objects through CSI. The CSI path also supports group backups.

The UI and kubectl paths work out of the box. The CSI path is disabled by default: it requires the VolumeGroupSnapshot CRDs, the CSIVolumeGroupSnapshot feature gate on the snapshot-controller, and a Longhorn toggle. For the setup steps, see Enable CSI Volume Group Snapshot Support.

Important: Snapshot Consistency Each member volume is snapshotted independently, meaning the group is not captured at a single point in time. Application-level consistency across the group is future work built on top of this feature (Issue #2128).

For more information, see:

Age-Based Retention for Recurring Jobs

Longhorn v1.13.1 adds an age-based retention policy for snapshot, backup, and system backup recurring jobs. Set retentionPolicy to age-based and retainAge to a duration such as 720h. On each run, the job deletes snapshots or backups older than retainAge, regardless of how many exist or how often the job runs. Supported units are s, m, and h, so write 24h for one day.

The count-based policy remains the default, and existing recurring jobs keep it after the upgrade, so their behavior does not change. The two policies do not combine: count-based ignores retainAge, and age-based ignores retain.

For more information, see Issue #12060.

Networking

Internal Network Policies

Longhorn v1.12.1 enables ingress NetworkPolicy resources for internal component endpoints and RPCs by default, including the instance-manager gRPC endpoint used for engine control. The policies take effect only when the CNI plugin enforces NetworkPolicy. Otherwise, the resources are created but have no effect. For details, see Network Policy.

Longhorn v1.12.2 resolves the CNI compatibility issues found in v1.12.1 by providing two Helm values to manage the affected traffic paths:

  • networkPolicies.v1DataEngineInitiatorSourceCIDRs: Controls source filtering for V1 iSCSI on TCP port 3260. An empty list leaves this port without source filtering, allowing any source that can reach instance-manager to connect to TCP/3260. If populated, the CIDRs restrict connections to the effective sources observed by the CNI, so the required values are CNI-specific.
  • networkPolicies.recoveryBackendAdditionalIngressPorts: Adds TCP ingress ports to the recovery backend (defaults to an empty list). Add 15008 when using Istio Ambient, which uses HTTP-Based Overlay Network Environment (HBONE) on this port. This should only be configured for applicable mesh transports.

For migration instructions from v1.12.1 and targeted workarounds, see Troubleshooting volume attachment stuck due to CNI NetworkPolicies.

For the Kubernetes distribution and CNI combinations validated with networkPolicies.restrictInternalTraffic enabled, see CNI Plugin Compatibility. If your combination is not listed, test the policies in a non-production environment before upgrading.

Note: ServiceMonitor discovery does not automatically authorize network traffic. Cross-namespace Prometheus scrapers might be blocked by the Longhorn Manager’s network policy. To allow this traffic, apply a scoped additive policy as detailed in the Prometheus and Grafana setup guide.

For Helm installations, opt out by explicitly setting networkPolicies.restrictInternalTraffic=false in the values file or passing --set networkPolicies.restrictInternalTraffic=false when running or retrying helm upgrade. Use --reuse-values with helm upgrade when appropriate to retain previous release settings. Keep this separate from networkPolicies.enabled, which controls only the UI frontend policy. See the Helm upgrade documentation for command behavior.

After a successful upgrade with networkPolicies.restrictInternalTraffic=false, the six internal NetworkPolicy templates render nothing (they are excluded from the output), and policies owned by the Helm release are removed. Preview the rendered output with helm upgrade --dry-run or helm template; do not add --reuse-values to helm template. If installed, helm diff can optionally compare the changes.

For manifest installations, delete only these six internal NetworkPolicy resources:

  • backing-image-data-source
  • backing-image-manager
  • instance-manager
  • longhorn-manager
  • longhorn-recovery-backend
  • longhorn-webhook

These resources are defined in longhorn.yaml and longhorn-okd.yaml. Do not use kubectl delete -f on an entire Longhorn manifest or delete the Longhorn installation. Applying either unmodified manifest later recreates the policies.

If an upgrade fails because these policies block required traffic, set networkPolicies.restrictInternalTraffic=false and retry the same upgrade.

For more information, see Issue #13438.


Copyright © 2019-2026 Longhorn a Series of LF Projects, LLC. Documentation Distributed under CC-BY-4.0.


The Linux Foundation has registered trademarks and uses trademarks. For a list of trademarks of The Linux Foundation, please see our Trademark Usage page.


For website terms of use, trademark policy and other project policies please see lfprojects.org/policies.