DRBD 9.3.4 resync hangs in WFBitMapS when the resource Primary is a diskless node — correct durable fix?

Setup: 6-node hyperconverged CloudStack + KVM cluster. LINSTOR 1.34.0, DRBD 9.3.4 (DKMS), drbd-utils 9.34.0, Ubuntu 24.04 / kernel 6.8, LVM-thin backend, 3-replica resource group. Many VMs run as LINSTOR diskless clients, so a resource’s Primary is often a diskless node.

Problem: When a diskful replica needs to (re)sync while the resource’s Primary is diskless (new-volume initial sync, or a replica rejoining after a reboot), the resync wedges and never completes:

  • drbdsetup status --statistics: SyncTarget, rs-in-flight:0, no progress.
  • dmesg: Becoming WFBitMapS because primary is diskless, Can not start OV/resync since it is already active (-8), postponing this until current resync finished.
  • Survives disconnect/connect, invalidate, down/up, and a full node reboot (restarts then re-wedges).
  • Reproduced on an all-9.3.4 cluster. Impact: the volume can’t reach ≥2 UpToDate; a brand-new disk is unusable until corrected.

What we found: making the Primary diskful clears it immediately —
linstor resource toggle-disk <vm-node> <res> --storage-pool <pool> — the WFBitMapS step goes away and the resync finishes. We believe this matches the open GitHub report LINBIT/drbd #148 (diskless-Primary + WFBitMapS), though our error signature (already active -8) differs from that report’s (receive_bitmap in Established).

Questions:

  1. Is DrbdOptions/auto-diskful <minutes> (+ auto-diskful-allow-cleanup) the recommended durable mitigation for diskless-client deployments, so the Primary is never diskless long enough to wedge? Any guidance on the timer value?
  2. Is a DRBD-level fix for the diskless-Primary resync hang expected after 9.3.4?
  3. During the brief window before auto-diskful triggers, is there a safer approach than toggle-disk to recover a wedged resync?

Can share full drbdsetup status, dmesg, and linstor output on request.