Linstor: all secondary resources vanished after upgrading drbd-utils and linstor-proxmox

Something very strange is going on here. I’m on a three-node Linstor+Proxmox cluster (for historical reasons the nodes are called “virtual1”, “virtual5” and “virtual6”) and about to upgrade everything to latest.

SHORT VERSION

Everything was fine yesterday.

This morning I did an “apt-update” which upgraded only two packages: drbd-utils (9.34.3->9.35.0) and linstor-proxmox (8.3.1->8.3.2). Now I find all my drbd resources are missing everywhere except on the primary. As a result, the primary cannot connect to the secondaries. e.g.

root@virtual1:/home/brian# linstor v l -r pm-4b8bd822 -p
+--------------------------------------------------------------------------------------------------------------+
| Resource    | Node     | StoragePool | VolNr | MinorNr | DeviceName    | Allocated | InUse |    State | Repl |
|==============================================================================================================|
| pm-4b8bd822 | virtual1 | sata_ssd    |     0 |    1017 | /dev/drbd1017 | 50.02 GiB |       |  Unknown |      |
| pm-4b8bd822 | virtual5 | sata_ssd    |     0 |    1017 | /dev/drbd1017 | 50.02 GiB |       |  Unknown |      |
| pm-4b8bd822 | virtual6 | sata_ssd    |     0 |    1017 | /dev/drbd1017 | 50.02 GiB | InUse | UpToDate |      |
+--------------------------------------------------------------------------------------------------------------+

root@virtual6:/home/brian# drbdadm status pm-4b8bd822
pm-4b8bd822 role:Primary
  disk:UpToDate open:yes
  virtual1 connection:Connecting
  virtual5 connection:Connecting

root@virtual5:/home/brian# drbdadm status pm-4b8bd822
pm-4b8bd822: No such resource
Command 'drbdsetup status pm-4b8bd822' terminated with exit code 10

root@virtual1:/home/brian# drbdadm status pm-4b8bd822
pm-4b8bd822: No such resource
Command 'drbdsetup status pm-4b8bd822' terminated with exit code 10

root@virtual6:/home/brian# ls -l /dev/drbd1017
brw-rw---- 1 root disk 147, 1017 Oct  7 02:43 /dev/drbd1017
root@virtual5:/home/brian# ls -l /dev/drbd1017
ls: cannot access '/dev/drbd1017': No such file or directory
root@virtual1:/home/brian# ls -l /dev/drbd1017
ls: cannot access '/dev/drbd1017': No such file or directory

systemctl status shows “degraded”, with a failure of “drbd-graceful-shutdown”:

root@virtual1:/home/brian# systemctl --failed
  UNIT                           LOAD   ACTIVE SUB    DESCRIPTION
● drbd-graceful-shutdown.service loaded failed failed ensure all DRBD resources shut down gracefully at system shut down

LOAD   = Reflects whether the unit definition was properly loaded.
ACTIVE = The high-level unit activation state, i.e. generalization of SUB.
SUB    = The low-level unit activation state, values depend on unit type.
1 loaded units listed.

root@virtual1:/home/brian# journalctl -eu drbd-graceful-shutdown --no-pager
...
Jun 22 21:41:29 virtual1 systemd[1]: Finished drbd-graceful-shutdown.service - ensure all DRBD resources shut down gracefully at system shut down.
Oct 07 01:54:47 virtual1 systemd[1]: Stopping drbd-graceful-shutdown.service - ensure all DRBD resources shut down gracefully at system shut down...
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: linstor_db: State change failed: (-12) Device is held open by someone
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: additional info from kernel:
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: failed to demote
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: /dev/drbd1002 open_cnt:1, writable:1; list of openers follows
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: drbd1002 opened by mount (pid 657672) at 2025-10-20 10:00:55.657
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: pm-335713f4: State change failed: (-12) Device is held open by someone
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: additional info from kernel:
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: failed to demote
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: /dev/drbd1007 open_cnt:1, writable:1; list of openers follows
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: drbd1007 opened by kvm (pid 6625) at 2025-10-07 11:14:10.582
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: pm-de1bd0da: State change failed: (-12) Device is held open by someone
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: additional info from kernel:
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: failed to demote
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: /dev/drbd1012 open_cnt:1, writable:1; list of openers follows
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: drbd1012 opened by kvm (pid 13502) at 2025-10-07 11:22:59.215
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: pm-f00a673d: State change failed: (-12) Device is held open by someone
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: additional info from kernel:
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: failed to demote
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: /dev/drbd1011 open_cnt:1, writable:1; list of openers follows
Oct 07 01:54:47 virtual1 drbd-service-shim.sh[2912938]: drbd1011 opened by kvm (pid 6076) at 2025-10-07 11:13:35.689
Oct 07 01:54:48 virtual1 systemd[1]: drbd-graceful-shutdown.service: Control process exited, code=exited, status=11/n/a
Oct 07 01:54:48 virtual1 systemd[1]: drbd-graceful-shutdown.service: Failed with result 'exit-code'.
Oct 07 01:54:48 virtual1 systemd[1]: Stopped drbd-graceful-shutdown.service - ensure all DRBD resources shut down gracefully at system shut down.

(Those timestamps are UTC-7).

My guess is that the upgrade of either drbd-utils or proxmox-linstor triggered a “drbd-graceful-shutdown”, which tried to close everything; the ones it couldn’t close were the ones which KVM had open (fortunately!)

Now I need to tell linstor to recreate the missing resources, but I can’t see how. It’s not linstor resource activate:

root@virtual1:/home/brian# linstor resource activate virtual5 pm-4b8bd822
INFO:
    Resource is already activated. Noop
root@virtual1:/home/brian# linstor r l -r pm-4b8bd822 -p
+--------------------------------------------------------------------------------------------------+
| ResourceName | Node     | Layers       | Usage | Conns                         |    State | Vote |
|==================================================================================================|
| pm-4b8bd822  | virtual1 | DRBD,STORAGE |       |                               |  Unknown | Yes  |
| pm-4b8bd822  | virtual5 | DRBD,STORAGE |       |                               |  Unknown | Yes  |
| pm-4b8bd822  | virtual6 | DRBD,STORAGE | InUse | Connecting(virtual5,virtual1) | UpToDate | Yes  |
+--------------------------------------------------------------------------------------------------+

Nor is it “mkavail”:

root@virtual1:/home/brian# linstor resource mkavail virtual5 pm-4b8bd822
SUCCESS:
    Resource already deployed as requested
root@virtual1:/home/brian# linstor r l -r pm-4b8bd822 -p
+--------------------------------------------------------------------------------------------------+
| ResourceName | Node     | Layers       | Usage | Conns                         |    State | Vote |
|==================================================================================================|
| pm-4b8bd822  | virtual1 | DRBD,STORAGE |       |                               |  Unknown | Yes  |
| pm-4b8bd822  | virtual5 | DRBD,STORAGE |       |                               |  Unknown | Yes  |
| pm-4b8bd822  | virtual6 | DRBD,STORAGE | InUse | Connecting(virtual5,virtual1) | UpToDate | Yes  |
+--------------------------------------------------------------------------------------------------+

Anyway: I’m now in a state where I’m running without redundancy, because the drbd resources have vanished from the non-primary nodes; I need linstor to recreate all the missing resources.

I am wondering if restarting the satellites might do this, but I don’t want to risk breaking things unnecessarily.

Any clues please??

ADDITIONAL INFO

/var/log/apt/history.log shows (with timestamps UTC-7):

### virtual1
Start-Date: 2026-10-07  01:54:47
Commandline: apt dist-upgrade
Requested-By: brian (20040)
Upgrade: linstor-proxmox:amd64 (8.3.1-1, 8.3.2-1), drbd-utils:amd64 (9.34.3-1, 9.35.0-1)
End-Date: 2026-10-07  01:54:55

### virtual5
Start-Date: 2026-10-07  01:54:54
Commandline: apt dist-upgrade
Requested-By: brian (20040)
Upgrade: linstor-proxmox:amd64 (8.3.1-1, 8.3.2-1), drbd-utils:amd64 (9.34.3-1, 9.35.0-1)
End-Date: 2026-10-07  01:55:04

### virtual6
Start-Date: 2026-10-07  01:54:58
Commandline: apt dist-upgrade
Requested-By: brian (20040)
Upgrade: linstor-proxmox:amd64 (8.3.1-1, 8.3.2-1), drbd-utils:amd64 (9.34.3-1, 9.35.0-1)
End-Date: 2026-10-07  01:55:07

Note that I did not upgrade drbd itself, nor linstor: these are pinned.

root@virtual1:/home/brian# dpkg-query -l | grep -v ^ii
Desired=Unknown/Install/Remove/Purge/Hold
| Status=Not/Inst/Conf-files/Unpacked/halF-conf/Half-inst/trig-aWait/Trig-pend
|/ Err?=(none)/Reinst-required (Status,Err: uppercase=bad)
||/ Name                                             Version                                     Architecture Description
+++-================================================-===========================================-============-================================================================================
hi  drbd-dkms                                        9.2.15-1                                    all          RAID 1 over TCP/IP for Linux module source
hi  drbd-reactor                                     1.9.0-1                                     amd64        Monitors DRBD resources via plugins.
hi  linstor-common                                   1.32.3-1                                    all          DRBD distributed resource management utility
hi  linstor-controller                               1.32.3-1                                    all          DRBD distributed resource management utility
hi  linstor-satellite                                1.32.3-1                                    all          DRBD distributed resource management utility
hi  openjdk-17-jre-headless:amd64                    17.0.16+8-1~deb12u1                         amd64        OpenJDK Java runtime, using Hotspot JIT (headless)

Now the drbd resources exist only on the “primary” node where the VM is running. It is attempting to connect to the other nodes, but because they don’t exist there, there’s nothing to connect to.

root@virtual1:/home/brian# linstor r l -p
+--------------------------------------------------------------------------------------------------+
| ResourceName | Node     | Layers       | Usage | Conns                         |    State | Vote |
|==================================================================================================|
| linstor_db   | virtual1 | DRBD,STORAGE | InUse | Connecting(virtual5,virtual6) | UpToDate | Yes  |
| linstor_db   | virtual5 | DRBD,STORAGE |       |                               |  Unknown | Yes  |
| linstor_db   | virtual6 | DRBD,STORAGE |       |                               |  Unknown | Yes  |
| pm-4b8bd822  | virtual1 | DRBD,STORAGE |       |                               |  Unknown | Yes  |
| pm-4b8bd822  | virtual5 | DRBD,STORAGE |       |                               |  Unknown | Yes  |
| pm-4b8bd822  | virtual6 | DRBD,STORAGE | InUse | Connecting(virtual5,virtual1) | UpToDate | Yes  |
... truncated because of discourse message size limit ...
+--------------------------------------------------------------------------------------------------+

root@virtual1:/home/brian# linstor v l -p
+------------------------------------------------------------------------------------------------------------------------+
| Resource    | Node     | StoragePool          | VolNr | MinorNr | DeviceName    |  Allocated | InUse |    State | Repl |
|========================================================================================================================|
| linstor_db  | virtual1 | sata_ssd             |     0 |    1002 | /dev/drbd1002 |    204 MiB | InUse | UpToDate |      |
| linstor_db  | virtual5 | sata_ssd             |     0 |    1002 | /dev/drbd1002 |    204 MiB |       |  Unknown |      |
| linstor_db  | virtual6 | sata_ssd             |     0 |    1002 | /dev/drbd1002 |    204 MiB |       |  Unknown |      |
...
| pm-4b8bd822 | virtual1 | sata_ssd             |     0 |    1017 | /dev/drbd1017 |  50.02 GiB |       |  Unknown |      |
| pm-4b8bd822 | virtual5 | sata_ssd             |     0 |    1017 | /dev/drbd1017 |  50.02 GiB |       |  Unknown |      |
| pm-4b8bd822 | virtual6 | sata_ssd             |     0 |    1017 | /dev/drbd1017 |  50.02 GiB | InUse | UpToDate |      |
...
+------------------------------------------------------------------------------------------------------------------------+

root@virtual1:/home/brian# ls -l /dev/drbd????
brw-rw---- 1 root disk 147, 1002 Oct 20  2025 /dev/drbd1002
brw-rw---- 1 root disk 147, 1007 Oct  7 02:12 /dev/drbd1007
brw-rw---- 1 root disk 147, 1011 Oct  7 02:12 /dev/drbd1011
brw-rw---- 1 root disk 147, 1012 Oct  7 02:12 /dev/drbd1012

root@virtual5:/home/brian# ls -l /dev/drbd????
brw-rw---- 1 root disk 147, 1008 Oct  7 02:12 /dev/drbd1008
brw-rw---- 1 root disk 147, 1009 Oct  7 02:12 /dev/drbd1009
brw-rw---- 1 root disk 147, 1015 Oct  7 02:12 /dev/drbd1015

root@virtual6:/home/brian# ls -l /dev/drbd????
brw-rw---- 1 root disk 147, 1001 Oct  7 02:12 /dev/drbd1001
brw-rw---- 1 root disk 147, 1003 Oct  7 02:12 /dev/drbd1003
brw-rw---- 1 root disk 147, 1005 Oct  7 02:09 /dev/drbd1005
brw-rw---- 1 root disk 147, 1006 Oct  7 02:12 /dev/drbd1006
brw-rw---- 1 root disk 147, 1013 Oct  7 02:12 /dev/drbd1013
brw-rw---- 1 root disk 147, 1017 Oct  7 02:12 /dev/drbd1017

root@virtual1:/home/brian# drbdadm status
linstor_db role:Primary
  disk:UpToDate open:yes
  virtual5 connection:Connecting
  virtual6 connection:Connecting

pm-335713f4 role:Primary
  disk:UpToDate open:yes
  virtual5 connection:Connecting
  virtual6 connection:Connecting

pm-de1bd0da role:Primary
  disk:UpToDate open:yes
  virtual5 connection:Connecting
  virtual6 connection:Connecting

pm-f00a673d role:Primary
  disk:UpToDate open:yes
  virtual5 connection:Connecting
  virtual6 connection:Connecting

root@virtual5:/home/brian# drbdadm status
pm-4628dd67 role:Primary
  disk:UpToDate open:yes
  virtual1 connection:Connecting
  virtual6 connection:Connecting

pm-ca812bf8 role:Primary
  disk:UpToDate open:yes
  virtual1 connection:Connecting
  virtual6 connection:Connecting

pm-eb40ca57 role:Primary
  disk:UpToDate open:yes
  virtual1 connection:Connecting
  virtual6 connection:Connecting

root@virtual6:/home/brian# drbdadm status
pm-4b8bd822 role:Primary
  disk:UpToDate open:yes
  virtual1 connection:Connecting
  virtual5 connection:Connecting

pm-72c2f543 role:Primary
  disk:UpToDate open:yes
  virtual1 connection:Connecting
  virtual5 connection:Connecting

pm-9711fa32 role:Primary
  disk:UpToDate open:yes
  virtual1 connection:Connecting
  virtual5 connection:Connecting

pm-a565d280 role:Primary
  disk:UpToDate open:yes
  virtual1 connection:Connecting
  virtual5 connection:Connecting

pm-af4db968 role:Primary
  disk:UpToDate open:yes
  virtual1 connection:Connecting
  virtual5 connection:Connecting

pm-d5eb1b28 role:Primary
  disk:UpToDate open:yes
  virtual1 connection:Connecting
  virtual5 connection:Connecting

Yet, I can’t see linstor advising any fixes:

root@virtual1:/home/brian# linstor advise resource -p
+---------------------------------+
| Resource | Issue | Possible fix |
|=================================|
+---------------------------------+
root@virtual1:/home/brian# linstor advise maintenance virtual1 -p
+---------------------------------+
| Resource | Issue | Possible fix |
|=================================|
+---------------------------------+

Linstor logs don’t look interesting:

root@virtual1:/home/brian# journalctl -S yesterday -eu linstor-satellite | grep -v SpaceInfo
Oct 06 01:15:00 virtual1 Satellite[657601]: 2026-10-06 01:15:00.502 [MainWorkerPool-16] INFO  LINSTOR/Satellite/13fa08 SYSTEM - Storage pool sata_ssd reports capacity 3750756352 kiB, allocated space 1639997440 kiB
Oct 06 01:15:00 virtual1 Satellite[657601]: 2026-10-06 01:15:00.502 [MainWorkerPool-16] INFO  LINSTOR/Satellite/13fa08 SYSTEM - SpaceTracking: Satellite aggregate capacity is 3750756352 kiB, allocated capacity is 2110758912 kiB, no errors
Oct 06 03:00:57 virtual1 Satellite[657601]: 2026-10-06 03:00:57.912 [MainWorkerPool-15] INFO  LINSTOR/Satellite/3428cb SYSTEM - LogArchive: Running log archive on directory: /var/log/linstor-satellite
Oct 06 03:00:57 virtual1 Satellite[657601]: 2026-10-06 03:00:57.912 [MainWorkerPool-15] INFO  LINSTOR/Satellite/3428cb SYSTEM - LogArchive: No logs to archive.
root@virtual1:/home/brian# journalctl -S yesterday -eu linstor-controller | egrep -v '/QryAllSizeInfo|/metrics' | tail
Oct 06 02:37:10 virtual1 Controller[657680]: 2026-10-06 02:37:10.056 [MainWorkerPool-10] INFO  LINSTOR/Controller/13fad7 SYSTEM - Assembled error reports; count 4
Oct 06 02:38:10 virtual1 Controller[657680]: 2026-10-06 02:38:10.056 [MainWorkerPool-2] INFO  LINSTOR/Controller/13fad9 SYSTEM - Assembled error reports; count 4
Oct 06 02:39:10 virtual1 Controller[657680]: 2026-10-06 02:39:10.055 [MainWorkerPool-9] INFO  LINSTOR/Controller/0910bc SYSTEM - Assembled error reports; count 4
Oct 06 02:40:10 virtual1 Controller[657680]: 2026-10-06 02:40:10.055 [MainWorkerPool-10] INFO  LINSTOR/Controller/13fadf SYSTEM - Assembled error reports; count 4
Oct 06 02:41:10 virtual1 Controller[657680]: 2026-10-06 02:41:10.055 [MainWorkerPool-2] INFO  LINSTOR/Controller/13fae1 SYSTEM - Assembled error reports; count 4
Oct 06 02:42:10 virtual1 Controller[657680]: 2026-10-06 02:42:10.055 [MainWorkerPool-10] INFO  LINSTOR/Controller/0910c4 SYSTEM - Assembled error reports; count 4
Oct 06 02:43:10 virtual1 Controller[657680]: 2026-10-06 02:43:10.055 [MainWorkerPool-10] INFO  LINSTOR/Controller/13fae7 SYSTEM - Assembled error reports; count 4
Oct 06 02:44:10 virtual1 Controller[657680]: 2026-10-06 02:44:10.055 [MainWorkerPool-2] INFO  LINSTOR/Controller/13fae9 SYSTEM - Assembled error reports; count 4
Oct 06 02:45:10 virtual1 Controller[657680]: 2026-10-06 02:45:10.055 [MainWorkerPool-9] INFO  LINSTOR/Controller/13faeb SYSTEM - Assembled error reports; count 4
Oct 06 02:46:10 virtual1 Controller[657680]: 2026-10-06 02:46:10.054 [MainWorkerPool-9] INFO  LINSTOR/Controller/13faef SYSTEM - Assembled error reports; count 4

I don’t think the “error reports” are relevant, there’s nothing this year:

root@virtual1:/home/brian# linstor error-reports list -p | head -5
+----------------------------------------------------------------------------------------------------------------------------------------+
| Id                    | Datetime            | Node       | Exception                                                                   |
|========================================================================================================================================|
| 68152067-6AA5E-000000 | 2025-05-06 07:25:38 | S|virtual1 | StorageException: Failed to get physical devices for volume group: sata_ssd |
| 68152067-6AA5E-000001 | 2025-05-06 07:25:38 | S|virtual1 | ApiRcException: Failed to query free space from storage pool                |
root@virtual1:/home/brian# linstor error-reports list -p | tail -5
| 6863FBAC-4149C-000327 | 2025-08-22 11:02:32 | S|virtual6 | ApiRcException: Failed to query free space from storage pool                |
| 6863FBA3-6AA5E-000330 | 2025-08-22 11:02:32 | S|virtual1 | ApiRcException: Failed to query free space from storage pool                |
| 6863FBB0-B63BE-000328 | 2025-08-22 11:02:32 | S|virtual5 | ApiRcException: Failed to query free space from storage pool                |
| 6863FBAC-4149C-000328 | 2025-08-22 11:02:32 | S|virtual6 | ApiRcException: Failed to query free space from storage pool                |
+----------------------------------------------------------------------------------------------------------------------------------------+

Logs in dmesg show things disconnecting. For example, take resource drbd1017 / pm-4b8bd822, which is primary on virtual6.

root@virtual1:/home/brian# dmesg | egrep 'drdb1017|pm-4b8bd822'
...
[22421606.626005] drbd pm-4b8bd822/0 drbd1017 virtual6: Resync done (total 21 sec; paused 0 sec; 1736564 K/sec)
[22421606.626369] drbd pm-4b8bd822/0 drbd1017 virtual6: pdsk( Inconsistent -> UpToDate ) repl( PausedSyncS -> Established ) [resync-finished]
[22421606.639262] drbd pm-4b8bd822/0 drbd1017 virtual6: resync-susp( peer -> no ) [peer-state]
[22434199.952877] drbd pm-4b8bd822: Preparing remote state change 2663799604: 2->all role( Primary )
[22434199.962073] drbd pm-4b8bd822 virtual6: Committing remote state change 2663799604 (primary_nodes=5)
[22434199.962407] drbd pm-4b8bd822 virtual6: peer( Secondary -> Primary ) [remote]
[22434200.988005] drbd pm-4b8bd822: Preparing cluster-wide state change 2934531304: 0->all role( Secondary )
[22434200.989426] drbd pm-4b8bd822: State change 2934531304: primary_nodes=4, weak_nodes=FFFFFFFFFFFFFFF8
[22434200.989835] drbd pm-4b8bd822: Committing cluster-wide state change 2934531304 (2ms)
[22434200.990226] drbd pm-4b8bd822: role( Primary -> Secondary ) [kvm:810599 auto-demote]
[31528363.040934] drbd pm-4b8bd822: Preparing cluster-wide state change 3498089884: 0->1 conn( Disconnecting )
[31528363.042160] drbd pm-4b8bd822: State change 3498089884: primary_nodes=4, weak_nodes=FFFFFFFFFFFFFFF9
[31528363.042552] drbd pm-4b8bd822: Committing cluster-wide state change 3498089884 (2ms)
[31528363.042859] drbd pm-4b8bd822 virtual5: conn( Connected -> Disconnecting ) peer( Secondary -> Unknown ) [down]
[31528363.043101] drbd pm-4b8bd822/0 drbd1017 virtual5: pdsk( UpToDate -> DUnknown ) repl( Established -> Off ) [down]
[31528363.043480] drbd pm-4b8bd822 virtual5: Terminating sender thread
[31528363.043995] drbd pm-4b8bd822 virtual5: Starting sender thread (peer-node-id 1)
[31528363.059533] drbd pm-4b8bd822 virtual5: Connection closed
[31528363.059884] drbd pm-4b8bd822 virtual5: helper command: /sbin/drbdadm disconnected
[31528363.068976] drbd pm-4b8bd822 virtual5: helper command: /sbin/drbdadm disconnected exit code 0
[31528363.069287] drbd pm-4b8bd822 virtual5: conn( Disconnecting -> StandAlone ) [disconnected]
[31528363.069547] drbd pm-4b8bd822 virtual5: Terminating receiver thread
[31528363.069829] drbd pm-4b8bd822 virtual5: Terminating sender thread
[31528363.088211] drbd pm-4b8bd822: Preparing cluster-wide state change 1529379123: 0->2 conn( Disconnecting ) disk( Outdated )
[31528363.089869] drbd pm-4b8bd822: State change 1529379123: primary_nodes=4, weak_nodes=FFFFFFFFFFFFFFF9
[31528363.090163] drbd pm-4b8bd822 virtual6: Cluster is now split
[31528363.090534] drbd pm-4b8bd822: Committing cluster-wide state change 1529379123 (2ms)
[31528363.090823] drbd pm-4b8bd822 virtual6: conn( Connected -> Disconnecting ) peer( Primary -> Unknown ) [down]
[31528363.091076] drbd pm-4b8bd822/0 drbd1017: disk( UpToDate -> Outdated ) quorum( yes -> no ) [down]
[31528363.091387] drbd pm-4b8bd822/0 drbd1017 virtual6: pdsk( UpToDate -> DUnknown ) repl( Established -> Off ) [down]
[31528363.091811] drbd pm-4b8bd822 virtual6: Terminating sender thread
[31528363.092306] drbd pm-4b8bd822 virtual6: Starting sender thread (peer-node-id 2)
[31528363.108529] drbd pm-4b8bd822 virtual6: Connection closed
[31528363.108873] drbd pm-4b8bd822 virtual6: helper command: /sbin/drbdadm disconnected
[31528363.117899] drbd pm-4b8bd822 virtual6: helper command: /sbin/drbdadm disconnected exit code 0
[31528363.118191] drbd pm-4b8bd822 virtual6: conn( Disconnecting -> StandAlone ) [disconnected]
[31528363.118466] drbd pm-4b8bd822 virtual6: Terminating receiver thread
[31528363.118777] drbd pm-4b8bd822 virtual6: Terminating sender thread
[31528363.154213] drbd pm-4b8bd822/0 drbd1017: disk( Outdated -> Detaching ) [down]
[31528363.154561] drbd pm-4b8bd822/0 drbd1017: disk( Detaching -> Diskless ) [go-diskless]
[31528363.155185] drbd pm-4b8bd822/0 drbd1017: drbd_bm_resize called with capacity == 0
[31528363.199183] drbd /unregistered/pm-4b8bd822: Terminating worker thread
root@virtual6:/home/brian# dmesg | egrep 'drdb1017|pm-4b8bd822'
...
[4748492.397763] drbd pm-4b8bd822: Committing cluster-wide state change 2663799604 (3ms)
[4748492.398777] drbd pm-4b8bd822: role( Secondary -> Primary ) [kvm:2544356 auto-promote]
[4748493.431501] drbd pm-4b8bd822: Preparing remote state change 2934531304: 0->all role( Secondary )
[4748493.441045] drbd pm-4b8bd822 virtual1: Committing remote state change 2934531304 (primary_nodes=4)
[4748493.441927] drbd pm-4b8bd822 virtual1: peer( Primary -> Secondary ) [remote]
[13842520.253544] drbd pm-4b8bd822: Preparing remote state change 3498089884: 0->1 conn( Disconnecting )
[13842520.263272] drbd pm-4b8bd822 virtual1: Committing remote state change 3498089884 (primary_nodes=4)
[13842520.300659] drbd pm-4b8bd822: Preparing remote state change 1529379123: 0->2 conn( Disconnecting ) disk( Outdated )
[13842520.309802] drbd pm-4b8bd822 virtual1: Committing remote state change 1529379123 (primary_nodes=4)
[13842520.310172] drbd pm-4b8bd822 virtual1: conn( Connected -> TearDown ) peer( Secondary -> Unknown ) [remote]
[13842520.310531] drbd pm-4b8bd822/0 drbd1017 virtual1: pdsk( UpToDate -> Outdated ) repl( Established -> Off ) [remote]
[13842520.311009] drbd pm-4b8bd822 virtual1: Terminating sender thread
[13842520.311413] drbd pm-4b8bd822 virtual1: Starting sender thread (peer-node-id 0)
[13842520.328518] drbd pm-4b8bd822 virtual1: Connection closed
[13842520.328946] drbd pm-4b8bd822 virtual1: helper command: /sbin/drbdadm disconnected
[13842520.338006] drbd pm-4b8bd822 virtual1: helper command: /sbin/drbdadm disconnected exit code 0
[13842520.338403] drbd pm-4b8bd822 virtual1: conn( TearDown -> Unconnected ) [disconnected]
[13842520.338806] drbd pm-4b8bd822 virtual1: Restarting receiver thread
[13842520.339195] drbd pm-4b8bd822 virtual1: conn( Unconnected -> Connecting ) [connecting]
[13842521.204489] drbd pm-4b8bd822/0 drbd1017: new current UUID: C8A21055D7D19D49 weak: FFFFFFFFFFFFFFF9
[13842528.129077] drbd pm-4b8bd822: Preparing remote state change 1331942977: 1->2 conn( Disconnecting ) disk( Outdated )
[13842528.138573] drbd pm-4b8bd822 virtual5: Committing remote state change 1331942977 (primary_nodes=4)
[13842528.138951] drbd pm-4b8bd822 virtual5: conn( Connected -> TearDown ) peer( Secondary -> Unknown ) [remote]
[13842528.139271] drbd pm-4b8bd822/0 drbd1017 virtual5: pdsk( UpToDate -> Outdated ) repl( Established -> Off ) [remote]
[13842528.139714] drbd pm-4b8bd822 virtual5: Terminating sender thread
[13842528.140067] drbd pm-4b8bd822 virtual5: Starting sender thread (peer-node-id 1)
[13842528.154048] drbd pm-4b8bd822 virtual5: Connection closed
[13842528.154360] drbd pm-4b8bd822 virtual5: helper command: /sbin/drbdadm disconnected
[13842528.163321] drbd pm-4b8bd822 virtual5: helper command: /sbin/drbdadm disconnected exit code 0
[13842528.163616] drbd pm-4b8bd822 virtual5: conn( TearDown -> Unconnected ) [disconnected]
[13842528.163900] drbd pm-4b8bd822 virtual5: Restarting receiver thread
[13842528.164179] drbd pm-4b8bd822 virtual5: conn( Unconnected -> Connecting ) [connecting]
[13842529.434199] drbd pm-4b8bd822/0 drbd1017: new current UUID: ADE94B041FBC7C57 weak: FFFFFFFFFFFFFFFB

Timestamp of 31528363.040934 on virtual1 corresponds to 09:03:37 UTC this morning, and the timestamp of 13842520.253544 on virtual6 corresponds to 08:55:13 UTC. The package updates were done at 08:54:47 on virtual1 and 08:55:07 on virtual6.

Looking for clues about “drbd-graceful-shutdown”

root@virtual1:/home/brian# dpkg-query -L drbd-utils | xargs grep graceful 2>/dev/null
/lib/systemd/system/drbd-graceful-disconnect.service:# to a networkd restart or netplan apply), it gracefully disconnects
/lib/systemd/system/drbd-graceful-disconnect.service:Description=DRBD graceful disconnect on network down, reconnect on network up
/lib/systemd/system/drbd-graceful-disconnect.service:Documentation=man:drbd-graceful-disconnect.service(7)
/lib/systemd/system/drbd-graceful-disconnect.service:ExecStart=/usr/lib/drbd/scripts/drbd-service-shim.sh graceful-reconnect
/lib/systemd/system/drbd-graceful-disconnect.service:ExecStop=/usr/lib/drbd/scripts/drbd-service-shim.sh graceful-disconnect
/lib/systemd/system/drbd-graceful-disconnect.service:# drbd-graceful-disconnect.service for some reason, you can still
/lib/systemd/system/drbd-graceful-down.service:Description=DRBD graceful down on system shutdown
/lib/systemd/system/drbd-graceful-down.service:Documentation=man:drbd-graceful-down.service(7)
/lib/systemd/system/drbd-graceful-down.service:Before=drbd-graceful-disconnect.service
/lib/systemd/system/drbd-graceful-down.service:# drbd-graceful-down.service for some reason, you can still
/lib/systemd/system/drbd@.service:After=drbd-graceful-disconnect.service drbd-graceful-down.service
/lib/udev/rules.d/65-drbd.rules:ENV{SYSTEMD_WANTS}="drbd-graceful-disconnect.service drbd-graceful-down.service"
/usr/lib/drbd/scripts/drbd-service-shim.sh:GRACEFUL_STATE=/run/drbd/graceful-disconnect
/usr/lib/drbd/scripts/drbd-service-shim.sh:graceful-disconnect)
/usr/lib/drbd/scripts/drbd-service-shim.sh:graceful-reconnect)

(I find “drbd-graceful-down” but not “drbd-graceful-shutdown”; perhaps it changed name with the update)

FYI, I was able to reproduce the problem with some VMs, and it was fixed by systemctl restart linstor-satellite, which gave me enough confidence to do this in production.

This has left me with several resources that have frozen their SyncTarget at a particular percentage, e.g.

root@virtual6:/home/brian# linstor v l -p
+-----------------------------------------------------------------------------------------------------------------------------------------------------------+
| Resource    | Node     | StoragePool          | VolNr | MinorNr | DeviceName    |  Allocated | InUse  |              State | Repl                         |
|===========================================================================================================================================================|
| linstor_db  | virtual1 | sata_ssd             |     0 |    1002 | /dev/drbd1002 |    204 MiB | InUse  |           UpToDate | virtual5: Established        |
|             |          |                      |       |         |               |            |        |                    | virtual6: SyncSource         |
|-----------------------------------------------------------------------------------------------------------------------------------------------------------|
| linstor_db  | virtual5 | sata_ssd             |     0 |    1002 | /dev/drbd1002 |    204 MiB | Unused |           UpToDate | virtual6: PausedSyncS(0.15%) |
|             |          |                      |       |         |               |            |        |                    | virtual1: Established        |
|-----------------------------------------------------------------------------------------------------------------------------------------------------------|
| linstor_db  | virtual6 | sata_ssd             |     0 |    1002 | /dev/drbd1002 |    204 MiB | Unused | SyncTarget(20.22%) | virtual5: PausedSyncT(0.15%) |
|             |          |                      |       |         |               |            |        |                    | virtual1: SyncTarget(20.22%) |
|-----------------------------------------------------------------------------------------------------------------------------------------------------------|
...
|-----------------------------------------------------------------------------------------------------------------------------------------------------------|
| pm-4628dd67 | virtual1 | sata_ssd             |     0 |    1015 | /dev/drbd1015 | 250.06 GiB | Unused |           UpToDate | virtual5: Established        |
|             |          |                      |       |         |               |            |        |                    | virtual6: PausedSyncS(0.01%) |
|-----------------------------------------------------------------------------------------------------------------------------------------------------------|
| pm-4628dd67 | virtual5 | sata_ssd             |     0 |    1015 | /dev/drbd1015 | 250.06 GiB | InUse  |           UpToDate | virtual6: SyncSource         |
|             |          |                      |       |         |               |            |        |                    | virtual1: Established        |
|-----------------------------------------------------------------------------------------------------------------------------------------------------------|
| pm-4628dd67 | virtual6 | sata_ssd             |     0 |    1015 | /dev/drbd1015 | 250.06 GiB | Unused |  SyncTarget(1.85%) | virtual5: SyncTarget(1.85%)  |
|             |          |                      |       |         |               |            |        |                    | virtual1: PausedSyncT(0.01%) |
|-----------------------------------------------------------------------------------------------------------------------------------------------------------|
| pm-4b8bd822 | virtual1 | sata_ssd             |     0 |    1017 | /dev/drbd1017 |  50.02 GiB | Unused | SyncTarget(18.45%) | virtual5: PausedSyncT(0.01%) |
|             |          |                      |       |         |               |            |        |                    | virtual6: SyncTarget(18.45%) |
|-----------------------------------------------------------------------------------------------------------------------------------------------------------|
| pm-4b8bd822 | virtual5 | sata_ssd             |     0 |    1017 | /dev/drbd1017 |  50.02 GiB | Unused |           UpToDate | virtual6: Established        |
|             |          |                      |       |         |               |            |        |                    | virtual1: PausedSyncS(0.01%) |
|-----------------------------------------------------------------------------------------------------------------------------------------------------------|
| pm-4b8bd822 | virtual6 | sata_ssd             |     0 |    1017 | /dev/drbd1017 |  50.02 GiB | InUse  |           UpToDate | virtual5: Established        |
|             |          |                      |       |         |               |            |        |                    | virtual1: SyncSource         |
|-----------------------------------------------------------------------------------------------------------------------------------------------------------|
...

(plus some more). The actual percentage completion figures from drdbadm status disagree:

root@virtual1:/home/brian# drbdadm status linstor_db
linstor_db role:Primary
  disk:UpToDate open:yes
  virtual5 role:Secondary
    peer-disk:UpToDate
  virtual6 role:Secondary
    replication:SyncSource peer-disk:Inconsistent done:97.65

root@virtual6:/home/brian# drbdadm status pm-4b8bd822
pm-4b8bd822 role:Primary
  disk:UpToDate open:yes
  virtual1 role:Secondary
    replication:SyncSource peer-disk:Inconsistent done:18.45
  virtual5 role:Secondary
    peer-disk:UpToDate

but they are still frozen. However I expect I can fix these with disconnects/reconnects, or at worst forced resyncs.

Regards,

Brian.


For reference, here’s how I tested it. Unfortunately, previous versions of the drbd-utils and linstor packages are not available in the linbit-ubuntu-linbit-drbd9-stack PPA, so I couldn’t reproduce it exactly, but instead just used drbdadm down <resource> manually on each of the nodes to simulate the problem I had in live.

Install 3 VMs with ubuntu 26.04: vmhost{1..3}
(without secureboot to avoid the MOK dance)
and add /etc/hosts entries for all three

Attach an extra drive as sdb
incus storage volume attach --create default vmhost1-lvm vmhost1 sdb
...

Inside VMs: create volume group
pvcreate /dev/sdb
vgcreate vg0 /dev/sdb

Inside VMs: install drbd9
add-apt-repository ppa:linbit/linbit-drbd9-stack
apt update
apt install drbd-dkms drbd-utils
modprobe drbd
cat /sys/module/drbd/version  # 9.3.4

Inside VMs: install linstor
apt install linstor-client linstor-satellite
echo -e '[global]\ncontroller=vmhost1' >/etc/linstor/linstor-client.conf

On vmhost1 only:
apt install linstor-controller
linstor node create vmhost1 --node-type combined  # ignore ZFS errors
linstor node create vmhost2
linstor node create vmhost3
linstor storage-pool create lvm vmhost1 vg0 vg0
linstor storage-pool create lvm vmhost2 vg0 vg0
linstor storage-pool create lvm vmhost3 vg0 vg0
linstor resource-group create --description "LVM 3 way replica" --storage-pool vg0 --place-count 3 rg0
linstor volume-group create rg0
linstor rg spawn-resources rg0 testvol 200M

root@vmhost1:~# linstor v l
╭─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ Resource │ Node    │ StoragePool │ VolNr │ MinorNr │ DeviceName    │ Allocated │ InUse  │    State │ Repl           │
╞═════════════════════════════════════════════════════════════════════════════════════════════════════════════════════╡
│ testvol  │ vmhost1 │ vg0         │     0 │    1000 │ /dev/drbd1000 │   204 MiB │ Unused │ UpToDate │ Established(2) │
│ testvol  │ vmhost2 │ vg0         │     0 │    1000 │ /dev/drbd1000 │   204 MiB │ Unused │ UpToDate │ Established(2) │
│ testvol  │ vmhost3 │ vg0         │     0 │    1000 │ /dev/drbd1000 │   204 MiB │ Unused │ UpToDate │ Established(2) │
╰─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

Inside a screen session on vmhost1, hold /dev/drbd1000 open:
sleep 86400 3</dev/drbd1000

On all nodes, manually delete the drbd resources to replicate the problem
(it fails on vmhost1)
drbdadm down testvol

root@vmhost1:~# linstor v l
╭───────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ Resource │ Node    │ StoragePool │ VolNr │ MinorNr │ DeviceName    │ Allocated │ InUse  │    State │ Repl │
╞═══════════════════════════════════════════════════════════════════════════════════════════════════════════╡
│ testvol  │ vmhost1 │ vg0         │     0 │    1000 │ /dev/drbd1000 │   204 MiB │ Unused │ UpToDate │      │
│ testvol  │ vmhost2 │ vg0         │     0 │    1000 │ /dev/drbd1000 │   204 MiB │        │  Unknown │      │
│ testvol  │ vmhost3 │ vg0         │     0 │    1000 │ /dev/drbd1000 │   204 MiB │        │  Unknown │      │
╰───────────────────────────────────────────────────────────────────────────────────────────────────────────╯

Then on all nodes, systemctl restart linstor-satellite, wait a few seconds and check again:

root@vmhost1:~# linstor v l
╭─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ Resource │ Node    │ StoragePool │ VolNr │ MinorNr │ DeviceName    │ Allocated │ InUse  │    State │ Repl           │
╞═════════════════════════════════════════════════════════════════════════════════════════════════════════════════════╡
│ testvol  │ vmhost1 │ vg0         │     0 │    1000 │ /dev/drbd1000 │   204 MiB │ Unused │ UpToDate │ Established(2) │
│ testvol  │ vmhost2 │ vg0         │     0 │    1000 │ /dev/drbd1000 │   204 MiB │ Unused │ UpToDate │ Established(2) │
│ testvol  │ vmhost3 │ vg0         │     0 │    1000 │ /dev/drbd1000 │   204 MiB │ Unused │ UpToDate │ Established(2) │
╰─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

It might be possible to reproduce the exact circumstance around the upgrade of drbd-utils using debian plus the PVE kernel plus the proxmox repo, since this does appear to keep all the old versions lying around.

# On a proxmox node
$ apt-cache show drbd-utils | grep Version: | head
Version: 9.35.0-1
Version: 9.34.3-1
Version: 9.34.0-1
Version: 9.33.1-1
Version: 9.33.0-1
Version: 9.32.0-1
Version: 9.31.0-1
Version: 9.30.0-1
Version: 9.29.0-1
Version: 9.28.0-1

Hello Brian,

It does seem to be an impact from the utils upgrade based on what you are seeing:

The restart of the satellite service is a valid remediation option in that case, and it does seem like it worked for you here. I didn’t see you mention the runtime drop-in so I assume you did not use it, but let me know if you did and it didn’t do the trick in this case as it is supposed to prevent the resources going down.

Have you gotten the resyncs to move along? I believe that in your output one of them is unpaused, so hoping that saturation of the link has caused them to pause but that they will resume once the traffic dies down from the ongoing resync operation.

Whoops: that is very clearly documented, but I did not see it, and hence did not use the run-time drop-in. This will teach me to read the release notes in future, even for packages which I had wrongly assumed would be unlikely to have side effects. I’ll also add drbd-utils to my set of “hold” packages so that future updates are controlled.

The stuck drbd replication didn’t clear, and I/O completely froze on one of the nodes which I ended up having to reboot. Things are working again now and I’ll have another go today at draining the nodes and upgrading them.

Many thanks for taking the trouble to read and analyse my post.