# File corruption with NFS HA Cluster

**URL:** <https://forums.linbit.com/t/file-corruption-with-nfs-ha-cluster/192>\
**Category:** DRBD\
**Created:** [July 24, 2024, 11:29am UTC](https://forums.linbit.com/t/file-corruption-with-nfs-ha-cluster/192 "2024-07-24T11:29:23Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Argadonis](https://avatars.discourse-cdn.com/v4/letter/a/bc79bd/32.png) [@Argadonis](https://forums.linbit.com/u/Argadonis)\
**Post date:** [July 24, 2024, 11:29am UTC](https://forums.linbit.com/t/file-corruption-with-nfs-ha-cluster/192/1 "2024-07-24T11:29:23Z")

</div>

Hello everyone,

I’ve created a test environment with 3 nodes based on AlmaLinux 9 and followed the how-to guide “[NFS High Availability Clustering Using DRBD and Pacemaker on RHEL 9](https://linbit.com/tech-guide/drbd9-nfs-rhel9/)”. And I have a NFS share visible on the network.

The issue is with testing the failover as is stated in Chapter 5.

I did not create a file using “dd” but instead I copied a video file of ~2Gb and when I did “hard” reset, as exemplified in the aforementioned chapter, the cluster did its thing and jumped to one of the available nodes and the file continued copying but the resulted file got corrupted (checked the file hash as seen in the following screenshot).  
 ![Screenshot 2024-07-24 at 14.31.12](https://canada1.discourse-cdn.com/flex003/uploads/linbit/original/1X/1a1740211d940ad346817a6bb40bc9c187d9ed5a.png)

If I do the same test but instead of “hard” reset I do a normal “reboot” of the primary node the resulting file is good (checked with file hash).

I do know the reason why the file gets corrupted as the primary DRBD node was “hard” killed before was able to send the data to secondary DRBD node.

So my question is : there is a way to configure DRBD to mitigate this situation ?

Best regards

---

<div class="post-metadata">

**Author:** ![Devin](https://yyz1.discourse-cdn.com/flex003/user_avatar/forums.linbit.com/devin/32/26_2.png) [@Devin](https://forums.linbit.com/u/Devin)\
**Post date:** [July 24, 2024, 5:42pm UTC](https://forums.linbit.com/t/file-corruption-with-nfs-ha-cluster/192/2 "2024-07-24T17:42:18Z")

</div>

I suspect it might just be write buffers not getting flushed to disk with the hard reboot versus the soft one.

Try mounting the filesystem with `-o sync` option. In Pacemaker speak that would mean adding `options=sync` as a parameter to the OCF:heartbeat:Filesystem resource.

This will likely cause a noticeable hit to write performance, but test with the above and see if you can recreate.

---

<div class="post-metadata">

**Author:** ![Argadonis](https://avatars.discourse-cdn.com/v4/letter/a/bc79bd/32.png) [@Argadonis](https://forums.linbit.com/u/Argadonis)\
**Post date:** [July 24, 2024, 6:09pm UTC](https://forums.linbit.com/t/file-corruption-with-nfs-ha-cluster/192/3 "2024-07-24T18:09:57Z")

</div>

Thank you @Devin . That was the missing piece of information.

Once the `sync` option was added to the options present in the pcs command from Chapter 4.2 of the How-to Guide the file copied while doing a “hard” reset has the same hash as the original file.

The performance hit is low as I the test nodes are on a system with NVMe (I will do some test and get the numbers and see exactly the real hit on performance). I will test it on a SATA/SAS system and get back with real life stats.

---

<div class="post-metadata">

**Author:** ![vik-t](https://avatars.discourse-cdn.com/v4/letter/v/ba8739/32.png) [@vik-t](https://forums.linbit.com/u/vik-t)\
**Post date:** [September 30, 2024, 3:01pm UTC](https://forums.linbit.com/t/file-corruption-with-nfs-ha-cluster/192/4 "2024-09-30T15:01:54Z")

</div>

Hi @Argadonis, I’m curious, were you able to test it on SATA/SAS and did you collect any stats?

@Devin Do I understand correctly that in a system with enterprise disks that support PLP this would not have happened, and the `-o sync` option is not necessary?

---

<div class="post-metadata">

**Author:** ![Devin](https://yyz1.discourse-cdn.com/flex003/user_avatar/forums.linbit.com/devin/32/26_2.png) [@Devin](https://forums.linbit.com/u/Devin)\
**Post date:** [September 30, 2024, 3:40pm UTC](https://forums.linbit.com/t/file-corruption-with-nfs-ha-cluster/192/5 "2024-09-30T15:40:16Z")

</div>

As I understand it PLP is just capacitor backed disk cache. While this would certainly help in the event of a power outage, I don’t think it would do anything to protect any writes saved in system memory and not yet flushed down to the disk.

I suspect you would still want the `-o sync` option set.
