Associate
- Joined
- 22 Oct 2012
- Posts
- 1,326
Characteristics:
The fault
1–2. 2026-06-26 & 2026-06-27 During normal multi-GPU training: Xid 79 on serial xxxxxx, within minutes to a few hours depending on how much I lowered the power limit on the affected card.
3–4. 2026-06-29: Under standard Linux's gpu-burn (replicating machine learning workloads) FP32 stress on that card alone (other cards idle): Xid 79 again, reproducibly; once after ~1 hour, the card experienced thermal-slowdown before dropping.
5. Most recent, 2026-06-30: with the card moved to a different physical PCIe slot (now bus 0000:02:00.0), it failed identically during normal multi-GPU training - Xid 79 on serial xxxxxx. The fault follows the card across slots, this eliminates the bus as the issue.
The other cards always perform flawlessly and without crashing.
Fault Summary:
Under sustained compute load, this card repeatedly falls off the PCIe bus, logged by NVIDIA's own driver as Xid 79 ("GPU has fallen off the bus"), requiring a full node reboot to recover. It fails at its intended purpose (multi-GPU training) - this is reproducible on demand. It also fails during gpu-burn on linux, with only that card under load and the others idle.
Case Cooling:
The case is very well cooled - the primary case which houses the card has:
Intake:
2x 200mm Noctua
3x 140mm NZXT
1x 120mm Silverstone
Exhaust:
1x120mm NZXT
2x140mm NZXT
The Lian-Li OD-11XL separates both cards into separate thermal zones (one card is at the front of the case)
Soo... what do you guys reckon? I think the GPU is unrecoverable?
- Thermal anomaly: this card runs its fan at a constant 100% while i) reporting the lowest core temperature and clocks of the three cards, and ii) flags HW_THERMAL_SLOWDOWN indicating heat is not coupling correctly from the die to the cooler. A healthy card behaves nothing like this. Tellingly, the replacement-position card now runs ~50% fan in the exact slot, whereas this one sat at 100%.
- Power dose-response: time-to-failure scales inversely with power limit — ~7h+ stable at 250W, ~4h at 450W, ~5m-1.5h at 600W (stock). A clean thermal-dependent signature.
- Degradation over time: the crashes are becoming more and more frequent, even when TDP limiting the card
- Cause excluded as anything but the card: Two identical RTX PRO 6000 cards sit in the same sealed, heavily-cooled chassis (multiple 200/140/120mm intakes and exhausts, ~26°C ambient (well within rated operating range, see below), on the same PSU, same CPU/PCIe host, under the same workloads.
- They never fail. The failure is specific to serial xxxx and is not attributable to cooling, ambient temperature, host PCIe, or power delivery.
- The card has now been tested and failed in two different physical slots, ruling out motherboard, slots seating or cables
- The card has been tested in solo operation ruling out the PSU, via gpu-burn and showing the issue is not code-specific
The fault
1–2. 2026-06-26 & 2026-06-27 During normal multi-GPU training: Xid 79 on serial xxxxxx, within minutes to a few hours depending on how much I lowered the power limit on the affected card.
3–4. 2026-06-29: Under standard Linux's gpu-burn (replicating machine learning workloads) FP32 stress on that card alone (other cards idle): Xid 79 again, reproducibly; once after ~1 hour, the card experienced thermal-slowdown before dropping.
5. Most recent, 2026-06-30: with the card moved to a different physical PCIe slot (now bus 0000:02:00.0), it failed identically during normal multi-GPU training - Xid 79 on serial xxxxxx. The fault follows the card across slots, this eliminates the bus as the issue.
The other cards always perform flawlessly and without crashing.
Fault Summary:
Under sustained compute load, this card repeatedly falls off the PCIe bus, logged by NVIDIA's own driver as Xid 79 ("GPU has fallen off the bus"), requiring a full node reboot to recover. It fails at its intended purpose (multi-GPU training) - this is reproducible on demand. It also fails during gpu-burn on linux, with only that card under load and the others idle.
Case Cooling:
The case is very well cooled - the primary case which houses the card has:
Intake:
2x 200mm Noctua
3x 140mm NZXT
1x 120mm Silverstone
Exhaust:
1x120mm NZXT
2x140mm NZXT
The Lian-Li OD-11XL separates both cards into separate thermal zones (one card is at the front of the case)
Soo... what do you guys reckon? I think the GPU is unrecoverable?