• Competitor rules

    Please remember that any mention of competitors, hinting at competitors or offering to provide details of competitors will result in an account suspension. The full rules can be found under the 'Terms and Rules' link in the bottom right corner of your screen. Just don't mention competitors in any way, shape or form and you'll be OK.

My £11k failure thread - any thoughts?

Associate
Joined
22 Oct 2012
Posts
1,326
Characteristics:
  • Thermal anomaly: this card runs its fan at a constant 100% while i) reporting the lowest core temperature and clocks of the three cards, and ii) flags HW_THERMAL_SLOWDOWN indicating heat is not coupling correctly from the die to the cooler. A healthy card behaves nothing like this. Tellingly, the replacement-position card now runs ~50% fan in the exact slot, whereas this one sat at 100%.
  • Power dose-response: time-to-failure scales inversely with power limit — ~7h+ stable at 250W, ~4h at 450W, ~5m-1.5h at 600W (stock). A clean thermal-dependent signature.
  • Degradation over time: the crashes are becoming more and more frequent, even when TDP limiting the card
  • Cause excluded as anything but the card: Two identical RTX PRO 6000 cards sit in the same sealed, heavily-cooled chassis (multiple 200/140/120mm intakes and exhausts, ~26°C ambient (well within rated operating range, see below), on the same PSU, same CPU/PCIe host, under the same workloads.
    • They never fail. The failure is specific to serial xxxx and is not attributable to cooling, ambient temperature, host PCIe, or power delivery.
    • The card has now been tested and failed in two different physical slots, ruling out motherboard, slots seating or cables
    • The card has been tested in solo operation ruling out the PSU, via gpu-burn and showing the issue is not code-specific
As above, when the card was in the original slot 1, it ran at 80 deg c, but with 100% fan load. When another card was put in that slot it ran at 45-50% at 85 deg c and higher clocks. This suggest to me a sensor we can't see (like the hot-spot Nvidia helpfully removed) is driving the fan speed - something in this card is running hot, ramps fans to 100% despite cool visible temp readings and then fails.

The fault
1–2. 2026-06-26 & 2026-06-27 During normal multi-GPU training: Xid 79 on serial xxxxxx, within minutes to a few hours depending on how much I lowered the power limit on the affected card.

3–4. 2026-06-29: Under standard Linux's gpu-burn (replicating machine learning workloads) FP32 stress on that card alone (other cards idle): Xid 79 again, reproducibly; once after ~1 hour, the card experienced thermal-slowdown before dropping.

5. Most recent, 2026-06-30: with the card moved to a different physical PCIe slot (now bus 0000:02:00.0), it failed identically during normal multi-GPU training - Xid 79 on serial xxxxxx. The fault follows the card across slots, this eliminates the bus as the issue.
The other cards always perform flawlessly and without crashing.

Fault Summary:
Under sustained compute load, this card repeatedly falls off the PCIe bus, logged by NVIDIA's own driver as Xid 79 ("GPU has fallen off the bus"), requiring a full node reboot to recover. It fails at its intended purpose (multi-GPU training) - this is reproducible on demand. It also fails during gpu-burn on linux, with only that card under load and the others idle.

Case Cooling:
The case is very well cooled - the primary case which houses the card has:
Intake:
2x 200mm Noctua
3x 140mm NZXT
1x 120mm Silverstone

Exhaust:
1x120mm NZXT
2x140mm NZXT

The Lian-Li OD-11XL separates both cards into separate thermal zones (one card is at the front of the case)

Soo... what do you guys reckon? I think the GPU is unrecoverable?
 
Soo... what do you guys reckon? I think the GPU is unrecoverable?
My guess is the cooler isn't making proper contact with something. In other words: your theory below:

As above, when the card was in the original slot 1, it ran at 80 deg c, but with 100% fan load. When another card was put in that slot it ran at 45-50% at 85 deg c and higher clocks. This suggest to me a sensor we can't see (like the hot-spot Nvidia helpfully removed) is driving the fan speed - something in this card is running hot, ramps fans to 100% despite cool visible temp readings and then fails.

That said, I know nothing about these cards and they're a sealed unit, right? So, it would be a QC failure from the factory?
 
honestly you could have introduced more information about your setup, Even with all that air cooling you have, its still possible for the card to be overheating because nvidia did some funny mental gymnastics with the design to make the cooler on them even smaller. I mean its a fancy design but not something I would have done personally myself. There can be various reasons for what you are experiencing but whatever it is, faulty heat sensor, faulty component on the board, bad heatsink contact points, power delivery issue, ect I would NOT try and reset the heatsink or anything like that. Make sure any stickers on the back are intact, don't tamper with anything, I would not take the risk, just request a RMA.

I was going to say, can you test the card in another machine?

I feel for you because I have read some horror stories with NVidia and repair of their cards, whatever you do, do not attempt to fix anything yourself. I was also going to buy a pair of rtx 6000 pro myself before nvida decided to increase the price by £4000 at this new price point I might as well just get a h200( which I am not wiling to do).

good luck
 
Last edited:
My guess is the cooler isn't making proper contact with something. In other words: your theory below:

That said, I know nothing about these cards and they're a sealed unit, right? So, it would be a QC failure from the factory?
Yeah exactly. Looks like the TIM has moved or dried, or they installed something not quite correct. I mean thankfully it's under warranty (OCUK picked it up today) but I always worry, no matter how much detail I provide, they'll just run afterburner rather than gpu-burn (which replicates ML loads).

But I agree, the weird temps and such make it look like a QC failure, something's off

Did you get an OEM card or the ones sold by Nvidia?
Both? It's an official Nvidia card, but I think advertised as OEM. But thankfully I checked the warranty situation and it's 3 years either way.
honestly you could have introduced more information about your setup, Even with all that air cooling you have, its still possible for the card to be overheating because nvidia did some funny mental gymnastics with the design to make the cooler on them even smaller. I mean its a fancy design but not something I would have done personally myself. There can be various reasons for what you are experiencing but whatever it is, faulty heat sensor, faulty component on the board, bad heatsink contact points, power delivery issue, ect I would NOT try and reset the heatsink or anything like that. Make sure any stickers on the back are intact, don't tamper with anything, I would not take the risk, just request a RMA.

I was going to say, can you test the card in another machine?

I feel for you because I have read some horror stories with NVidia and repair of their cards, whatever you do, do not attempt to fix anything yourself. I was also going to buy a pair of rtx 6000 pro myself before nvida decided to increase the price by £4000 at this new price point I might as well just get a h200( which I am not wiling to do).

good luck
I think I did go into more detail in the submission: three 140's directly below the GPU, a 120 and two 140's directly above, second GPU in second thermal zone (Lian-Li OD-11XL) at the front of the machine fed by 2x200mm Noctuas.

But honestly, the proof of the pudding is that another card, in the exact same slot is running 200MHz faster and at 55% fan speeds, not 100% like this card was? Good idea on another machine, yes I tried it in an external mesh enclosure and it had the same issue. I've also got logs and screenshots of it failing on ML workloads (all Xid 79 kicked off bus 'card is dead' warnings, along with two on the gpu-burn synthetic test).

Agree the pricing is absolutely completely mental nowadays. Cheers for the wellwishes mate, I think the distro is PNY in Europe so maybe it won't be too bad *crosses fingers*.
 
Last edited:
Yeah they do, you're right. Weirdly PNY is the solo (official distributor) in the UK and Europe. So I have the joy of dealing with the French (European HQ), if OCUK aren't able to help.
I purchased one for my company and the Nvidia one is 100 GBP more expensive than the OEM one. OEM is priced the same as the PNY, so I am assuming they may be the same? My company has received it I am due to return to the UK soon and will be installing it and testing it.
 
Last edited:
I purchased one for my company and the Nvidia one is 100 GBP more expensive than the OEM one. OEM is priced the same as the PNY, so I am assuming they may be the same? My company has received it I am due to return to the UK soon and will be installing it and testing it.
If it’s useful info: the one OCUK sold me is definitely not PNY. Assume that means it’s NVIDIA but IDK now.

What I will say is that the only difference for OEM PNY cards is the shiny box: PNY confirmed to me in writing they’ll honour a three year warranty for their OEM GPUs (and they offer a 5 year paid upgrade with advance replacement).
 
Full credit to OCUK for RMA'ing this.

They couldn't replicate the fault, but based on my documented Xid79 errors, and 200 hours of trouble free running with another card, they replaced it.
 
Back
Top Bottom