TMnight wrote: ↑19 Jul 2026, 10:56
exqo wrote: ↑18 Jul 2026, 17:15
Thank you for the acknowledgement and for escalating this as a high-priority case.
I want to raise one point clearly, without any hostility, because I think it matters for how this case is handled going forward.
Over the past months I have effectively been acting as a beta tester: capturing kernel oops traces, running MemTest86 across multiple sessions, decoding faulting instructions from register dumps, and installing and validating two custom kernel builds on my production unit. I accepted that role willingly, and I would accept it again, because I wanted to help you find the root cause.
But I am a customer, not a tester. This started with the C-state mitigation, which did not resolve the issue. Then the v1 kernel, which did not resolve it. Then the v2 kernel, which improved the frequency from every 10-18 hours to every 4-5 days, but still leaves me with a unit that dies and requires a hard power cut on a weekly basis. Each iteration has been a test run on my own hardware, with my own data at risk.
Seven months is a long time to operate a storage appliance that cannot be trusted to stay online, and my drives are taking the cost of it.
I will of course test a v3 kernel as soon as your development team has one, and I remain fully cooperative. But I would appreciate an indication of timeframe, even approximate, rather than an open-ended wait.
I would also note that a hardware exchange would benefit both sides. You mentioned earlier that you have not been able to reproduce the issue perfectly in your laboratory. My unit reproduces it reliably, roughly every four to five days. Sending me a replacement and taking this one back would give your development team a system that actually exhibits the fault, under their own instrumentation, rather than relying on logs and remote testing through me.
Thank you again for the work done so far and I do not doubt the engineering effort behind it.