AI: Network
Fabric resiliency and failure handling
A training job spreads across hundreds of GPUs, so the fabric must keep running through device failures. The good news of scale: *the bigger the fabric, the higher the chance something fails - but the smaller each failure's blast radius.* Spine failure - graded, not catastrophic In a non-blocking