Beau
Okay, Jo. So we've built this thing. We have our Transformer block, we've got the training loop humming with AMP and a fancy learning rate scheduler, we can even generate text with a KV-cache. But... it's all running on one GPU. It feels like we built a high-performance engine and put it in a go-kart.