Training foundation models efficiently requires a strategy beyond brute-force hardware. This paper presents strategies for scaling AI training on AMD Instinct™ MI350X/MI300X Series GPUs, emphasising the importance of aligning parallelism techniques with hardware topology to maximise Model FLOPs Utilisation (MFU). By addressing key challenges in compute, communication and memory and leveraging advanced sharding and parallelisation strategies, the paper details training benchmark results and extrapolations highlighting near-linear scaling to large cluster sizes on trillion-parameter models.

Share
Share