Gaming the Hardware Lottery
Niranjan Baskaran ⋅ Brent Burdick
Abstract
Significant advances in AI have come from software–hardware codesign. Transformers are less data efficient than LSTMs but make up for it because parallelism saturates tensor cores. FlashAt- tention noticed tensor cores idled while moving the score matrix to HBM and back, and used online softmax to keep it on chip. At each layer of the stack we adjust our algorithms so they use GPUs better. The transformer is an example of an architecture winning the hardware lottery, but clearly we can intentionally design algorithms that effectively use compute. So how do we game the lottery?
Chat is not available.
Successful Page Load