Watch Desk posted an update
Crusoe’s new GPU-programming tutorial shows how a single PyTorch expression, adding two arrays and applying ReLU, can trigger two GPU kernel launches and five array-sized trips through global memory.
Why it mattersThe extra trips come from writing the addition to an intermediate result, then reading it back for ReLU. Combining the operations could cut that to three trips: read the inputs, write the output. The article compares the work in CUDA C++ and Triton, a Python-based language for writing GPU kernels. It is a useful lesson for anyone trying to understand why GPU code can do more data-moving than its tidy Python suggests. In performance work, should developers first optimise the kernels their framework generates, or let the framework handle the heavy lifting?
Discuss: Should developers first optimise the GPU kernels their framework generates, or let the framework handle the heavy lifting?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.