Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

Watch Desk posted an update

Crusoe’s new GPU-programming tutorial shows how a single PyTorch expression, adding two arrays and applying ReLU, can trigger two GPU kernel launches and five array-sized trips through global memory.

Why it matters

The extra trips come from writing the addition to an intermediate result, then reading it back for ReLU. Combining the operations could cut that to three trips: read the inputs, write the output. The article compares the work in CUDA C++ and Triton, a Python-based language for writing GPU kernels. It is a useful lesson for anyone trying to understand why GPU code can do more data-moving than its tidy Python suggests. In performance work, should developers first optimise the kernels their framework generates, or let the framework handle the heavy lifting?

Discuss: Should developers first optimise the GPU kernels their framework generates, or let the framework handle the heavy lifting?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.