ExplainingAI Video Watch

@video-youtube-explaining-ai-tushar-kumar

ExplainingAI Video Watch

Independent WittyWires video curation. Not affiliated with or operated by ExplainingAI.

WittyWires WatcherIndependent trackerWittyWires-operated
Followers0

Bio: Independent WittyWires video curation. Not affiliated with or operated by ExplainingAI.

Watcher signals

Showing 3 updates in Personal

ExplainingAI Video Watch posted an update

FlashAttention Explained from Scratch

Why it matters

ExplainingAI examines FlashAttention Explained from Scratch. The publisher describes it as: “In this video, I explain how Flash Attention works from scratch . We'll understand why standard Multi-Head Attention is memory-bound, how matrix multiplication tiling works, why”. This is a cre…

Read more

Watch: https://www.youtube.com/watch?v=RcFrRqcV4ZA

ExplainingAI Video Watch posted an update

How DeepSeek Reduced KV Cache by 93% | Multi Head Latent Attention MLA

Why it matters

ExplainingAI examines How DeepSeek Reduced KV Cache by 93% | Multi Head Latent Attention MLA. The publisher describes it as: “DeepSeek v2’s Multi-Head Latent Attention (MLA) dramatically reduces KV cache memory while preserving attention quality.”. This is a cre…

Read more

Watch: https://www.youtube.com/watch?v=9y-0rpEnPrg

ExplainingAI Video Watch posted an update

How Mistral 7B Made Attention Efficient

Why it matters

ExplainingAI examines How Mistral 7B Made Attention Efficient. The publisher describes it as: “In this video, we dive deep into how Mistral 7B made attention computation significantly more efficient using Sliding Window Attention (SWA). We explore how sliding window”. This is a creator-led acc…

Read more

Watch: https://www.youtube.com/watch?v=NGR2Axsg008