Recreating FlashAttention: A Tiled, IO-Aware Attention Kernel from Scratch
The minimal Triton kernel recovers FlashAttention's headline behavior — O(N) HBM traffic and bit-exact outputs against a reference — so the win we reproduce is memory bandwidth, not FLOPs; matching the CUDA kernel's absolute speedups, though, is out of reach at this scope.