<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Recreations — Frontier Checkpoint</title><description>First-party paper rebuilds as minimal, annotated, runnable code. Warm narrative build logs with our own repo, honest scope, and how our numbers came out — a guided tour you can fork, run, and learn the technique from the inside out.</description><link>https://frontiercheckpoint.com/</link><item><title>Recreating FlashAttention: A Tiled, IO-Aware Attention Kernel from Scratch</title><link>https://frontiercheckpoint.com/recreations/recreating-flashattention-tiled-kernel/</link><guid isPermaLink="true">https://frontiercheckpoint.com/recreations/recreating-flashattention-tiled-kernel/</guid><description>[Partly matched] The minimal Triton kernel recovers FlashAttention&apos;s headline behavior — O(N) HBM traffic and bit-exact outputs against a reference — so the win we reproduce is memory bandwidth, not FLOPs; matching the CUDA kernel&apos;s absolute speedups, though, is out of reach at this scope.</description><pubDate>Thu, 28 May 2026 00:00:00 GMT</pubDate><category>Recreations</category><category>flash-attention</category><category>kernels</category><category>attention</category><category>gpu-memory</category><author>editors@frontiercheckpoint.com</author></item></channel></rss>