34 min read
A Tinyblog about Tinygrad
A deep dive through Tinygrad's compiler and runtime — rangeify, memory planning, beam search — and where its GPU-shaped assumptions start to bind.
Pavle Padjin
34 min read
A deep dive through Tinygrad's compiler and runtime — rangeify, memory planning, beam search — and where its GPU-shaped assumptions start to bind.
Pavle Padjin
6 min read
The weights dropped Friday. By Sunday, our harness had beaten weeks of hand-tuning, dead ends included.
Momcilo Mrkaic, Vladimir Zeljkovic
1 min read
Fifth hardware world, same one-line command.
Vladimir Zeljkovic
1 min read
Tiling, layout, fusion, scheduling and precision choices multiply fast.
Vladimir Zeljkovic
1 min read
An agent optimizing for tok/s will happily break numerics to get there.
Vladimir Zeljkovic
1 min read
Optimization expertise is scarce, non-transferable and perishable. It shouldn't be.
Momcilo Mrkaic, Pavle Padjin, Vladimir Zeljkovic
1 min read
A layer-by-layer roofline of five frontier MoE models across four accelerators.
Pavle Padjin, Vladimir Zeljkovic
1 min read
The runtime is 11k lines with no vendor SDK in the hot path.
Pavle Padjin, Vladimir Zeljkovic
1 min read
Decode latency cut 34% at EP=32. We didn't suggest it, and for two days we didn't believe it.
Vladimir Zeljkovic
1 min read
Kernel changes, driver bumps and framework upgrades quietly cost throughput.
Momcilo Mrkaic
1 min read
We quantized GLM-5 six ways and measured both the Pareto curve and the eval damage.
Momcilo Mrkaic, Pavle Padjin
1 min read
We're backed by a16z speedrun. What we're spending it on and what we're deliberately not building.
Momcilo Mrkaic