The people who can make a frontier model run near peak on a given accelerator number in the low thousands worldwide. Their knowledge is scarce, hard to transfer, and goes stale with the next hardware generation.
The three properties
Scarce. Hiring for it competes with every lab and every vendor.
Non-transferable. What you learn about one memory hierarchy rarely survives the move to the next one intact.
Perishable. A new accelerator or a new attention variant can invalidate a year of tuning.
That combination is why so much deployed inference runs far below what the hardware can do — and why we think the work belongs in a system that can redo it on demand.