any
model
New models ship every week.
Most are only performant on one hardware.
Observation #1
Optimizing a model is a slow, manual, chip-specific work. Teams do it once and for one stack, so most hardware never gets used well.
any
hardware
any
model
Observation #2
AI accelerators have exceptional raw performance, but software determines how much of it gets used.
Hardware is not the bottleneck.
Software is.
any
hardware
any
model
Hoid is the agentic
AI compiler.
A fast route between your model and your silicon.
Breakthrough
Optimization has always meant scarce, hardware-specific expertise, applied by hand. Hoid hardwires that expertise into software.
any
hardware
Method
Push silicon to its peak.
Hoid's agentic AI compiler profiles, benchmarks, and navigates the optimization space for whichever hardware world it enters.
Stats
Models should run at their best in hours, not weeks
That shouldn't depend on the silicon you use.
Team
FLOP/s-first team.
Kernels, compilers, agents. We live in the gap between models and hardware.
FAQ
How does Hoid work?
Hoid has a low latency runtime for specific hardware provider and an agentic harness which discovers best optimizations for any given model.
How can we work with you?
You can just book a call with us and we can work together!
How does a pilot with Hoid work?
You choose a model and target hardware, and we agree on the performance targets that matter for your workload. Book a call to discuss it!
Can I use Hoid if I already have inhouse inference?
Yes! We have many ways to bring performance optimizations to you. Book a call and we can discuss it.
How is Hoid different from a traditional compiler?
Hoid's Agentic Harness optimizes model by finding bottlenecks and fixing them in a continuous loop, unlike traditional AI compilers which define fixed heuristics.
Does model need to be open source?
No, we work with custom model optimizations too.
Why do I need Hoid?
If you're not satisfied with your model performance, Hoid is for you! Hoid works like an autonomous performance engineer, it finds bottlenecks, profiles and updates the model to deliver peak performance.
Didn't find your answer?
We're obsessed with optimizing AI inference.
Ask us the hard questions.