Deep dives on kernels, serving infrastructure, benchmarks and the open models we run, written by the team building Mixlayer.
We built MixBench to test whether today's best models can write fast, correct fused CUDA kernels across 87 real inference workloads. The short answer: not yet, and the gap tells us a lot about where inference optimization is headed.
Research and engineering notes from the Mixlayer team. No spam, unsubscribe anytime.