The Mixlayer blog

Research, engineering and product from the inference frontier.

Deep dives on kernels, serving infrastructure, benchmarks and the open models we run, written by the team building Mixlayer.

ResearchMixBench
Featured·Jun 18, 2026·9 min read

MixBench: frontier LLMs still can't write fast fused kernels

We built MixBench to test whether today's best models can write fast, correct fused CUDA kernels across 87 real inference workloads. The short answer: not yet, and the gap tells us a lot about where inference optimization is headed.

A
Ada Reyes
Performance Research
Company
Jun 12, 2026 · 4 min

Mixlayer earns SOC 2 Type II and ISO 27001

We completed our SOC 2 Type II and ISO 27001 audits, validating the controls behind every layer of the platform.

P
Priya N
Benchmarks
Jun 9, 2026 · 7 min

Benchmarking inference at scale: coding agents

Real-world inference benchmarks for coding agents, 31% more tokens per second at lower tail latency.

M
Marcus L
Engineering
Jun 3, 2026 · 11 min

Serving MiniMax-M3 with 1M-token context

How we served MiniMax-M3 efficiently with KV-block-major paging and no quality regressions.

W
Wei Z
Research
May 27, 2026 · 8 min

Speculative decoding that actually ships

A practical look at draft-model selection and acceptance rates across the workloads our customers run in production.

A
Ada R
Product
May 20, 2026 · 5 min

Introducing the Frontier Gateway

One tuned, observable layer to route every model call, across providers, with automatic failover.

S
Sam K
Engineering
May 14, 2026 · 9 min

Cutting cold starts to under 900ms

The snapshotting and weight-streaming tricks that let us autoscale GPU replicas without the cold-start tax.

W
Wei Z
Benchmarks
May 6, 2026 · 6 min

The fastest embeddings, measured honestly

Throughput and tail-latency numbers for our embedding endpoints, with the methodology to reproduce them.

M
Marcus L
Company
Apr 29, 2026 · 5 min

Why we're betting on open models

Our thesis on why open-weight models win for serious inference, and what that means for the stack we build.

P
Priya N
Company
Mar 14, 2026 · 4 min

Hello Mixlayer

Introducing Mixlayer: a platform for high-quality open source AI inference, built from the ground up to bring better infrastructure, protocols, and APIs to open models.

Z
Zack Angelo

New posts, straight to your inbox.

Research and engineering notes from the Mixlayer team. No spam, unsubscribe anytime.

Subscribe →