OpenLake Blogs
Technical notes on GPU inference, storage systems, and high throughput AI infrastructure.
Inference
KV Cache
RDMA
Storage
Search
⌘K

Deep Dive · Research
Compressing the Incompressible
BF16 was built to resist compression. Its exponent field had other ideas.
7 min read

Deep Dive · Inference
Breaking the KV Wall
A 256K conversation once filled an entire H100. Now it fills nothing.
8 min read

Deep Dive · Announcement
Introducing OpenLake
OpenLake: Fast, Durable Storage for LLM Inference and Training
5 min read