Blog

Technical writing for inference operators.

Best practices for AI inference, GPU optimization for LLM inference, and field guidance for teams running GenAI inference in production.

Aug 14, 2026 4 min read

An FDE’s playbook for diagnosing TTFT delays

Most discussions about AI inference performance focus on model speed. In practice, however, users experience performance through a different metric: Time to First Token (TTFT). TTFT is the elapsed…

Mar 16, 2026 5 min read

Inference is Underrated

Scan the AI headlines and you'd be forgiven for thinking the only thing that matters is the next training run. $5 billion clusters. Millions of GPU-hours. A new…