Skip to content
Ashish's Engineering Lab

Engineering Notes

High-signal thoughts on infrastructure, artificial intelligence, and the intricacies of building resilient systems at scale.

  • 3 min readAI Engineering

    The True Cost of LLM Latency

    Time-to-first-token and total generation time are different products. Streaming, deadline propagation, and why your timeout budget is probably wrong.

  • 3 min readInfrastructure

    Optimizing Vector Search at Scale

    Reducing latency in large vector indexes: quantization, graph pruning, and why the page cache is usually a better tiering strategy than the one you were about to write.