skip to content

Pradeep - notes on AI/ML, system design, and building software

Featured

Understanding LLM Token Caching

The missing details you need to run production prompt caches: provider minimums, write costs, TTL/eviction, and how the KV cache actually maps to your bill.

12 min read

AdTurbine

AdTurbine is a platform for generating UGC (User Generated Content) ad videos using AI. The system follows a pipeline architecture.

fastapinext.jsdockerllm

Ragforge

Enterprise patterns for building reliable, scalable RAG systems with clear service boundaries and governance layers.

fastapipgvectorprometheusnext.jsdocker