Writing · RSS

Writing

Notes from building, not tutorials from docs. RAG, agents, evaluation, and the infrastructure that makes AI reliable in production.

First posts are on the way. Here's what's being written right now.

Up next

Now writing

Current series · Production RAG

  1. Production RAG: the complete guidePillar
  2. Chunking strategies for RAG: fixed-size vs recursive vs semantic (benchmarked)Soon
  3. pgvector vs a dedicated vector database: when do you actually need one?Soon
  4. How to evaluate RAG retrieval quality (the metrics that matter)Soon
  5. Does reranking actually improve RAG answers? (I measured it)Soon
  6. Choosing an embedding model: OpenAI vs open-source, benchmarkedSoon
  7. Hybrid search in production: combining BM25 and vector searchSoon
  8. Why your RAG returns irrelevant chunks (and how to debug it)Soon
  9. Cutting RAG costs: caching embeddings and LLM callsSoon
  10. Metadata filtering in RAG: the cheap accuracy win most people skipSoon
  11. Late-interaction retrieval (ColBERT): is it worth the complexity?Soon