AI Engineer · LLM infra × backend

Vaibhav
Agarwal

I build reliable AI systems — where LLM infrastructure meets real backend engineering.

Writing about RAG, agents, evaluation, and the infrastructure that makes AI reliable in production. Notes from building, not tutorials from docs.

01 / Work

Selected projects

All projects →

LLM Interpretability

Black-box interpretability for language models: benchmarking, an open-source library, and a novel probing method.

Case study in progress

MCP Servers

Model Context Protocol servers that expose structured data layers to LLM tools cleanly and reliably.

Case study in progress

AI News Desk

A personal aggregation pipeline that tracks the AI landscape and auto-drafts posts for distribution.

Case study in progress

02 / Writing

Recent writing

All writing →

Current series · Production RAG

  1. Production RAG: the complete guidePillar
  2. Chunking strategies for RAG: fixed-size vs recursive vs semantic (benchmarked)Soon
  3. pgvector vs a dedicated vector database: when do you actually need one?Soon
  4. How to evaluate RAG retrieval quality (the metrics that matter)Soon
  5. Does reranking actually improve RAG answers? (I measured it)Soon
03 / Connect

Say hello

Working on retrieval, agents, or putting LLMs into production? I'm always up for comparing notes. Email is best: vaibhavkumaragarwal2020@gmail.com