LLM Interpretability
Black-box interpretability for language models: benchmarking, an open-source library, and a novel probing method.
Case study in progress
I build reliable AI systems — where LLM infrastructure meets real backend engineering.
Writing about RAG, agents, evaluation, and the infrastructure that makes AI reliable in production. Notes from building, not tutorials from docs.
Black-box interpretability for language models: benchmarking, an open-source library, and a novel probing method.
Case study in progress
Model Context Protocol servers that expose structured data layers to LLM tools cleanly and reliably.
Case study in progress
A personal aggregation pipeline that tracks the AI landscape and auto-drafts posts for distribution.
Case study in progress
Current series · Production RAG
Working on retrieval, agents, or putting LLMs into production? I'm always up for comparing notes. Email is best: vaibhavkumaragarwal2020@gmail.com