Tomcy Thomas ~[t-om-c]
0
0
about
blogs
play
cv
~blogs
// Notes on building, designing and breaking things.
Cutting Input Tokens by 82%: Optimizing an Agentic RAG Pipeline
Our agentic RAG pipeline burned ~12,000 input tokens per turn. We cut it to ~2,200 on average — an 82% reduction — while keeping citations and confidence scoring.
July 27, 2026.
8 min read
AI
RAG
LLM
Optimization
© 2026, sktomsi ✦ tomcy thomas