LLM routing: when GPT-4 is overkill
A simple per-document model picker that cut my LLM bill by 40% without hurting quality.
When you ship an LLM-powered product, your first instinct is to default to the best model — GPT-4o for everything. It's a fine default until your monthly bill hits four digits and you realize half your traffic is translating recipe PDFs that GPT-4o is laughably overqualified for.
On BookahTranslate I built a small routing layer that picks a model based on three signals: document kind, user tier, and document size. It looks unsophisticated. It works.
def pick_model(doc: Document, tier: UserTier) -> ModelChoice:
if tier == UserTier.ECONOMY and doc.kind in {
DocKind.TXT, DocKind.PLAIN_PDF
}:
return ModelChoice.GOOGLE_TRANSLATE
if doc.kind == DocKind.SCANNED_PDF:
return ModelChoice.GPT_4O # vision + OCR fusion
if doc.kind == DocKind.LEGAL or doc.has_dense_terminology:
return ModelChoice.CLAUDE_SONNET
if doc.pages > 200 and tier != UserTier.PREMIUM:
return ModelChoice.CLAUDE_HAIKU # cheap for bulk
return ModelChoice.GPT_4O
Why each branch
- Plain PDFs and TXT on the Economy tier: Google Translate is dramatically cheaper, and for a recipe or news article the quality difference is imperceptible.
- Scanned PDFs: need vision-capable model. GPT-4o currently wins on OCR + translation in one pass.
- Legal / dense technical: Claude Sonnet handles terminology consistency over long contexts better than GPT-4o in my A/B sampling.
- Long documents on lower tiers: Claude Haiku is ~10x cheaper than GPT-4o and the quality drop on simple prose is hard to notice unless you're reading line-by-line.
What I'd warn against
Don't build a 'smart router' that asks an LLM to pick the LLM. I tried that. It added latency, costs, and one more failure mode. A handful of explicit branches based on cheap signals (file type, page count, user tier) is the right tool. If you can't articulate the decision in 8 lines of Python, you don't understand your traffic well enough to route it yet.
Observability matters more than the picker
Every translation logs the picked model, token counts, latency, and a quality flag (filled in async by a small Claude pass that re-reads the output). I review the worst-quality 1% weekly. Most of the time the fix is 'route this document profile to a stronger model,' not 'improve the prompt.'