BookahTranslate / AthenaDev
AI-powered document translation SaaS. PDF, EPUB, DOCX up to 300+ pages with layout preservation, powered by GPT-4 and Claude.
Outcomes
Context
Translation of long-form documents (50–300 pages) with formatting preserved is a real pain: Google Translate flattens PDFs into plain text, DeepL has no PDF support for end users, and the few SaaS options that do support PDF are either slow, expensive, or lose the layout entirely.
I wanted a service where you drop a PDF, EPUB, or DOCX and get back a readable translation with the original layout intact — figures, tables, code blocks, math, footnotes. And I wanted to ship it as a real SaaS with subscription billing, not a toy demo.
What I built
- Three-component architecture: Web App (Flask + SQLAlchemy), Translation API (Flask, isolated for long-running jobs), and a Telegram Bot — all sharing nothing except a Redis queue.
- LLM routing layer: GPT-4 / GPT-4o for high-quality literary text, Anthropic Claude for technical / legal content, Google Translate for cheap bulk. Choice driven by document type, user tier, and document size.
- Three-tier subscription system with page-based billing, bonus pages, one-time payments, and auto-renewal via YooKassa (Russian payment provider, Stripe-equivalent). Webhook verification by IP whitelist and signature.
- OAuth 2.0 via Google, VK, and Telegram (deep-link). Plus email/password with rate-limited registration and password validation.
- PDF processing pipeline: BabelDOC for layout-preserving translation, pdf2zh as fallback, ReportLab for output, OCR via Tesseract when documents are scanned.
- Async job processing through Celery + Redis with progress streaming back to the browser — handles 300-page documents without hitting Gunicorn timeouts.
- Provider fallback at the call site: timeouts, rate limits, provider outages and malformed responses each degrade to a secondary model rather than failing the user's job. Every fallback is logged with the error class that triggered it, so the failure pattern is visible rather than inferred.
- An evaluation set gates every prompt and model change: accuracy, relevance, hallucination rate, latency and cost per page, with per-case regression against a stored baseline. A change that improves the average while breaking three specific document types does not ship.
- Versioned prompts and request/response logging with pipeline tracing — when a translation comes out wrong, the exact prompt, model and intermediate chunks that produced it are recoverable.
- 80+ pytest tests (auth, billing, translation, user isolation) plus Playwright e2e covering the full register-to-translate-to-pay journey.
Technical decisions
Database-per-context separation: three independent PostgreSQL databases (translate_bot, translate_bot_staging, translate_web). Failure in one doesn't propagate. Same pattern for Redis instances.
Race condition protection: page deductions go through SELECT FOR UPDATE inside an atomic transaction; subscription activation is enforced by a partial unique index ('only one active subscription per user').
Security defaults: JWT with database-backed blacklist (immediate revocation on logout), Fernet encryption for user-supplied API keys, Flask-Limiter on registration (5/hour) and login (10/min) backed by Redis.
Cost engineering: smart model routing cut average per-page LLM cost by ~40% vs. always using GPT-4. Bulk Google Translate path saves another tier of cost for users on Economy plan. Response caching removes repeat spend on identical chunks, which matters for documents with boilerplate sections.
A malformed 200 is still a failure. Provider responses are validated for shape before being accepted, because a truncated or empty translation that arrives with HTTP 200 is worse than an error — it reaches the customer looking finished.
Refuse rather than half-deliver: when every provider in the chain fails, the job fails loudly and the page is refunded. A partially translated document handed over as complete destroys trust far more cheaply than an honest error does.
From the code
def use_pages_atomic(subscription_id: int, pages_count: int) -> tuple[bool, int]:
"""Deduct pages with row-level lock to prevent concurrent over-spend."""
sub = (
db.session.query(UserSubscription)
.with_for_update() # SELECT ... FOR UPDATE
.filter(UserSubscription.id == subscription_id)
.first()
)
if not sub or sub.status not in ("active", "grace_period"):
return False, 0
if sub.pages_remaining < pages_count:
db.session.rollback()
return False, sub.pages_remaining
sub.pages_remaining -= pages_count
db.session.add(PageTransaction(
subscription_id=sub.id,
pages_delta=-pages_count,
balance_after=sub.pages_remaining,
))
db.session.commit() # releases lock
return True, sub.pages_remaining
def pick_model(doc: Document, tier: UserTier) -> ModelChoice:
if tier == UserTier.ECONOMY and doc.kind in {DocKind.TXT, DocKind.PLAIN_PDF}:
return ModelChoice.GOOGLE_TRANSLATE
if doc.kind == DocKind.SCANNED_PDF:
return ModelChoice.GPT_4O # needs vision + OCR fusion
if doc.kind == DocKind.LEGAL or doc.has_dense_terminology:
return ModelChoice.CLAUDE_SONNET
if doc.pages > 200 and tier != UserTier.PREMIUM:
return ModelChoice.CLAUDE_HAIKU # cheap for bulk
return ModelChoice.GPT_4O
# Each failure class is retried differently: a timeout deserves another
# attempt, a 400 does not. The chain degrades to a cheaper model rather
# than returning an error to someone who already paid for the page.
RETRYABLE = (ProviderTimeout, RateLimited, ProviderUnavailable)
def translate_chunk(chunk: str, chain: list[ModelChoice]) -> Translation:
failures: list[str] = []
for model in chain:
for attempt in range(MAX_ATTEMPTS):
try:
result = call_provider(model, chunk, timeout=timeout_for(model))
if not result.is_well_formed():
# A malformed response is a failure even with HTTP 200.
raise MalformedResponse(model)
if failures:
log.warning("recovered after %s", ",".join(failures))
metrics.incr("translate.fallback", tags={"to": model.name})
return result
except RETRYABLE as exc:
failures.append(f"{model.name}:{type(exc).__name__}")
sleep(backoff(attempt)) # exponential, jittered
except MalformedResponse as exc:
failures.append(f"{model.name}:malformed")
break # retrying won't fix the prompt
except ProviderRefused:
failures.append(f"{model.name}:refused")
break # next model, not next attempt
# Every provider is down: fail loudly and refund the page, never
# hand the user a half-translated document as if it succeeded.
raise AllProvidersFailed(failures)
def gate(report: RunReport, baseline: Baseline) -> GateResult:
"""Block a deploy when quality drops. Published as llm-eval-harness."""
violations = []
if report.hallucination_rate > MAX_HALLUCINATION_RATE:
violations.append(f"hallucination {report.hallucination_rate:.1%}")
if report.mean_accuracy < MIN_ACCURACY:
violations.append(f"accuracy {report.mean_accuracy:.3f}")
if report.cost_per_page > baseline.cost_per_page * COST_CEILING:
violations.append(f"cost/page {report.cost_per_page:.4f}")
# The averages above can all pass while specific documents break.
# Per-case comparison is what actually catches a regression.
regressions = [
case.id for case in report.cases
if baseline.passed(case.id) and not case.passed
]
return GateResult(
passed=not violations and not regressions,
violations=violations,
regressions=regressions,
)
Result
In production on bookahtranslate.tech (web) and via @BookahTranslateBot (Telegram). Built and operated solo, paying users, organic growth.
Acts as my live laboratory for AI-native product engineering: every PR ships to a real service with real users, real billing, real failure modes — not a sandbox.
The reliability patterns here are published as standalone, readable repositories: the release gate as llm-eval-harness (47 tests), grounded retrieval that refuses rather than guesses as rag-grounded (40 tests), and tool calling as mcp-toolserver (50 tests). Running a paid service is what taught me which of these actually matter.