← all work
2024 — Now

BookahTranslate / AthenaDev

CTO

AI-powered document translation SaaS. PDF, EPUB, DOCX up to 300+ pages with layout preservation, powered by GPT-4 and Claude.

Outcomes

3
Components shipped solo
300+
Pages per document supported
80+
Tests (>70% coverage)
3
Auth providers (Google, VK, Telegram)

Context

Translation of long-form documents (50–300 pages) with formatting preserved is a real pain: Google Translate flattens PDFs into plain text, DeepL has no PDF support for end users, and the few SaaS options that do support PDF are either slow, expensive, or lose the layout entirely.

I wanted a service where you drop a PDF, EPUB, or DOCX and get back a readable translation with the original layout intact — figures, tables, code blocks, math, footnotes. And I wanted to ship it as a real SaaS with subscription billing, not a toy demo.

What I built

  • Three-component architecture: Web App (Flask + SQLAlchemy), Translation API (Flask, isolated for long-running jobs), and a Telegram Bot — all sharing nothing except a Redis queue.
  • LLM routing layer: GPT-4 / GPT-4o for high-quality literary text, Anthropic Claude for technical / legal content, Google Translate for cheap bulk. Choice driven by document type, user tier, and document size.
  • Three-tier subscription system with page-based billing, bonus pages, one-time payments, and auto-renewal via YooKassa (Russian payment provider, Stripe-equivalent). Webhook verification by IP whitelist and signature.
  • OAuth 2.0 via Google, VK, and Telegram (deep-link). Plus email/password with rate-limited registration and password validation.
  • PDF processing pipeline: BabelDOC for layout-preserving translation, pdf2zh as fallback, ReportLab for output, OCR via Tesseract when documents are scanned.
  • Async job processing through Celery + Redis with progress streaming back to the browser — handles 300-page documents without hitting Gunicorn timeouts.
  • Provider fallback at the call site: timeouts, rate limits, provider outages and malformed responses each degrade to a secondary model rather than failing the user's job. Every fallback is logged with the error class that triggered it, so the failure pattern is visible rather than inferred.
  • An evaluation set gates every prompt and model change: accuracy, relevance, hallucination rate, latency and cost per page, with per-case regression against a stored baseline. A change that improves the average while breaking three specific document types does not ship.
  • Versioned prompts and request/response logging with pipeline tracing — when a translation comes out wrong, the exact prompt, model and intermediate chunks that produced it are recoverable.
  • 80+ pytest tests (auth, billing, translation, user isolation) plus Playwright e2e covering the full register-to-translate-to-pay journey.

Technical decisions

Database-per-context separation: three independent PostgreSQL databases (translate_bot, translate_bot_staging, translate_web). Failure in one doesn't propagate. Same pattern for Redis instances.

Race condition protection: page deductions go through SELECT FOR UPDATE inside an atomic transaction; subscription activation is enforced by a partial unique index ('only one active subscription per user').

Security defaults: JWT with database-backed blacklist (immediate revocation on logout), Fernet encryption for user-supplied API keys, Flask-Limiter on registration (5/hour) and login (10/min) backed by Redis.

Cost engineering: smart model routing cut average per-page LLM cost by ~40% vs. always using GPT-4. Bulk Google Translate path saves another tier of cost for users on Economy plan. Response caching removes repeat spend on identical chunks, which matters for documents with boilerplate sections.

A malformed 200 is still a failure. Provider responses are validated for shape before being accepted, because a truncated or empty translation that arrives with HTTP 200 is worse than an error — it reaches the customer looking finished.

Refuse rather than half-deliver: when every provider in the chain fails, the job fails loudly and the page is refunded. A partially translated document handed over as complete destroys trust far more cheaply than an honest error does.

From the code

python
Atomic page deduction (race-safe)
def use_pages_atomic(subscription_id: int, pages_count: int) -> tuple[bool, int]:
    """Deduct pages with row-level lock to prevent concurrent over-spend."""
    sub = (
        db.session.query(UserSubscription)
        .with_for_update()  # SELECT ... FOR UPDATE
        .filter(UserSubscription.id == subscription_id)
        .first()
    )
    if not sub or sub.status not in ("active", "grace_period"):
        return False, 0
    if sub.pages_remaining < pages_count:
        db.session.rollback()
        return False, sub.pages_remaining

    sub.pages_remaining -= pages_count
    db.session.add(PageTransaction(
        subscription_id=sub.id,
        pages_delta=-pages_count,
        balance_after=sub.pages_remaining,
    ))
    db.session.commit()  # releases lock
    return True, sub.pages_remaining
python
LLM router by document profile
def pick_model(doc: Document, tier: UserTier) -> ModelChoice:
    if tier == UserTier.ECONOMY and doc.kind in {DocKind.TXT, DocKind.PLAIN_PDF}:
        return ModelChoice.GOOGLE_TRANSLATE

    if doc.kind == DocKind.SCANNED_PDF:
        return ModelChoice.GPT_4O   # needs vision + OCR fusion

    if doc.kind == DocKind.LEGAL or doc.has_dense_terminology:
        return ModelChoice.CLAUDE_SONNET

    if doc.pages > 200 and tier != UserTier.PREMIUM:
        return ModelChoice.CLAUDE_HAIKU   # cheap for bulk

    return ModelChoice.GPT_4O
python
Provider fallback — a user's job should not die with one vendor
# Each failure class is retried differently: a timeout deserves another
# attempt, a 400 does not. The chain degrades to a cheaper model rather
# than returning an error to someone who already paid for the page.
RETRYABLE = (ProviderTimeout, RateLimited, ProviderUnavailable)

def translate_chunk(chunk: str, chain: list[ModelChoice]) -> Translation:
    failures: list[str] = []

    for model in chain:
        for attempt in range(MAX_ATTEMPTS):
            try:
                result = call_provider(model, chunk, timeout=timeout_for(model))
                if not result.is_well_formed():
                    # A malformed response is a failure even with HTTP 200.
                    raise MalformedResponse(model)
                if failures:
                    log.warning("recovered after %s", ",".join(failures))
                    metrics.incr("translate.fallback", tags={"to": model.name})
                return result

            except RETRYABLE as exc:
                failures.append(f"{model.name}:{type(exc).__name__}")
                sleep(backoff(attempt))          # exponential, jittered
            except MalformedResponse as exc:
                failures.append(f"{model.name}:malformed")
                break                            # retrying won't fix the prompt
            except ProviderRefused:
                failures.append(f"{model.name}:refused")
                break                            # next model, not next attempt

    # Every provider is down: fail loudly and refund the page, never
    # hand the user a half-translated document as if it succeeded.
    raise AllProvidersFailed(failures)
python
Release gate — a prompt change must prove it did no harm
def gate(report: RunReport, baseline: Baseline) -> GateResult:
    """Block a deploy when quality drops. Published as llm-eval-harness."""
    violations = []

    if report.hallucination_rate > MAX_HALLUCINATION_RATE:
        violations.append(f"hallucination {report.hallucination_rate:.1%}")
    if report.mean_accuracy < MIN_ACCURACY:
        violations.append(f"accuracy {report.mean_accuracy:.3f}")
    if report.cost_per_page > baseline.cost_per_page * COST_CEILING:
        violations.append(f"cost/page {report.cost_per_page:.4f}")

    # The averages above can all pass while specific documents break.
    # Per-case comparison is what actually catches a regression.
    regressions = [
        case.id for case in report.cases
        if baseline.passed(case.id) and not case.passed
    ]

    return GateResult(
        passed=not violations and not regressions,
        violations=violations,
        regressions=regressions,
    )

Result

In production on bookahtranslate.tech (web) and via @BookahTranslateBot (Telegram). Built and operated solo, paying users, organic growth.

Acts as my live laboratory for AI-native product engineering: every PR ships to a real service with real users, real billing, real failure modes — not a sandbox.

The reliability patterns here are published as standalone, readable repositories: the release gate as llm-eval-harness (47 tests), grounded retrieval that refuses rather than guesses as rag-grounded (40 tests), and tool calling as mcp-toolserver (50 tests). Running a paid service is what taught me which of these actually matter.

Stack

PythonFlaskSQLAlchemyCeleryPostgreSQLRedisOpenAI GPT-4Anthropic ClaudeGoogle TranslateYooKassaTelegram Bot APINginxLinuxpytestPlaywright