About

How 13 years of engineering led to LLM products in production — and why the AI work holds up precisely because of that foundation.

I build LLM products that reach production. Since 2024 that has mostly meant BookahTranslate — an AI document-translation SaaS I built and operate solo, with paying subscribers. The pipeline runs on GPT-4 and Claude and handles PDF, EPUB and DOCX up to 300+ pages while preserving layout: parsing, chunking, context assembly, translation, reassembly.

The interesting part of that work is never the prompt. It is automatic fallback between providers and models when one times out or rate-limits; response caching and model routing to keep inference costs down; and an evaluation set that gates every prompt or model change — 80+ tests, accuracy, relevance, hallucination rate, latency and cost, with automated regression.

Alongside it I do AI-implementation consulting for fintech, legal-tech and B2B SaaS teams: RAG over corporate knowledge bases — sentence-aligned chunking at ~3,500 characters, hybrid dense and lexical retrieval, calibrated confidence floors that keep unsupported answers under 10% on the eval set — agent architectures over custom MCP servers with strict tool-call schemas and multi-step tool loops, and rolling out Claude Code, Codex and Cursor across client engineering teams with quality gates before release.

A second product, Digital Psychologist, is where I compared fine-tuning against retrieval head to head: LangChain and FAISS for the retrieval layer with local sentence-transformer embeddings, and three interchangeable modes — retrieval-only, fine-tuned model, and hybrid — switchable by configuration, so the trade-off could be measured on one product rather than argued.

Two open-source tools came out of this work. pr-witness runs the base branch's tests against a pull request's code and checks the PR's own claims against the real run — the combination CI never runs, and the way an AI agent's "all tests pass" gets verified rather than trusted. ftgate checks whether a fine-tuned small model is served the bytes it was trained on; its first week measured what a known Ollama bug costs — the tool schema reaching the model as a Go struct dump — and found Qwen's official GGUFs embedding a pre-fix chat template, both reported upstream with measurements. Alongside them, three standalone demos — an eval harness with a CI release gate, an MCP server, and a RAG pipeline that refuses rather than guesses — all offline, with tests and green CI.

Underneath all of this is 13 years of engineering. I started in QA — manual, then leading a team and writing Python + Selenium automation — which is why I think about edge cases first and write tests as I go. Then backend at Smart Trade (Germany), 3.5 years at Bequant (an institutional crypto exchange: trading terminal at 10K events/min, a GraphQL layer that cut backend load 45%), and Coperniq, a US solar SaaS, where I led 4 engineers, drove the monolith → microservices migration and cut API p50 from 1.2s to 180ms.

That foundation is the reason the AI work holds up. Anyone can call an LLM API; the difficulty is making it survive contact with real users, real failure modes and a real bill.

Education

  • 2024 — Russian State Social University, Jurisprudence / Social services
  • 2020 — RUDN University, Jurisprudence / Corporate Law
  • 2015 — RUDN University, Mathematics and Computer Science

Languages

  • Russian — native
  • English — C1, daily work language
  • French — A2