athenadev.tech

AthenaDev

AI integration consultancy. Embedding Claude Code, MCP servers, and LLM workflows into mid-size product teams.

AthenaDev is my consultancy for embedding AI into the day-to-day of product engineering. Not 'add a chatbot' projects. Real workflow integration: Claude Code rollouts to dev teams, MCP servers connecting LLMs to internal tools, RAG systems for institutional knowledge, AI-driven code review pipelines.

Fixed-scope projectsRetainer (4–20h/week)Advisory / fractional CTO for AI

What we do

Claude Code rollout for product teams

Take a team from 'we hear about Claude Code' to 'half our PRs start as agent-driven drafts'. Setup, custom slash commands, project-specific hooks, MCP integrations, and the boring-but-critical training piece that determines adoption.

  • →Per-project settings.json with sane permission defaults
  • →Custom skills, hooks, and slash commands for your stack
  • →MCP server connecting Claude to your Linear/Jira/Slack/internal APIs
  • →Adoption metrics dashboard (active users, PR-with-AI-assist ratio)

Custom MCP servers

Model Context Protocol is the right primitive for connecting LLMs to your stack. I build production-grade MCP servers in TypeScript or Python: auth, rate-limiting, observability, schema validation. Self-hosted or deployed as a managed service.

  • →Bidirectional tool integration (read + write)
  • →Type-safe tool schemas (Zod / Pydantic) with runtime validation
  • →Per-user auth scoping (no 'agent acts as admin' incidents)
  • →Audit log + dashboards for tool-call traffic

RAG over internal knowledge

Production RAG, not a demo. Document ingestion pipeline, embedding strategy (chunking, hierarchy, hybrid lexical + dense), retrieval evaluation, citations in every answer, output guardrails. Built on whatever vector store fits — Postgres + pgvector, Qdrant, Pinecone — chosen on cost/latency, not hype.

  • →Document ingestion with versioning + incremental updates
  • →Hybrid retrieval (BM25 + dense embeddings) + reranking
  • →Citation-grounded answers with link-back to source
  • →Evaluation harness (precision/recall on a labeled set)

LLM observability & cost engineering

Most teams ship LLM features blind. I install the observability layer (Langfuse, OpenTelemetry, custom dashboards), the cost-tracking layer (per-feature, per-user, per-model), and the model-routing layer that lets you pick the cheapest model that still passes your quality bar.

  • →Per-request traces with prompt, completion, tokens, latency
  • →Cost attribution by team / feature / customer
  • →Smart model router (rules or learned)
  • →Alerts on quality regressions and cost spikes

Selected engagements

Clients are anonymized — standard consulting practice. Profile and context can be verified under NDA.

Fintech · payment processing · 8 weeks engagement + retainer · 1 dev team (12 engineers) → org-wide

Mid-size fintech (Series B, ~80 engineers)

Problem

Engineering leadership saw Claude Code being used ad-hoc by individual devs but had no organizational adoption, no shared config, and no visibility. Some teams built clever workflows; others were stuck running Claude in default mode with no permissions guardrails.

Concrete asks: standardize on a per-repo configuration, build org-wide skills/slash-commands for their stack (Go + React + Postgres), wire Claude into their internal tools (Linear, GitHub, their internal feature-flag service), and produce metrics that engineering leadership could review monthly.

Approach

  • Started with one pilot team of 12 engineers. Wrote a baseline .claude/settings.json with allowlist-based Bash permissions, denied destructive commands by default (rm -rf, force-push, db drops), and pre-approved their internal CLI tooling.
  • Built 6 custom skills: 'review-pr' (runs their lint + test + custom static-analysis), 'add-feature-flag' (talks to their flag service via MCP), 'rollback-migration', 'check-staging', 'ship-it' (PR + assign reviewers based on CODEOWNERS), 'why-flaky' (their flaky-test investigation playbook).
  • Built an internal MCP server in TypeScript exposing their Linear, GitHub, and feature-flag APIs. Per-user OAuth — agent acts with the engineer's permissions, not a service account.
  • Hook-based PR description generation: on git commit, a post-commit hook calls Claude with the diff to draft a PR description in their format. Saves ~5 min per PR.
  • Monthly adoption review: dashboard showing active users, sessions per dev per week, PR-with-AI-assist ratio, and tool-call breakdown by skill.

Decisions

Per-user auth on every MCP tool, not a shared service account. The cost is more setup; the benefit is no 'agent silently does production things' incidents. Audit logs map every tool call to a real engineer.

Made the rollout opt-in for the first 4 weeks. Engineers who ignored it kept ignoring it. Engineers who tried it became internal advocates. Forcing adoption from above poisons the well — let the network effect do the work.

Built the metrics dashboard before the rollout, not after. Leadership pre-committed to specific success metrics. Avoided the 'did this even help?' debate later.

From the code

json
settings.json with team-safe defaults
{
  "permissions": {
    "allow": [
      "Bash(npm:*)",
      "Bash(go test:*)",
      "Bash(git status)",
      "Bash(git diff:*)",
      "Bash(./scripts/dev-cli:*)"
    ],
    "deny": [
      "Bash(rm -rf:*)",
      "Bash(git push --force:*)",
      "Bash(git push -f:*)",
      "Bash(psql:*drop*)",
      "Bash(make deploy:*)"
    ],
    "ask": [
      "Bash(git push:*)",
      "Bash(./scripts/migrate:*)"
    ]
  },
  "hooks": {
    "PostToolUse": [
      {
        "matcher": "Edit|Write",
        "hooks": [{
          "type": "command",
          "command": "./scripts/run-precommit.sh"
        }]
      }
    ]
  }
}
typescript
MCP tool: scoped Linear search
server.tool(
  "linear_search_issues",
  "Search Linear issues with the current user's permissions.",
  {
    query: z.string().describe("Linear search query syntax"),
    limit: z.number().int().min(1).max(50).default(20),
  },
  async ({ query, limit }, { session }) => {
    // session.linearToken is the user's OAuth token, not a service account.
    const client = new LinearClient({ accessToken: session.linearToken });
    const issues = await client.issues({ filter: parseQuery(query), first: limit });

    audit.log({
      tool: "linear_search_issues",
      userId: session.userId,
      query,
      resultCount: issues.nodes.length,
    });

    return issues.nodes.map((i) => ({
      id: i.identifier,
      title: i.title,
      state: i.state?.name,
      url: i.url,
    }));
  }
);

Outcome

Pilot team adoption went from ~3 active users to 11 of 12 within the engagement. Average time-to-first-PR for new hires dropped from 4 days to 1.5. The 'review-pr' skill caught a class of bugs (missing migration in PR) that their CI hadn't been configured to flag.

After the engagement they pulled the same playbook into 4 more teams. I stay on a retainer for ongoing skill-building and MCP server maintenance.

3 → 11/12
Active users in pilot team
−63%
Time-to-first-PR for new hires
6
Custom skills shipped
4
Teams adopted post-engagement
B2B SaaS · workflow automation · 5 weeks · Solo + their engineering lead

B2B SaaS company (~30 engineers, Series A)

Problem

Engineering lead asked for an AI-driven first-pass code reviewer on GitHub PRs. The trial-and-error stage was over — they had piloted three off-the-shelf tools (Copilot Review, CodeRabbit, Greptile) and weren't satisfied with the signal-to-noise ratio. Too many cosmetic nits, not enough useful catches.

What they actually wanted: a reviewer that knew their codebase conventions (lint rules, internal patterns, deprecated utilities), flagged risky changes (auth, billing, migrations), and posted exactly one comment per PR — a summary, not 14 separate inline nits.

Approach

  • GitHub Action triggered on PR open + synchronize. Pulls the diff, the changed files in full, and (for risky paths) the touched files' callers via a quick grep pass.
  • Routed Claude Sonnet — Haiku was too imprecise on their TS codebase; Opus was overkill for routine PRs. Sonnet on a 4K-token max response kept costs at ~$0.07/PR average.
  • Wrote a custom system prompt with their actual conventions: 'use ourBigDecimal for money, never Number'; 'never call /v1/payments directly, go through PaymentService'; 'migrations need a reversible down() within 24h'. About 1200 tokens of repo-specific rules.
  • Risk-tier classification: PR touching auth/, billing/, or migrations/ → 'high', triggers a stricter prompt and tags @security-team. PR touching only tests or docs → 'low', shorter response template.
  • Output: one structured comment per PR with 'Summary', 'Risk', 'Issues found' (with severity), 'Suggestions'. No inline noise. The team's own engineers handle the line-by-line review — Claude does the framing.
  • PII/secret filter on input: any line matching common secret patterns (API keys, tokens) is stripped before sending to the API. Their security team signed off on this before we shipped.

Decisions

One comment per PR, not inline. The team had review fatigue from previous tools; consolidation respected their attention. Engineers can ignore one comment they disagree with; they can't ignore 14.

Risk-tier routing instead of running every PR through the strictest prompt. Cost dropped ~3x and false-positive rate dropped because the high-tier prompt was reserved for changes that genuinely warranted scrutiny.

Shipped with a feedback button (thumbs up/down + optional reason) on every Claude comment. Used the negative feedback weekly to update the system prompt — most fixes were 'this rule isn't actually a rule, our codebase does X here'.

From the code

yaml
GitHub Action entry point
name: AI code review

on:
  pull_request:
    types: [opened, synchronize, reopened]

jobs:
  review:
    runs-on: ubuntu-latest
    permissions:
      contents: read
      pull-requests: write
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0

      - uses: actions/setup-node@v4
        with: { node-version: "20" }

      - name: Classify risk tier
        id: tier
        run: node scripts/risk-tier.mjs >> "$GITHUB_OUTPUT"

      - name: Strip secrets from diff
        run: node scripts/scrub-secrets.mjs > /tmp/diff.txt

      - name: Run AI reviewer
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
          RISK_TIER: ${{ steps.tier.outputs.tier }}
        run: node scripts/review.mjs /tmp/diff.txt
typescript
Risk-tier classifier
const HIGH_RISK_PATHS = [
  /^src\/auth\//,
  /^src\/billing\//,
  /^src\/payments\//,
  /^migrations\//,
  /\.env(\..+)?$/,
];

const LOW_RISK_PATHS = [
  /\.test\.tsx?$/,
  /^docs\//,
  /^\.github\/workflows\//,
];

export function classifyTier(changedFiles: string[]): "high" | "normal" | "low" {
  if (changedFiles.some((f) => HIGH_RISK_PATHS.some((re) => re.test(f)))) {
    return "high";
  }
  if (changedFiles.every((f) => LOW_RISK_PATHS.some((re) => re.test(f)))) {
    return "low";
  }
  return "normal";
}

Outcome

After 4 weeks in production: 'useful catch' rate (engineer thumbs-up) was 71%. Time-to-first-human-review on PRs dropped from ~4h to ~45min (engineers triage Claude's summary first, then commit to a deep review).

Engineering lead's own statement at our wrap-up: 'this is the first AI tool we've used that actually feels like a junior reviewer, not a linter pretending to be one.'

71%
Useful-catch rate (engineer 👍)
−81%
Time-to-first-human-review
$0.07
Avg LLM cost per PR
1
Comment per PR (no inline noise)

How we work

1 · Diagnostic call (free, ~45 min)

Where you actually are with AI today, what's painful, what would 'good' look like. No deck, just questions.

2 · Scoped proposal

1-2 page proposal with concrete deliverables, timeline, and success metrics. Fixed-scope or retainer.

3 · Pilot (typically 2–4 weeks)

Ship the smallest end-to-end thing that proves the approach. Pilot with one team, not the whole org.

4 · Rollout + adoption work

Configuration, custom skills, internal documentation, training sessions. Adoption is 60% of the value — I do this part too.

5 · Retainer (optional)

Ongoing maintenance, expansion to new teams, monthly metric reviews. Most engagements continue here.

Get in touch

If you'd like to talk through whether Claude Code or RAG fits your team — reach out. Diagnostic call is free, no obligation.

bogachev.vitaliy91test@gmail.com · @tenkuioo · all contacts