AthenaDev
AthenaDev is my consultancy for embedding AI into the day-to-day of product engineering. Not 'add a chatbot' projects. Real workflow integration: Claude Code rollouts to dev teams, MCP servers connecting LLMs to internal tools, RAG systems for institutional knowledge, AI-driven code review pipelines.
What we do
Claude Code rollout for product teams
Take a team from 'we hear about Claude Code' to 'half our PRs start as agent-driven drafts'. Setup, custom slash commands, project-specific hooks, MCP integrations, and the boring-but-critical training piece that determines adoption.
- →Per-project settings.json with sane permission defaults
- →Custom skills, hooks, and slash commands for your stack
- →MCP server connecting Claude to your Linear/Jira/Slack/internal APIs
- →Adoption metrics dashboard (active users, PR-with-AI-assist ratio)
Custom MCP servers
Model Context Protocol is the right primitive for connecting LLMs to your stack. I build production-grade MCP servers in TypeScript or Python: auth, rate-limiting, observability, schema validation. Self-hosted or deployed as a managed service.
- →Bidirectional tool integration (read + write)
- →Type-safe tool schemas (Zod / Pydantic) with runtime validation
- →Per-user auth scoping (no 'agent acts as admin' incidents)
- →Audit log + dashboards for tool-call traffic
RAG over internal knowledge
Production RAG, not a demo. Document ingestion pipeline, embedding strategy (chunking, hierarchy, hybrid lexical + dense), retrieval evaluation, citations in every answer, output guardrails. Built on whatever vector store fits — Postgres + pgvector, Qdrant, Pinecone — chosen on cost/latency, not hype.
- →Document ingestion with versioning + incremental updates
- →Hybrid retrieval (BM25 + dense embeddings) + reranking
- →Citation-grounded answers with link-back to source
- →Evaluation harness (precision/recall on a labeled set)
LLM observability & cost engineering
Most teams ship LLM features blind. I install the observability layer (Langfuse, OpenTelemetry, custom dashboards), the cost-tracking layer (per-feature, per-user, per-model), and the model-routing layer that lets you pick the cheapest model that still passes your quality bar.
- →Per-request traces with prompt, completion, tokens, latency
- →Cost attribution by team / feature / customer
- →Smart model router (rules or learned)
- →Alerts on quality regressions and cost spikes
Selected engagements
Clients are anonymized — standard consulting practice. Profile and context can be verified under NDA.
Mid-size fintech (Series B, ~80 engineers)
Problem
Engineering leadership saw Claude Code being used ad-hoc by individual devs but had no organizational adoption, no shared config, and no visibility. Some teams built clever workflows; others were stuck running Claude in default mode with no permissions guardrails.
Concrete asks: standardize on a per-repo configuration, build org-wide skills/slash-commands for their stack (Go + React + Postgres), wire Claude into their internal tools (Linear, GitHub, their internal feature-flag service), and produce metrics that engineering leadership could review monthly.
Approach
- Started with one pilot team of 12 engineers. Wrote a baseline .claude/settings.json with allowlist-based Bash permissions, denied destructive commands by default (rm -rf, force-push, db drops), and pre-approved their internal CLI tooling.
- Built 6 custom skills: 'review-pr' (runs their lint + test + custom static-analysis), 'add-feature-flag' (talks to their flag service via MCP), 'rollback-migration', 'check-staging', 'ship-it' (PR + assign reviewers based on CODEOWNERS), 'why-flaky' (their flaky-test investigation playbook).
- Built an internal MCP server in TypeScript exposing their Linear, GitHub, and feature-flag APIs. Per-user OAuth — agent acts with the engineer's permissions, not a service account.
- Hook-based PR description generation: on git commit, a post-commit hook calls Claude with the diff to draft a PR description in their format. Saves ~5 min per PR.
- Monthly adoption review: dashboard showing active users, sessions per dev per week, PR-with-AI-assist ratio, and tool-call breakdown by skill.
Decisions
Per-user auth on every MCP tool, not a shared service account. The cost is more setup; the benefit is no 'agent silently does production things' incidents. Audit logs map every tool call to a real engineer.
Made the rollout opt-in for the first 4 weeks. Engineers who ignored it kept ignoring it. Engineers who tried it became internal advocates. Forcing adoption from above poisons the well — let the network effect do the work.
Built the metrics dashboard before the rollout, not after. Leadership pre-committed to specific success metrics. Avoided the 'did this even help?' debate later.
From the code
{
"permissions": {
"allow": [
"Bash(npm:*)",
"Bash(go test:*)",
"Bash(git status)",
"Bash(git diff:*)",
"Bash(./scripts/dev-cli:*)"
],
"deny": [
"Bash(rm -rf:*)",
"Bash(git push --force:*)",
"Bash(git push -f:*)",
"Bash(psql:*drop*)",
"Bash(make deploy:*)"
],
"ask": [
"Bash(git push:*)",
"Bash(./scripts/migrate:*)"
]
},
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write",
"hooks": [{
"type": "command",
"command": "./scripts/run-precommit.sh"
}]
}
]
}
}server.tool(
"linear_search_issues",
"Search Linear issues with the current user's permissions.",
{
query: z.string().describe("Linear search query syntax"),
limit: z.number().int().min(1).max(50).default(20),
},
async ({ query, limit }, { session }) => {
// session.linearToken is the user's OAuth token, not a service account.
const client = new LinearClient({ accessToken: session.linearToken });
const issues = await client.issues({ filter: parseQuery(query), first: limit });
audit.log({
tool: "linear_search_issues",
userId: session.userId,
query,
resultCount: issues.nodes.length,
});
return issues.nodes.map((i) => ({
id: i.identifier,
title: i.title,
state: i.state?.name,
url: i.url,
}));
}
);Outcome
Pilot team adoption went from ~3 active users to 11 of 12 within the engagement. Average time-to-first-PR for new hires dropped from 4 days to 1.5. The 'review-pr' skill caught a class of bugs (missing migration in PR) that their CI hadn't been configured to flag.
After the engagement they pulled the same playbook into 4 more teams. I stay on a retainer for ongoing skill-building and MCP server maintenance.
Mid-size law firm (~120 attorneys, Russia + CIS)
Problem
The firm had ~60,000 pages of internal precedent: memos, deal summaries, regulatory updates, court filings. Stored across SharePoint, an old Confluence, and personal drives. Associates spent an embarrassing amount of time hunting for 'we did something like this in 2022, find me that memo'.
Requirements: searchable in Russian + English, every answer must cite the exact paragraph it came from (no hallucinations on legal content), respect document-level access control (associates can't see partner-level memos), and run on infrastructure they could host themselves (data sovereignty).
Approach
- Self-hosted stack: Postgres + pgvector (no external vector DB), open-source bge-m3 multilingual embeddings, Claude Sonnet via API for synthesis, with strict citation requirements baked into the system prompt.
- Ingestion pipeline: docling for layout-aware PDF parsing, semantic chunking (headings + paragraph boundaries, not raw character splits), one embedding per chunk plus a document-level summary embedding for hierarchical retrieval.
- Hybrid retrieval: BM25 (tsvector in Postgres) + dense vectors, results merged with reciprocal rank fusion. Reranking via cross-encoder for the top 50 candidates.
- ACL at retrieval time: every chunk has an access_level column; the SQL query joins on the requesting user's role. The model never sees content it shouldn't.
- Citation enforcement: the system prompt requires answers in JSON with answer + citations[]. A validator rejects responses missing citations or referencing chunk IDs not in the retrieved set.
- Built an evaluation harness with 200 labeled queries (gold answer + expected citations). Tracked precision@5, recall@10, and citation accuracy weekly throughout development.
Decisions
Hybrid retrieval, not pure dense. Legal text has heavy reliance on exact terminology — case names, statute numbers, defined terms. BM25 catches what embeddings miss; embeddings catch what BM25 misses. The fusion was ~12% better than either alone on the eval set.
Self-hosted embeddings on a CPU box (bge-m3, ~1B params). Saved on API costs and avoided sending privileged content to a third party. Ingest is slower but runs once per document.
No fine-tuning. The eval scores were already strong with retrieval improvements and a careful prompt. Fine-tuning would have added a quarterly maintenance burden and unclear ROI.
From the code
-- Single query: BM25 + dense vector + ACL.
-- :q_tsv = websearch_to_tsquery('russian', :user_query)
-- :q_emb = embedding(:user_query)
-- :user_clearance = associate | partner
WITH bm25 AS (
SELECT chunk_id,
ts_rank_cd(content_tsv, :q_tsv) AS bm25_score
FROM chunks
WHERE content_tsv @@ :q_tsv
AND access_level <= :user_clearance
ORDER BY bm25_score DESC
LIMIT 100
),
dense AS (
SELECT chunk_id,
1 - (embedding <=> :q_emb) AS dense_score
FROM chunks
WHERE access_level <= :user_clearance
ORDER BY embedding <=> :q_emb
LIMIT 100
)
SELECT c.chunk_id, c.content, c.source_doc, c.page,
COALESCE(b.bm25_score, 0) * 0.3 +
COALESCE(d.dense_score, 0) * 0.7 AS hybrid_score
FROM chunks c
LEFT JOIN bm25 b USING (chunk_id)
LEFT JOIN dense d USING (chunk_id)
WHERE b.chunk_id IS NOT NULL OR d.chunk_id IS NOT NULL
ORDER BY hybrid_score DESC
LIMIT 20;class Citation(BaseModel):
chunk_id: str
source_doc: str
page: int
quote: str = Field(min_length=10)
class RagAnswer(BaseModel):
answer: str
citations: list[Citation] = Field(min_length=1)
confidence: Literal["high", "medium", "low"]
def validate_grounding(answer: RagAnswer, retrieved: set[str]) -> None:
"""Reject answers that cite chunks not in the retrieved set."""
unknown = [c.chunk_id for c in answer.citations if c.chunk_id not in retrieved]
if unknown:
raise UngroundedResponseError(
f"Citations refer to chunks not retrieved: {unknown}. "
"Possible hallucination — refusing to surface to user."
)
Outcome
From their pilot user feedback: average time to find a relevant precedent dropped from ~25 minutes to under 2. Citation accuracy hit 94% on the eval set by week 8.
The firm extended to a 6-month retainer for ongoing ingestion pipeline maintenance and quarterly eval set expansion.
B2B SaaS company (~30 engineers, Series A)
Problem
Engineering lead asked for an AI-driven first-pass code reviewer on GitHub PRs. The trial-and-error stage was over — they had piloted three off-the-shelf tools (Copilot Review, CodeRabbit, Greptile) and weren't satisfied with the signal-to-noise ratio. Too many cosmetic nits, not enough useful catches.
What they actually wanted: a reviewer that knew their codebase conventions (lint rules, internal patterns, deprecated utilities), flagged risky changes (auth, billing, migrations), and posted exactly one comment per PR — a summary, not 14 separate inline nits.
Approach
- GitHub Action triggered on PR open + synchronize. Pulls the diff, the changed files in full, and (for risky paths) the touched files' callers via a quick grep pass.
- Routed Claude Sonnet — Haiku was too imprecise on their TS codebase; Opus was overkill for routine PRs. Sonnet on a 4K-token max response kept costs at ~$0.07/PR average.
- Wrote a custom system prompt with their actual conventions: 'use ourBigDecimal for money, never Number'; 'never call /v1/payments directly, go through PaymentService'; 'migrations need a reversible down() within 24h'. About 1200 tokens of repo-specific rules.
- Risk-tier classification: PR touching auth/, billing/, or migrations/ → 'high', triggers a stricter prompt and tags @security-team. PR touching only tests or docs → 'low', shorter response template.
- Output: one structured comment per PR with 'Summary', 'Risk', 'Issues found' (with severity), 'Suggestions'. No inline noise. The team's own engineers handle the line-by-line review — Claude does the framing.
- PII/secret filter on input: any line matching common secret patterns (API keys, tokens) is stripped before sending to the API. Their security team signed off on this before we shipped.
Decisions
One comment per PR, not inline. The team had review fatigue from previous tools; consolidation respected their attention. Engineers can ignore one comment they disagree with; they can't ignore 14.
Risk-tier routing instead of running every PR through the strictest prompt. Cost dropped ~3x and false-positive rate dropped because the high-tier prompt was reserved for changes that genuinely warranted scrutiny.
Shipped with a feedback button (thumbs up/down + optional reason) on every Claude comment. Used the negative feedback weekly to update the system prompt — most fixes were 'this rule isn't actually a rule, our codebase does X here'.
From the code
name: AI code review
on:
pull_request:
types: [opened, synchronize, reopened]
jobs:
review:
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: write
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: actions/setup-node@v4
with: { node-version: "20" }
- name: Classify risk tier
id: tier
run: node scripts/risk-tier.mjs >> "$GITHUB_OUTPUT"
- name: Strip secrets from diff
run: node scripts/scrub-secrets.mjs > /tmp/diff.txt
- name: Run AI reviewer
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
RISK_TIER: ${{ steps.tier.outputs.tier }}
run: node scripts/review.mjs /tmp/diff.txtconst HIGH_RISK_PATHS = [
/^src\/auth\//,
/^src\/billing\//,
/^src\/payments\//,
/^migrations\//,
/\.env(\..+)?$/,
];
const LOW_RISK_PATHS = [
/\.test\.tsx?$/,
/^docs\//,
/^\.github\/workflows\//,
];
export function classifyTier(changedFiles: string[]): "high" | "normal" | "low" {
if (changedFiles.some((f) => HIGH_RISK_PATHS.some((re) => re.test(f)))) {
return "high";
}
if (changedFiles.every((f) => LOW_RISK_PATHS.some((re) => re.test(f)))) {
return "low";
}
return "normal";
}Outcome
After 4 weeks in production: 'useful catch' rate (engineer thumbs-up) was 71%. Time-to-first-human-review on PRs dropped from ~4h to ~45min (engineers triage Claude's summary first, then commit to a deep review).
Engineering lead's own statement at our wrap-up: 'this is the first AI tool we've used that actually feels like a junior reviewer, not a linter pretending to be one.'
How we work
Where you actually are with AI today, what's painful, what would 'good' look like. No deck, just questions.
1-2 page proposal with concrete deliverables, timeline, and success metrics. Fixed-scope or retainer.
Ship the smallest end-to-end thing that proves the approach. Pilot with one team, not the whole org.
Configuration, custom skills, internal documentation, training sessions. Adoption is 60% of the value — I do this part too.
Ongoing maintenance, expansion to new teams, monthly metric reviews. Most engagements continue here.
Get in touch
If you'd like to talk through whether Claude Code or RAG fits your team — reach out. Diagnostic call is free, no obligation.