<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Notes on shipping LLM systems — Vitalii Bogachev</title>
    <link>https://talhayme.github.io/blog/</link>
    <description>Field notes from running LLM products in production: evaluation gates, provider fallback, retrieval that refuses to guess, and the incidents that taught me each one.</description>
    <language>en</language>
    <atom:link href="https://talhayme.github.io/blog/rss.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>3,355 times a coding agent edited a test</title>
      <link>https://talhayme.github.io/blog/3355-agent-test-edits</link>
      <guid isPermaLink="true">https://talhayme.github.io/blog/3355-agent-test-edits</guid>
      <description>SWE-chat publishes real coding-agent sessions, most of them Claude Code, with the commits they produced. I pulled every commit where the agent changed a test and the code together, ranked them, ran seven end to end against pr-witness, and found the one thing the cross-run could not see.</description>
      <pubDate>Sun, 11 Oct 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>What the trainer does to six tool-calling datasets</title>
      <link>https://talhayme.github.io/blog/six-tool-calling-datasets</link>
      <guid isPermaLink="true">https://talhayme.github.io/blog/six-tool-calling-datasets</guid>
      <description>The most-downloaded function-calling datasets, rendered through the model&apos;s own tokenizer and chat template. At max_seq_len=2048, 44% of the rows in hermes-function-calling-v1&apos;s two func_calling configs are cut before the answer. Plus two things met on the way: what a known Ollama bug costs (the model gets a Go struct dump instead of the tool schema), and a pre-fix template in Qwen&apos;s official GGUFs.</description>
      <pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>CI only ever runs the new tests against the new code</title>
      <link>https://talhayme.github.io/blog/ci-only-runs-new-tests</link>
      <guid isPermaLink="true">https://talhayme.github.io/blog/ci-only-runs-new-tests</guid>
      <description>An agent breaks a behaviour, edits the test to agree, and CI goes green. pr-witness runs the base branch&apos;s tests against the pull request&apos;s code, checks the description&apos;s claims against the real run, and signs the result. Built measuring-stick first; every number here was measured, including the ones that embarrassed me.</description>
      <pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Your Claude Code skills cost 7,155 tokens on every request</title>
      <link>https://talhayme.github.io/blog/skills-cost-context</link>
      <guid isPermaLink="true">https://talhayme.github.io/blog/skills-cost-context</guid>
      <description>Every installed skill injects its description into the system prompt of every turn, used or not. I measured mine: 64 skills, 7,155 tokens per request, half from one source, and one skill installed twice. Here is how to measure yours.</description>
      <pubDate>Fri, 09 Oct 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>What it actually takes to run an LLM product on your own</title>
      <link>https://talhayme.github.io/blog/shipping-llm-saas-solo</link>
      <guid isPermaLink="true">https://talhayme.github.io/blog/shipping-llm-saas-solo</guid>
      <description>A document-translation service on GPT-4 and Claude: the architecture, the reliability work nobody demos, and the engineering decisions that decide whether an AI product survives contact with paying customers.</description>
      <pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate>
    </item>
  </channel>
</rss>