<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>OCR Bench Blog</title>
    <link>https://ocrbench.app/blog/</link>
    <description>Guides and deep-dives on benchmarking LLMs for OCR and field-value extraction — accuracy, cost, and latency, measured on your own documents.</description>
    <language>en</language>
    <atom:link href="https://ocrbench.app/blog/rss.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>CER vs field-level accuracy: why character error rate misleads on structured extraction</title>
      <link>https://ocrbench.app/blog/cer-vs-field-accuracy/</link>
      <guid isPermaLink="true">https://ocrbench.app/blog/cer-vs-field-accuracy/</guid>
      <pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate>
      <description>Character error rate is the classic OCR metric — and the wrong one for structured field extraction. Play with a live invoice, watch CER and field-level accuracy disagree, and see why one maps to business impact and the other doesn&apos;t.</description>
      <category>ocr</category>
      <category>evaluation</category>
      <category>metrics</category>
      <category>accuracy</category>
    </item>
    <item>
      <title>LLM OCR pricing: a calculator for your real cost per document</title>
      <link>https://ocrbench.app/blog/llm-ocr-cost-calculator/</link>
      <guid isPermaLink="true">https://ocrbench.app/blog/llm-ocr-cost-calculator/</guid>
      <pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate>
      <description>Token-based pricing makes LLM document extraction hard to budget. This interactive calculator turns input/output tokens and per-million prices into a real cost per document, per 1,000 documents, and per year — with the exact formula shown.</description>
      <category>ocr</category>
      <category>llm</category>
      <category>cost</category>
      <category>pricing</category>
      <category>calculator</category>
    </item>
    <item>
      <title>How to choose an OCR and document-extraction model in 2026</title>
      <link>https://ocrbench.app/blog/how-to-choose-an-ocr-extraction-model-2026/</link>
      <guid isPermaLink="true">https://ocrbench.app/blog/how-to-choose-an-ocr-extraction-model-2026/</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>Traditional OCR, vision LLMs, or a specialized document-AI API? A practical, vendor-neutral framework for picking the model that reads your documents — scored on accuracy, cost, and latency, not on someone else&apos;s leaderboard.</description>
      <category>ocr</category>
      <category>llm</category>
      <category>document-extraction</category>
      <category>guides</category>
    </item>
    <item>
      <title>How OCR Bench scores field extraction: correct, wrong, missing, hallucinated</title>
      <link>https://ocrbench.app/blog/how-ocr-bench-scores-extraction/</link>
      <guid isPermaLink="true">https://ocrbench.app/blog/how-ocr-bench-scores-extraction/</guid>
      <pubDate>Tue, 16 Jun 2026 00:00:00 GMT</pubDate>
      <description>Diffing strings is a terrible way to grade an LLM. Here&apos;s the normalization-and-comparison engine OCR Bench uses to turn raw model output into a fair, per-field verdict — with dates, currency, and numbers handled properly.</description>
      <category>ocr</category>
      <category>llm</category>
      <category>evaluation</category>
      <category>normalization</category>
      <category>metrics</category>
    </item>
    <item>
      <title>Introducing OCR Bench: know which model reads your documents best</title>
      <link>https://ocrbench.app/blog/introducing-ocr-bench/</link>
      <guid isPermaLink="true">https://ocrbench.app/blog/introducing-ocr-bench/</guid>
      <pubDate>Wed, 10 Jun 2026 00:00:00 GMT</pubDate>
      <description>Stop picking document-extraction models by vibes. OCR Bench runs every model that matters over your own documents and ranks them on accuracy, cost, and latency — one scorecard, real numbers.</description>
      <category>ocr</category>
      <category>llm</category>
      <category>benchmarking</category>
      <category>document-extraction</category>
    </item>
  </channel>
</rss>
