|

How to Build an End-to-End OCR Pipeline with Baidu’s Unlimited-OCR for High-Resolution Images and Multi-Page PDF Parsing

In this tutorial, we construct a whole workflow for operating Baidu’s Unlimited-OCR mannequin on doc pictures and multi-page PDFs. We configure the GPU setting, set up the required dependencies, load the 3B-parameter vision-language mannequin with automated number of bfloat16 or float16, and generate structured pattern paperwork for testing. We then consider each the tiled Gundam inference mode and the quicker Base mode for single-page OCR earlier than extending the pipeline to multi-page PDF parsing with PyMuPDF and infer_multi(). Throughout the workflow, we protect long-context era settings, repetition controls, and structured output dealing with to course of dense layouts, tables, paragraphs, and cross-page content material in a reproducible end-to-end pipeline.

import subprocess, sys
def pip_install(*pkgs):
   subprocess.check_call([sys.executable, "-m", "pip", "install", "-q", *pkgs])
print(">> Installing dependencies (1-2 min)...")
pip_install(
   "transformers==4.57.1",
   "Pillow",
   "matplotlib",
   "einops",
   "addict",
   "easydict",
   "pymupdf",
   "psutil",
   "speed up",
)
print(">> Done.")
import os
import torch
from transformers import AutoModel, AutoTokenizer
assert torch.cuda.is_available(), (
   "No GPU detected! In Colab: Runtime -> Change runtime sort -> GPU."
)
gpu_name = torch.cuda.get_device_name(0)
print(f">> GPU: {gpu_name}")
use_bf16 = torch.cuda.is_bf16_supported()
DTYPE = torch.bfloat16 if use_bf16 else torch.float16
print(f">> Using dtype: {DTYPE}")
MODEL_NAME = "baidu/Unlimited-OCR"
print(">> Downloading mannequin (~6 GB for 3B params in BF16). First run takes some time...")
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME, trust_remote_code=True)
mannequin = AutoModel.from_pretrained(
   MODEL_NAME,
   trust_remote_code=True,
   use_safetensors=True,
   torch_dtype=DTYPE,
)
mannequin = mannequin.eval().cuda()
print(">> Model loaded and moved to GPU.")

We set up the required libraries and put together the Google Colab setting for Unlimited-OCR inference. We confirm {that a} CUDA-enabled GPU is on the market and mechanically select bfloat16 or float16 primarily based on {hardware} help. We then load the tokenizer and the 3B-parameter mannequin from Hugging Face, change them to analysis mode, and transfer them to the GPU.

from PIL import Image, ImageDraw, ImageFont
import textwrap
os.makedirs("inputs", exist_ok=True)
os.makedirs("outputs/single_gundam", exist_ok=True)
os.makedirs("outputs/single_base", exist_ok=True)
os.makedirs("outputs/multi_page", exist_ok=True)
def load_font(dimension):
   for path in [
       "/usr/share/fonts/truetype/dejavu/DejaVuSans-Bold.ttf",
       "/usr/share/fonts/truetype/dejavu/DejaVuSans.ttf",
   ]:
       if os.path.exists(path):
           return ImageFont.truetype(path, dimension)
   return ImageFont.load_default()
def make_sample_page(path, page_no):
   W, H = 1240, 1754
   img = Image.new("RGB", (W, H), "white")
   d = ImageDraw.Draw(img)
   title_f, head_f, body_f = load_font(48), load_font(34), load_font(26)
   d.textual content((80, 70), f"Quarterly Operations Report — Page {page_no}",
          fill="black", font=title_f)
   d.line([(80, 145), (W - 80, 145)], fill="black", width=3)
   physique = (
       "This doc demonstrates Unlimited-OCR's one-shot long-horizon "
       "parsing. The mannequin reads an whole web page — headings, paragraphs, "
       "and tables — and emits structured textual content in a single decoding move. "
       "Unlike basic OCR pipelines, no separate layout-analysis stage "
       "is required."
   )
   y = 190
   for line in textwrap.wrap(physique, width=72):
       d.textual content((80, y), line, fill="black", font=body_f)
       y += 40
   y += 30
   d.textual content((80, y), f"Table {page_no}: Regional Revenue (USD, thousands and thousands)",
          fill="black", font=head_f)
   y += 60
   rows = [
       ["Region",  "Q1",   "Q2",   "Q3"],
       ["North",   "12.4", "13.1", "15.0"],
       ["South",   "9.8",  "10.2", "11.7"],
       ["East",    "14.3", "13.9", "16.2"],
       ["West",    "11.1", "12.5", "12.9"],
   ]
   col_w, row_h, x0 = 260, 56, 80
   for r, row in enumerate(rows):
       for c, cell in enumerate(row):
           x = x0 + c * col_w
           d.rectangle([x, y, x + col_w, y + row_h], define="black", width=2)
           d.textual content((x + 14, y + 12), cell, fill="black", font=body_f)
       y += row_h
   y += 50
   footer = (
       f"Note {page_no}: Figures are illustrative. Multi-page mode stitches "
       "context throughout pages, so cross-page references stay coherent."
   )
   for line in textwrap.wrap(footer, width=72):
       d.textual content((80, y), line, fill="black", font=body_f)
       y += 40
   img.save(path)
   return path
IMAGE_PATH = make_sample_page("inputs/sample_page_1.png", 1)
PAGE_2     = make_sample_page("inputs/sample_page_2.png", 2)
PAGE_3     = make_sample_page("inputs/sample_page_3.png", 3)
print(f">> Sample pages written: {IMAGE_PATH}, {PAGE_2}, {PAGE_3}")
import matplotlib.pyplot as plt
plt.determine(figsize=(6, 8))
plt.imshow(Image.open(IMAGE_PATH))
plt.axis("off")
plt.title("Input doc (web page 1)")
plt.present()

We create the required enter and output directories and generate three reasonable pattern doc pages with PIL. We add headings, paragraphs, tables, and footnotes to take a look at the mannequin on structured, layout-rich content material. We additionally preview the primary generated web page with Matplotlib earlier than sending it to the OCR pipeline.

print("n" + "=" * 76)
print("STEP 4: Single picture — GUNDAM mode (tiled, excessive element)")
print("=" * 76)
mannequin.infer(
   tokenizer,
   immediate="<picture>doc parsing.",
   image_file=IMAGE_PATH,
   output_path="outputs/single_gundam",
   base_size=1024,
   image_size=640,
   crop_mode=True,
   max_length=32768,
   no_repeat_ngram_size=35,
   ngram_window=128,
   save_results=True,
)

We run single-image OCR utilizing Gundam mode, which mixes a world doc view with tiled picture crops. We allow crop_mode and use a smaller tile dimension to protect tremendous textual content and enhance recognition on dense doc layouts. We additionally configure long-output era and repetition controls to make sure the mannequin produces steady, structured outcomes.

print("n" + "=" * 76)
print("STEP 5: Single picture — BASE mode (single view, quicker)")
print("=" * 76)
mannequin.infer(
   tokenizer,
   immediate="<picture>doc parsing.",
   image_file=IMAGE_PATH,
   output_path="outputs/single_base",
   base_size=1024,
   image_size=1024,
   crop_mode=False,
   max_length=32768,
   no_repeat_ngram_size=35,
   ngram_window=128,
   save_results=True,
)

We course of the identical doc utilizing Base mode with a single 1024-pixel picture view. We flip off picture cropping to cut back inference complexity and enhance processing velocity for clear, clearly printed pages. We retain the identical output size and repetition-control settings to straight evaluate Base mode with Gundam mode.

print("n" + "=" * 76)
print("STEP 6: Multi-page / PDF parsing")
print("=" * 76)
import tempfile
import fitz
def pdf_to_images(pdf_path, dpi=300):
   """Rasterize each PDF web page to a PNG; return the checklist of picture paths."""
   doc = fitz.open(pdf_path)
   tmp_dir = tempfile.mkdtemp(prefix="pdf_ocr_")
   mat = fitz.Matrix(dpi / 72, dpi / 72)
   paths = []
   for i, web page in enumerate(doc):
       out = os.path.be part of(tmp_dir, f"page_{i + 1:04d}.png")
       web page.get_pixmap(matrix=mat).save(out)
       paths.append(out)
   doc.shut()
   return paths
SAMPLE_PDF = "inputs/sample_doc.pdf"
pdf = fitz.open()
for p in [IMAGE_PATH, PAGE_2, PAGE_3]:
   img_doc = fitz.open(p)
   rect = img_doc[0].rect
   pdf_bytes = img_doc.convert_to_pdf()
   img_pdf = fitz.open("pdf", pdf_bytes)
   web page = pdf.new_page(width=rect.width, peak=rect.peak)
   web page.show_pdf_page(rect, img_pdf, 0)
pdf.save(SAMPLE_PDF)
pdf.shut()
print(f">> Built pattern PDF: {SAMPLE_PDF}")
page_images = pdf_to_images(SAMPLE_PDF, dpi=300)
print(f">> Rasterized {len(page_images)} pages")
mannequin.infer_multi(
   tokenizer,
   immediate="<picture>Multi web page parsing.",
   image_files=page_images,
   output_path="outputs/multi_page",
   image_size=1024,
   max_length=32768,
   no_repeat_ngram_size=35,
   ngram_window=1024,
   save_results=True,
)

We create a three-page PDF from the generated doc pictures and rasterize every web page of the PDF right into a high-resolution PNG utilizing PyMuPDF. We move the ensuing page-image sequence to infer_multi() in order that the mannequin can parse the whole doc in a single long-horizon inference operation. We additionally widen the n-gram repetition window to preserve steady decoding throughout a number of pages.

print("n" + "=" * 76)
print("STEP 7: Saved outputs")
print("=" * 76)
TEXT_EXTS = {".txt", ".md", ".mmd", ".json"}
def show_outputs(root):
   print(f"n--- {root} ---")
   if not os.path.isdir(root):
       print("  (no output listing discovered)")
       return
   for dirpath, _, recordsdata in os.stroll(root):
       for fn in sorted(recordsdata):
           fp = os.path.be part of(dirpath, fn)
           dimension = os.path.getsize(fp)
           print(f"  {fp}  ({dimension:,} bytes)")
           if os.path.splitext(fn)[1].decrease() in TEXT_EXTS:
               with open(fp, "r", encoding="utf-8", errors="exchange") as f:
                   content material = f.learn()
               preview = content material[:1500]
               print("  " + "-" * 60)
               print("n".be part of("  | " + ln for ln in preview.splitlines()))
               if len(content material) > 1500:
                   print(f"  | ... [{len(content) - 1500:,} more chars]")
               print("  " + "-" * 60)
for out_dir in ["outputs/single_gundam", "outputs/single_base", "outputs/multi_page"]:
   show_outputs(out_dir)
print("""
============================================================================
DONE — CHEAT SHEET
============================================================================
Single picture, dense/small textual content .... infer(), gundam (640 + crop_mode=True)
Single picture, clear print ......... infer(), base   (1024, crop_mode=False)
Multi-page or PDF ................. infer_multi(), image_size=1024,
                                    ngram_window=1024
Long paperwork .................... maintain max_length=32768 and the
                                    no_repeat_ngram settings — they stop
                                    degeneration on lengthy outputs.
Your personal recordsdata .................... add through Colab sidebar, level
                                    image_file / pdf_to_images() at them.
============================================================================
""")

We examine the output directories created by the single-page and multi-page inference runs. We checklist each generated file and show previews of supported textual content, Markdown, MMD, and JSON artifacts. We conclude the workflow with a concise reference summarizing the really helpful inference modes for dense pictures, clear pages, and multi-page PDFs.

In conclusion, we accomplished a sensible OCR pipeline that handles each high-detail single-page paperwork and lengthy multi-page PDFs inside Google Colab. We in contrast Gundam and Base inference modes, rasterized PDFs into model-ready web page pictures, ran long-horizon doc parsing, and inspected the generated textual content, Markdown, and auxiliary artifacts straight from the output directories. We additionally configured the workflow to adapt to completely different GPU capabilities whereas retaining the era parameters required for steady long-document decoding. It gives us with a reusable basis for making use of Unlimited-OCR to stories, scanned types, technical paperwork, tables, and different layout-rich content material with out counting on a separate conventional OCR and layout-analysis stack.


Check out the Full Code hereAlso, be at liberty to observe us on Twitter and don’t neglect to be part of our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to associate with us for selling your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar and so on.? Connect with us

The put up How to Build an End-to-End OCR Pipeline with Baidu’s Unlimited-OCR for High-Resolution Images and Multi-Page PDF Parsing appeared first on MarkTechPost.

Similar Posts