Skip to main content
The OCR model is a vision-language reader. It takes the crops produced by layout and turns each region into machine-readable text. Convert on the Middleware API uses it for scans, photographed pages, and native PDFs that still need structure recovered.

Configuration