Models
OCR
Vision-language reading of page regions
The OCR model is a vision-language reader. It takes the crops produced by layout and turns each region into machine-readable text. Convert on the Middleware API uses it for scans, photographed pages, and native PDFs that still need structure recovered.