Garbage In, Garbage Out: Engineering Reliable AI Document Extraction Pipelines
DATA SCIENCE / AI
🌎 Mexico 2026
Your OCR problem isn’t really an OCR problem. It’s a pipeline problem. Most document AI pipelines don’t fail because of the model, they fail because of the inputs. When you send your LLMs blurry image…
Your OCR problem isn’t really an OCR problem. It’s a pipeline problem. Most document AI pipelines don’t fail because of the model, they fail because of the inputs. When you send your LLMs blurry images, bad crops, unstructured text, you get high costs, lost context, and unpredictable results. In this session, we'll walk through a real-world end-to-end document extraction pipeline and show how to enforce image quality at capture, extract structured data instead of raw text, and send smaller, smarter payloads to LLMs.
Sobre Daniel Monroy: Daniel Monroy is a bilingual Solutions Engineer at Apryse, specializing in PDF and document technologies. He works with organizations across Latin America and North America to design and demonstrate solutions for document generation, processing, digital signatures, OCR, data extraction, and much more. With a strong technical and customer-focused background, Daniel enjoys helping teams solve complex challenges and exploring the intersection of PDF technology, artificial intelligence, and enterprise document processing.