OCR users struggle to get reliable results from varied documents
People building OCR and digitization workflows encounter different failure modes across scanned PDFs, unusual layouts, accented languages, and poor-quality source material. They may need to compare engines, prepare pages, tune settings, or clean up text, but the evidence points to a lack of an easy way to evaluate and improve results on their own documents. A focused workbench could help without claiming to solve OCR for every language or document type.
For OCR developers and document-digitization practitioners. Mentioned from Mar 2024 to Dec 2025 on Hacker News and Stack Exchange.
6 different people described this problem in 6 separate discussions.
- Indie fit
- 5.0/10
- Pain
- 5.3/10
- Frequency
- 7.0/10
- Willingness to pay
- 0.0/10
- Momentum
- 5.0/10
- Who pays
- Professionals
- Competition
- Medium
- Build difficulty
- Medium
What people said
Quoted word for word. Follow a link to read the whole discussion.
Application for pdf deskew Looking for a tool that can deskew pdf pages. The pages are scanned images of text. The text (or all of it I'm interested in) is in two columns. This is prepatory for OCR. All the OCR tools I have tried get confused by the columns and put chunks of text in the wrong places.
There have been such a large number of OCR tools pop up over the past ~year; sorely in need for some benchmarks to compare them. Would love to see support for normal OCR tools like tesseract, EasyOCR, Microsoft Azure, etc
Build brief
See what to build and who will buy it
- 1 product idea with the smallest useful version and pricing
- 4 places to find your first customers
- 4 more quotes from people who have this problem
- Current workarounds, existing solutions and risks