Large personal document archives are hard to search and organize
People with large personal and research document collections struggle to find useful content across PDFs, scans, and image files. Folder browsing and filenames often aren’t enough, while OCR, metadata cleanup, and manual filing can take substantial time and still leave archives difficult to search. Existing tools may not fit an existing folder structure or provide reliable, low-friction indexing across file types.
For people with large personal and research document archives. Mentioned from Sep 2018 to Jul 2026 on Bluesky, product forums, GitHub, Hacker News and Stack Exchange.
+
45 different people described this problem in 43 separate discussions.
- Indie fit
- 6.0/10
- Pain
- 6.2/10
- Frequency
- 10.0/10
- Willingness to pay
- 3.0/10
- Momentum
- 4.7/10
- Who pays
- Consumers
- Competition
- High
- Build difficulty
- Medium
What people said
Quoted word for word. Follow a link to read the whole discussion.
I would pay for a foss paperless ngx fork with support for running in a readonly filesystem of arbitrary file structure, and giving me full text search with ocr for images, pdfs, and ideally descriptions of video files
For me it's worth done since I keep adding documents and using them, but it took more than an year to be "almost done enough" and it's still unfinished after 4 years
Build brief
See what to build and who will buy it
- 2 product ideas with the smallest useful version and pricing
- 5 places to find your first customers
- 43 more quotes from people who have this problem
- Current workarounds, existing solutions and risks