Volunteer archivists use budget cameras and machine learning to digitize 1,800 rare Urdu books

A group of friends in Pakistan spent a decade digitizing out-of-print Urdu books using Nikon cameras and manual post-processing, accumulating over 900,000 shutter actuations and 526,000 scans. They later developed a machine-learning process trained on their Photoshop edits to automate the correction of perspective and margins. The project highlights the challenges of preserving rare texts without formal funding or institutional support.
The Ibteda project began with modest equipment: a single Nikon D5300, LED bulbs, and a photocopier glass sheet used to flatten pages. The team purchased all books personally, with no institutional backing. Urdu's Nastaliq script presented particular difficulties—its dense dots and diacritics required distinguishing genuine text from photographic noise and blemishes. Each volume demanded unique treatment, as page thickness and binding affected perspective and margins.
After the team halted work in April 2026, they developed a machine-learning pipeline trained on their accumulated Photoshop corrections. This automated approach could potentially benefit similar preservation efforts worldwide, though the project's scale—902,000 total shutter actuations across both cameras—underscores the labor involved in such volunteer-driven archival work.
This project demonstrates that meaningful cultural preservation can occur outside