MobbleOpen in Mobble ⇢
Technology · Artificial intelligence · published 2026-08-30 · via Tom's Hardware

Volunteer archivists use budget cameras and machine learning to digitize 1,800 rare Urdu books

Image via Tom's Hardware
Image via Tom's Hardware

A group of friends in Pakistan spent a decade digitizing out-of-print Urdu books using Nikon cameras and manual post-processing, accumulating over 900,000 shutter actuations and 526,000 scans. They later developed a machine-learning process trained on their Photoshop edits to automate the correction of perspective and margins. The project highlights the challenges of preserving rare texts without formal funding or institutional support.

Expanded Detail

The Ibteda project began with modest equipment: a single Nikon D5300, LED bulbs, and a photocopier glass sheet used to flatten pages. The team purchased all books personally, with no institutional backing. Urdu's Nastaliq script presented particular difficulties—its dense dots and diacritics required distinguishing genuine text from photographic noise and blemishes. Each volume demanded unique treatment, as page thickness and binding affected perspective and margins.

After the team halted work in April 2026, they developed a machine-learning pipeline trained on their accumulated Photoshop corrections. This automated approach could potentially benefit similar preservation efforts worldwide, though the project's scale—902,000 total shutter actuations across both cameras—underscores the labor involved in such volunteer-driven archival work.

Context

This project demonstrates that meaningful cultural preservation can occur outside

Expanded detail and Context are AI-generated analysis; the linked article remains the authoritative source.
Read the full article at Tom's Hardware →
This summary is AI-generated and original to Mobble; the linked article is the authoritative source. Original headline: “DIY archivists push budget Nikons to 902,000 clicks to save 1,800 rare books — team trains neural net on Photoshop edits to process 526,000 scans.” Browse more stories.