Research · Leibniz Legible
Leibniz Legible: an open access layer for the digitized Leibniz Nachlass (project statement, September 2026)
Read the PDFDOI: 10.5281/zenodo.22782813
Abstract
The Leibniz Nachlass in Hannover, some 236,000 digitized page images, is the largest body of writing by any early-modern philosopher that remains mostly unread: after a century of work the Akademie-Ausgabe has printed roughly a quarter of it. Leibniz Legible is a one-person open project that builds an access layer over those images rather than an edition: every page machine-transcribed with a per-line confidence score and a full provenance record, searchable with tolerance for early-modern orthography, and browsable in an open IIIF viewer that loads the library's own images and shows the machine reading beside them with honest labels.
This statement records the state of the project in September 2026. The digitized corpus has been counted for the first time from the library's own metadata (2,225 works, 236,795 page images). The character error rate of the ERC PHILIUMM project's Leibniz handwriting model has been reproduced under a frozen protocol (7.95%, 95% CI 7.49–8.46, against a claimed 8.33%), and the same protocol gives the first published numbers for frontier vision-language models on Leibniz's hand (37% to 79% character error rate, zero-shot). The whole corpus has been segmented and read with that model: 236,210 pages, 13.5 million lines, each with a confidence and a run record, on one desktop computer in six weeks. A ground-truth factory has aligned the reading text of copyright-expired edition volumes back onto the machine lines and minted 297,424 training pairs, six times the target, at a precision that is measured on favourable material but only preliminarily audited on the corpus. The serving layer (a search index, a JSON API, IIIF Presentation 3 manifests with W3C annotations, and a viewer) is built, and the datasets are packaged with cards that carry their provenance, licence and error rates. The statement closes with what the next model can and cannot buy, what the crowd and language models can and cannot fix, and the roadmap to the public release.
Keywords
- Leibniz
- Leibniz Nachlass
- handwritten text recognition
- HTR
- digital humanities
- IIIF
- Leibniz Legible
- project statement
- access layer
- retro-alignment
- ground truth
- provenance
- open science
Cite this
Canonical deposit: doi.org/10.5281/zenodo.22782813. Select the BibTeX below to copy it.
@misc{atlas_leibniz_legible_project_statement,
author = {Atlas, Evan Tabak},
title = {Leibniz Legible: an open access layer for the digitized Leibniz Nachlass (project statement, September 2026)},
year = {2026},
month = {sep},
howpublished = {Zenodo},
doi = {10.5281/zenodo.22782813},
url = {https://doi.org/10.5281/zenodo.22782813}
}