Research · Leibniz Legible
Reproducing the PHILIUMM Leibniz HTR model, and benchmarking frontier vision-language models on Leibniz's hand (Leibniz Legible, Phase B1)
Read the PDFDOI: 10.5281/zenodo.22782817
Abstract
The first independent reproduction of the character error rate of the ERC PHILIUMM project's handwritten-text-recognition model for Leibniz's hand, measured with a frozen, documented protocol on the model's own 1,878-line validation split: 7.95% (95% CI 7.49 to 8.46) against the claimed 8.33%, word error rate 27.0% against 28.56%, one line in four character-perfect. The same protocol gives the first published numbers for zero-shot frontier vision-language models on the same lines, which are four to ten times worse than the fine-tuned model (36.6% to 78.7% character error rate for the well-behaved models), the larger models most often modernizing the archaic spelling.
Two caveats the report states itself: the validation split is Latin and French by construction, so German and Kurrent are unmeasured; and it is drawn from the dataset's manually corrected subset, so the number is a line-level figure on pre-segmented lines. The model, segmentation model and ground truth are the work of the ERC project PHILIUMM (Laboratoire SPHERE, Université Paris Cité and CNRS), released CC BY 4.0; Leibniz Legible builds on them and reports on them.
Keywords
- Leibniz
- Leibniz Nachlass
- handwritten text recognition
- HTR
- digital humanities
- IIIF
- Leibniz Legible
- reproduction
- character error rate
- benchmark
- vision-language models
- PHILIUMM
- Kraken
Cite this
Canonical deposit: doi.org/10.5281/zenodo.22782817. Select the BibTeX below to copy it.
@misc{atlas_leibniz_legible_philiumm_reproduction,
author = {Atlas, Evan Tabak},
title = {Reproducing the PHILIUMM Leibniz HTR model, and benchmarking frontier vision-language models on Leibniz's hand (Leibniz Legible, Phase B1)},
year = {2026},
month = {sep},
howpublished = {Zenodo},
doi = {10.5281/zenodo.22782817},
url = {https://doi.org/10.5281/zenodo.22782817}
}