OCR and Handwritten Text Recognition for Archives: What to Expect

To a computer, a volume of 1850s land records is a stack of pictures. Every page has been scanned, the images are sharp, and you can zoom in until the pen strokes look like brushwork. Not one word of it can be searched. For most archives and recording offices, turning those pictures into findable text is the obvious next step after imaging, and handwritten text recognition, or HTR, is the technology that has made it realistic for collections that will never get a full human transcription.
It is also the subject of more confident nonsense than anything else in my corner of the field. So before the project plan, some ground-clearing.
Myth and fact
Myth: Our OCR software will read the deed books. Fact: Optical character recognition was built for printed type, where every letter in a font looks the same each time it appears. Run it over a page of handwriting and you get gibberish. HTR works differently. It reads whole lines of script using models trained on examples of handwriting, so it can cope with letters that join, vary and change shape depending on their neighbors. Many projects need both, because a volume can switch from a clerk's pen to a typewriter partway through the twentieth century.
Myth: HTR is either perfect or useless. Fact: Accuracy runs along a spectrum, usually measured as a character error rate: the share of characters the machine got wrong compared with a careful human transcription. On a clean, consistent hand, a well-trained model can be good enough that someone skimming the output barely notices the mistakes. On faded, cramped or many-handed pages the output can be rough. Rough output still has uses. A transcript with an error every few words is no good as an edition and perfectly good for finding which page mentions a surname.
Myth: You have to train your own model from scratch. Fact: Platforms such as Transkribus and the open-source eScriptorium offer general models for common scripts and periods, and some do a respectable job on 19th-century American English hands with no training at all. Fine-tuning a general model on a few dozen pages of your own material often brings a large improvement for far less effort than starting from nothing.
Myth: Once the text exists, the work is done. Fact: Machine output has to be checked, tied to the right page images, loaded into something people can search, and stored with a record of which model produced it. A better model a few years on will produce different text, and you will want to know which version a researcher cited.
Myth: Full-text search makes a name index unnecessary. Fact: Names are the hardest words for a model to get right and the ones researchers care about most. Spellings wander (Boucher, Bushey, Bouchey) even when the clerk wrote clearly. A reviewed name index built from HTR output is often more useful to genealogists than raw full text, and the two work best together.
Myth: Volunteers will transcribe it for free. Fact: Crowdsourced transcription, through platforms such as FromThePage or Zooniverse or a society's own volunteers, can do wonderful work. It isn't free. Somebody has to write the guidelines, answer questions, review submissions and thank people, every week, for the life of the project.
Myth: If it was public on paper, it's fine to make it searchable. Fact: Search changes exposure. An identifying number on page 412 of a mid-century volume was effectively hidden from anyone who didn't already know where to look. Once it is indexed text, it's one query away from anyone, including automated scrapers. More on this below.
What to expect from 19th-century clerk hands
The good news about recording-office volumes is that they were written to be read. Clerks learned copybook hands (in much of 19th-century America, Spencerian or a related round hand), wrote carefully because mistakes had legal consequences, and repeated the same formulas page after page. Models love repetition. The boilerplate of a deed, with its "know all men by these presents" and "to have and to hold", gives the model easy wins and supplies context for the harder words around it.
The less good news:
- Clerks change. A new clerk after an election means a new hand, sometimes in the middle of a volume. Expect accuracy to shift when the hand does, and include samples of every major hand in your training material.
- Abbreviations and conventions. Superscript endings, "do." for ditto, "inst." and "ult." for this month and last, ampersands, and symbols for money and measures. Decide whether your transcription expands them or records them as written, then stick to it.
- Names, numbers and boundaries. Proper names, sums of money and metes-and-bounds descriptions ("thence north twelve degrees east forty rods to a stake") are where errors cluster and where they matter most.
- Ink and paper. Faded ink, show-through from the other side of the leaf, and the brown halo of iron gall ink corrosion all reduce accuracy. Good imaging helps more than any model setting.
- Layout. Marginal notes recording a later discharge of a mortgage, words squeezed between lines, pasted-in slips, and index volumes laid out as tables. Tables in particular need layout recognition that understands rows and columns, not just lines.
Our guide on reading a 19th-century deed book covers these conventions from the human side, and it's worth reading before you write transcription guidelines.
Keyword spotting when transcription falls short
When full transcription of a set of volumes comes out too rough to trust, keyword spotting is a useful fallback. Instead of committing to one reading of each word, the system scores how likely each word image is to match a search term and returns ranked hits, usually with a confidence level and a snippet of the page image. A researcher looking for a surname gets a list of likely spots and checks each one by eye. It suits name-heavy, poor-quality material especially well, and it needs less correction effort up front.
Planning the project, step by step
- Decide what you actually need. Search across volumes, a name index, a full transcription and a scholarly edition are four very different projects. Most archives need the first two.
- Choose a pilot. Pick a sample that covers the range: early and late volumes, the best and worst hands, an index volume, a damaged one. A few hundred pages is plenty to learn from.
- Check the images. HTR is only as good as the scans. Sharp focus, even lighting, flat pages and a clean gutter all matter; see scanning bound volumes without breaking them. Keep the masters in a preservation format, as described in our piece on file formats for long-term preservation, and run recognition on derivatives.
- Run layout analysis. Have the software find text regions and baselines, then correct them on your sample pages. Bad layout produces bad text no matter how good the model is.
- Make ground truth. Transcribe a set of pages by hand, carefully and consistently, following written guidelines for abbreviations, line breaks, deletions and illegible words. Consistency matters more than which rules you choose.
- Train and test honestly. Fine-tune on part of the ground truth and measure on pages the model has never seen. Testing on training pages flatters every model.
- Judge by the task, not only the score. Search your test volumes for twenty names you know are there. Whether a researcher can find them matters more than the error rate.
- Run at scale, then sample. Process the collection, spot-check across volumes and hands, and route the worst volumes to keyword spotting or human transcription.
- Publish and preserve. Deliver text alongside the images, ideally in a standard layout format such as ALTO or PAGE XML so each word stays tied to its position on the page. Record the model and version used for every batch.
Privacy and redaction of searchable text
Recording offices in particular should think hard before publishing the full text of 20th-century volumes. Land records can include recorded military discharge papers, which in the later decades of the century often carried Social Security numbers. Probate files name minors, and vital records can note causes of death. Some states require certain identifiers to be redacted from records published online, so check what yours says before anything goes public.
Options, from least to most restrictive:
- redact identifiers in both the text layer and the displayed image, and log each redaction
- publish full text for older volumes and only an index for more recent ones
- keep full text available to staff alone, for answering requests
- leave certain series out of online publication entirely
One more caution: read the data terms of any cloud recognition service before uploading images of restricted records. Some services reserve rights to use what you upload, and a public office may not be permitted to agree to that.
Cost and effort: where the hours go
I can't give you a per-page price that will hold up, and anyone who quotes one without seeing your volumes is guessing. I can tell you where the effort goes.
| Effort driver | Pushes effort down | Pushes effort up |
|---|---|---|
| Number of hands | One or two long-serving clerks | Dozens of writers, frequent turnover |
| Image quality | Sharp, flat, even, good contrast | Gutter shadow, show-through, faded ink |
| Layout | A single text block per page | Tables, marginalia, pasted slips |
| Accuracy target | Searchable for names and places | Publication-quality transcription |
| Review | Spot-checks on samples | A person checks every page |
| Language and script | One language, one script | Mixed languages, Latin phrases, shorthand |
In my experience the computing is the cheap part. Ground truth, review, layout correction and coordination are where projects spend their hours, so budget staff and volunteer time first and software second. Run the pilot, write down how long each step actually took, and build the estimate for the full project from those numbers rather than from anyone's brochure, mine included.


