A vendor claiming 98% property deed extraction accuracy has told you almost nothing about whether the legal description in your abstract is right. The figure usually averages across every field on every document, so easy fields like recording dates can hide errors in the one field that carries title risk. This article explains what accuracy means for deed data, why it swings from one document to the next, and how to test it on your own records before you rely on automation.
Why a Single Accuracy Percentage Hides the Risk
Extraction accuracy is almost always measured per field: for each value the system was supposed to capture, was it correct? A deed with a correct grantor, grantee, and recording date but a wrong legal description still contributes several correct fields to the tally. The headline percentage stays high while the one error that matters sits inside it.
Fields also differ sharply in consequence. A dropped middle initial is a nuisance that a reviewer fixes in seconds. A transposed section, township, or range call can place a conveyance on the wrong ground and become a title defect that surfaces months later. Counting both as “one error” treats them as equals, and they are not.
The metrics, in plain terms
- Precision: of the values the system extracted, what share were correct. Low precision means the output cannot be trusted without checking everything.
- Recall: of the values actually present on the document, what share the system captured. Low recall means reservations, exceptions, or other details get silently missed.
- Field-level accuracy: the share of individual fields that are correct across a set of documents.
- Document-level accuracy: the share of documents in which every required field is correct.
The gap between the last two is where headline numbers mislead. As an illustration, suppose a tool is 95% accurate per field and you extract 20 fields from each deed. If errors were spread evenly, the average deed would carry about one wrong value, and only roughly a third of documents (0.95 raised to the 20th power is about 36%) would come back fully clean. Real errors cluster rather than spread evenly, so the actual figure varies, but the lesson holds: a strong field-level rate can coexist with a weak document-level rate.
Precision and recall also trade off. A system that extracts only when it is sure can post high precision while missing a lot. Ask for both numbers, and ask which fields they cover.
Where Deed Extraction Goes Wrong
Errors tend to come from a handful of predictable sources. Knowing them tells you what to put in a test set.
Source quality
Extraction starts with optical character recognition (OCR), which converts page images into text. Poor scans, skewed pages, faded ink, and stamps or marginal notes laid over text all degrade that step, and everything downstream inherits the damage. Digit confusion is the classic failure: a 1 read as a 7, or a 5 as a 6, inside a book and page reference, an acreage figure, or a range number. These errors look plausible, which makes them hard to spot.
Handwriting and age
Typewritten instruments are manageable; handwritten ones are much harder. Older chains of title, and mineral and oil and gas conveyances from earlier decades, often combine handwriting, unusual phrasing, and degraded paper. A tool that performs well on modern recorded deeds may perform very differently on a 1940s lease assignment.
Legal descriptions
Legal descriptions are the hardest field. A metes and bounds description strings together bearings, distances, monuments, and calls that must stay in order; one missed line changes the parcel. Lot and block references depend on the correct plat and subdivision. Reservations and exceptions often run across page breaks, so a system that reads page by page can capture the start of an exception and lose the end.
Document variety
Warranty, quitclaim, special warranty, trustee’s, and sheriff’s deeds each use different structure and language, and counties add their own formats and recording stamps. Strength on one type does not guarantee strength on another.
Context errors
Some mistakes are not misreads but misunderstandings. Examples include swapping grantor and grantee, taking the acknowledgment date instead of the execution date, or missing that a party signed as trustee or life tenant. The text was read correctly; its meaning was not.
The Fields That Matter Most and How to Weight Them
Because fields carry different risk, a useful scoring scheme weights them. A simple three-tier structure works for most title and land teams, and you can adjust it to your product, whether that is a full abstract, a commitment, or a mineral ownership report.
- Tier 1, critical: legal description, grantor and grantee names, vesting and conveyance type, reservations, exceptions, and covenants. An error here can change who owns what.
- Tier 2, important: recording information (book and page or instrument number), execution and recording dates, consideration, and prior-reference citations. Errors here break the trail back to the source or distort the chain’s timeline.
- Tier 3, supporting: addresses, notary details, and preparer information. Mistakes are worth fixing but rarely change a conclusion.
Assign weights that reflect the gap in consequence. For example, you might count a Tier 1 error as ten times a Tier 3 error, or simply treat any Tier 1 error as automatically failing the document. The second approach is stricter and often closer to how an examiner thinks: a deed with a wrong legal description is not “mostly right.”
Normalize before you score
Scoring needs rules for what counts as a match. “John A. Smith” and “Smith, John A.” are the same person in different formats and should not be penalized. Abbreviation differences, capitalization, and punctuation in names or street suffixes usually belong in the same bucket. A transposed range number, a changed bearing, or a missing exception is a different category entirely and should count as a miss.
Write these rules down before testing. Otherwise the scoring drifts toward whatever makes the result look better, and comparisons between tools or between test rounds stop meaning anything.
How to Test Extraction Accuracy on Your Own Documents
Vendor demos use clean, representative documents. Your work does not. The only reliable measure is a benchmark built from your own files.
- Build the sample. Pull deeds that span your counties, eras, and instrument types. Include your worst scans and a few handwritten or oil and gas conveyances. If every document is clean, the result tells you little about your real workload. A few dozen documents is enough to start, though more gives steadier results.
- Create ground truth. Have an experienced abstractor record the correct value for every field you plan to score. Where possible, have a second reviewer confirm the Tier 1 fields independently. Ground truth that contains its own errors will punish or flatter the tool unfairly.
- Run the tool and score it. Compare each extracted value to ground truth and label it: exact match, acceptable variant (per your normalization rules), omission, or wrong value. Keep omissions and wrong values separate, since they reflect recall and precision respectively.
- Report two views. Calculate field-level accuracy by tier, and document-level accuracy, meaning the share of documents with no Tier 1 errors and the share with no errors at all.
- Study the error types. Counts alone hide patterns. Sort mistakes by county, document age, instrument type, and field. If errors cluster in pre-1960 documents or one county’s recording format, you know where to direct closer review.
- Retest on a schedule. New document sources, scan vendors, and model updates all change results. Rerun the same benchmark periodically, and add fresh samples so the set does not go stale.
When a vendor does quote an accuracy figure, ask for the test conditions: what documents, how many, which fields, how matches were defined, and the date of the test. A number without those details cannot be compared with anything, including your own results.
Building Accuracy Into the Workflow: Human Review and Confidence Signals
No benchmark result removes the need for review. Automation speeds up the first pass; a trained examiner remains responsible for the final abstract and commitment. The practical question is how to make that review fast and targeted rather than a full reread of every deed.
Source-linked extraction
The most useful design feature is traceability: every extracted value should point back to the exact spot on the page. A reviewer can then confirm a grantor name or a call in a metes and bounds description by glancing at the highlighted source, instead of hunting through the document. Verification takes seconds, and mistakes become much harder to miss.
Confidence scores and flags
A confidence score is the system’s own estimate of how likely a value is to be right. Used well, it routes work: high-confidence routine fields get a quick check, while low-confidence fields and all legal descriptions go to closer review. Treat the scores as a prioritization aid, not a guarantee, and test whether low scores actually line up with real errors in your benchmark.
Cross-checks
Some errors only show up when documents are read together. Useful automated or manual checks include:
- Legal description consistency across the chain, so a call that changes between deeds gets flagged.
- Grantee-to-grantor continuity, where each conveyance should start with the previous grantee.
- Acreage reasonableness, comparing stated acreage against what the description implies.
How this fits TitleTrackr
TitleTrackr’s platform is built around this review-first approach. Its document auto-extraction pulls key data from deeds and other instruments, and its instant abstracts assemble that data into a draft for the examiner to verify rather than a finished product to accept on faith. Feature availability changes, so confirm current capabilities, such as source linking and confidence flags, with the team as of 2026. The point is the division of labor: the software does the first pass, and you make the call.
Run Your Own Benchmark Before You Trust Any Number
Useful accuracy has three properties: it is weighted by field so a legal description error counts for more than a formatting quirk, it is tested on your own documents including the ugly ones, and it is backed by human review of whatever the software produces. A single percentage satisfies none of these.
The simplest next step is a small benchmark. Take a recent batch of deeds, have an experienced abstractor mark the correct values, run them through the tool, and score the results by tier. A TitleTrackr trial is one way to do that with your own files and see how extraction, source linking, and review work in practice. Learn more about our services


Leave a Reply