Trust & verification
Extraction accuracy
Read this number carefully
92% is measured over 25 expected terms across 7 documents. A sample that size has wide error bars: a single additional miss moves the headline figure by roughly 4 points. Treat it as an early internal benchmark that tells you the pipeline works and how it fails — not as evidence of accuracy on your portfolio. The product is built so that the extraction figure is not what you rely on: nothing enters the record as true until a person approves it against the cited page.
Early internal benchmark
92%
25 terms across 7 lease fixtures · last run 23 Aug 2026
Facts with citations
100%
Every extracted fact carries page + excerpt
Second-pass review
Always on
Adversarial check against raw document text
Verified before use
Every fact
No proposal is treated as true until approved
Methodology
- Sample
- 7 lease, amendment and assignment PDFs; 25 expected terms in total. Fixtures are fixed and versioned in the repository.
- Scoring rule
- Exact match against a hand-written answer key: dates must match to the day, money to the cent, percentages to the stated precision. A near miss counts as a miss. A term the model declines to answer counts as a miss, not as an abstention.
- Error types observed
- Effective-date vs commencement-date confusion on assignments; missed terms where the governing value sits in an exhibit rather than the body. No case so far of a confidently wrong value with a citation pointing at an unrelated page.
- What is not measured
- Scanned or handwritten documents, non-English leases, and portfolios at scale. Ground leases appear in the fixture set, but one fixture is not enough to claim an accuracy rate for that lease type. We will publish a larger benchmark once real customer volume supports one.
How accuracy is measured
- Fixed fixtures. Each benchmark document has a ground-truth answer key for commencement, expiration, rent, renewal notice, and CAM share.
- Adversarial second pass. A separate verifier checks every proposal against the document text and downgrades unsupported or contradicted terms.
- Confidence floors. Dates and money are held to a higher bar than text fields; anything below the threshold is flagged for human review.
- Append-only history. Verified facts cannot be silently edited; amendments create a new fact and preserve the prior version.
What happens to low-confidence terms
Terms that fail the second pass or fall below the confidence floor never enter the verified record. They land in the verification queue with a reason, so a person can confirm, correct, or reject them.
A missing excerpt, a contradiction with the text, or a value below the type-specific floor will all trigger human review.
Corrections made in the queue become worked examples for future extractions in the same workspace — not model fine-tuning, just better context.
Accuracy by field
Measured 23 Aug 2026 with google/gemini-3.7-flash. Seven lease, amendment and assignment PDFs run end to end through the production extraction pipeline, scored term by term against a hand-written answer key. Both misses were omissions caught in review, not wrong values presented as facts.
| Field | Terms measured | Accuracy | What makes it hard |
|---|---|---|---|
| Commencement date | 5 | 80% | Missed once, on an assignment where the date is the effective date of the transfer rather than a stated commencement |
| Expiration date | 5 | 100% | Includes an amendment that pushes the date |
| Base rent | 5 | 80% | Missed once, on an absolute-net industrial lease where the stated rent is followed immediately by an annual escalation and was read as a schedule rather than a base amount |
| Renewal notice window | 5 | 100% | Windows measured in days from expiration, including a reset by amendment and a 540-day ground-lease option |
| CAM / pro-rata share | 3 | 100% | Retail, office and absolute-net industrial shares |
| Security deposit | 1 | 100% | Single stated amount |
| Tenant entity | 1 | 100% | Tested through an assignment and assumption |
No field is treated as truth on the strength of these numbers. Every extracted term still waits for a person to confirm it against the cited page before it enters the record.
Benchmark fixtures
| Fixture | Expected terms | Purpose |
|---|---|---|
| Standard retail lease | 5 | Commencement, expiration, base rent, renewal notice, CAM share |
| Industrial NNN lease | 6 | Ten-year term, fixed escalations, deposit, right of first refusal |
| Office lease with abatement | 5 | Free-rent period, TI allowance, percentage escalation |
| Ground lease with options | 4 | Two five-year options, notice measured from expiration |
| Amendment extends term | 2 | Supersession and notice-window reset |
| Amendment resets rent | 1 | Mid-term rent change with expiration untouched |
| Assignment and assumption | 2 | New tenant entity extraction |
Request a Lease Truth Audit
Send up to three of your own lease documents. We run them through the same extraction and verification pipeline described on this page and return a verified fact sheet — every term cited to the page it came from — plus the deadlines those facts produce.
- Up to three documents: an original lease and its amendments work best
- A fact sheet with a page citation behind every extracted term
- The obligations those facts derive, with the arithmetic shown
- Anything the model could not support is flagged rather than guessed
Your documents are used only to produce your audit and are never used to train models.