Trust & verification

Extraction accuracy

Provenance treats every extracted term as a proposal until a person verifies it. The figure below is an early internal benchmark on a small fixed fixture set — it is not a measure of production accuracy across real portfolios, and we say so on purpose.

Read this number carefully

92% is measured over 25 expected terms across 7 documents. A sample that size has wide error bars: a single additional miss moves the headline figure by roughly 4 points. Treat it as an early internal benchmark that tells you the pipeline works and how it fails — not as evidence of accuracy on your portfolio. The product is built so that the extraction figure is not what you rely on: nothing enters the record as true until a person approves it against the cited page.

Early internal benchmark

92%

25 terms across 7 lease fixtures · last run 23 Aug 2026

Facts with citations

100%

Every extracted fact carries page + excerpt

Second-pass review

Always on

Adversarial check against raw document text

Verified before use

Every fact

No proposal is treated as true until approved

Methodology

Sample
7 lease, amendment and assignment PDFs; 25 expected terms in total. Fixtures are fixed and versioned in the repository.
Scoring rule
Exact match against a hand-written answer key: dates must match to the day, money to the cent, percentages to the stated precision. A near miss counts as a miss. A term the model declines to answer counts as a miss, not as an abstention.
Error types observed
Effective-date vs commencement-date confusion on assignments; missed terms where the governing value sits in an exhibit rather than the body. No case so far of a confidently wrong value with a citation pointing at an unrelated page.
What is not measured
Scanned or handwritten documents, non-English leases, and portfolios at scale. Ground leases appear in the fixture set, but one fixture is not enough to claim an accuracy rate for that lease type. We will publish a larger benchmark once real customer volume supports one.

How accuracy is measured

  • Fixed fixtures. Each benchmark document has a ground-truth answer key for commencement, expiration, rent, renewal notice, and CAM share.
  • Adversarial second pass. A separate verifier checks every proposal against the document text and downgrades unsupported or contradicted terms.
  • Confidence floors. Dates and money are held to a higher bar than text fields; anything below the threshold is flagged for human review.
  • Append-only history. Verified facts cannot be silently edited; amendments create a new fact and preserve the prior version.

What happens to low-confidence terms

Terms that fail the second pass or fall below the confidence floor never enter the verified record. They land in the verification queue with a reason, so a person can confirm, correct, or reject them.

A missing excerpt, a contradiction with the text, or a value below the type-specific floor will all trigger human review.

Corrections made in the queue become worked examples for future extractions in the same workspace — not model fine-tuning, just better context.

Accuracy by field

Measured 23 Aug 2026 with google/gemini-3.7-flash. Seven lease, amendment and assignment PDFs run end to end through the production extraction pipeline, scored term by term against a hand-written answer key. Both misses were omissions caught in review, not wrong values presented as facts.

FieldTerms measuredAccuracyWhat makes it hard
Commencement date580%Missed once, on an assignment where the date is the effective date of the transfer rather than a stated commencement
Expiration date5100%Includes an amendment that pushes the date
Base rent580%Missed once, on an absolute-net industrial lease where the stated rent is followed immediately by an annual escalation and was read as a schedule rather than a base amount
Renewal notice window5100%Windows measured in days from expiration, including a reset by amendment and a 540-day ground-lease option
CAM / pro-rata share3100%Retail, office and absolute-net industrial shares
Security deposit1100%Single stated amount
Tenant entity1100%Tested through an assignment and assumption

No field is treated as truth on the strength of these numbers. Every extracted term still waits for a person to confirm it against the cited page before it enters the record.

Benchmark fixtures

FixtureExpected termsPurpose
Standard retail lease5Commencement, expiration, base rent, renewal notice, CAM share
Industrial NNN lease6Ten-year term, fixed escalations, deposit, right of first refusal
Office lease with abatement5Free-rent period, TI allowance, percentage escalation
Ground lease with options4Two five-year options, notice measured from expiration
Amendment extends term2Supersession and notice-window reset
Amendment resets rent1Mid-term rent change with expiration untouched
Assignment and assumption2New tenant entity extraction

Request a Lease Truth Audit

Send up to three of your own lease documents. We run them through the same extraction and verification pipeline described on this page and return a verified fact sheet — every term cited to the page it came from — plus the deadlines those facts produce.

  • Up to three documents: an original lease and its amendments work best
  • A fact sheet with a page citation behind every extracted term
  • The obligations those facts derive, with the arithmetic shown
  • Anything the model could not support is flagged rather than guessed

Your documents are used only to produce your audit and are never used to train models.

We use your details only to respond to this request. See our privacy policy.