Due diligence at scale: sampling contract risk across hundreds of documents
Diligence timelines rarely allow a full read. The question is what to sample and how to know the sample was representative.
Why full review is not the benchmark
Diligence is bounded by a closing date, and the document set is whatever the target has accumulated. Treating a complete read as the standard and sampling as the compromise misdescribes the exercise. Even with unlimited time, reading everything would not be the right allocation of attention, because most contracts in a data room are immaterial to the investment thesis and reading them displaces work on the ones that are not.
The real question is coverage of risk rather than coverage of pages. A diligence exercise is trying to establish a small number of things: whether the revenue in the model is contracted or terminable at short notice, whether the IP is owned, whether change of control triggers anything, which obligations survive the transaction, and whether anything in the portfolio is materially off-market. Each is answerable from a subset of the documents, and the subsets are not the same.
That reframing changes what a sample is for. It is not a random draw meant to represent a population. It is a deliberate selection: the highest-value contracts by revenue, everything with a counterparty of strategic importance, everything signed on the counterparty's paper rather than the target's standard form, plus a random draw from the remaining bulk to test whether the standard form was actually used. The random layer is the one that catches the surprise.
Extracting comparable terms across inconsistent contracts
The obstacle to sampling at scale is not reading speed. It is that the same term appears in different words, different sections and different structures across contracts signed over a decade by people who no longer work there.
Extraction turns that set into a table: one row per contract, one column per term of interest — term length, renewal mechanism, termination rights, notice period, assignment and change of control, liability cap, exclusivity, most-favoured-nation, governing law. Once the portfolio is in that shape the analysis becomes comparison rather than reading, and outliers surface without anyone having decided in advance what to look for. This is the task LexTable is built around, and it is the step that makes the rest of the exercise tractable.
Two cautions. Extraction quality degrades with document quality: scanned agreements, executed amendments filed apart from the base contract, and unsigned drafts sitting in the same folder all produce plausible-looking rows that are wrong. Establishing which version governs is a human step and it comes first. Second, an empty cell is ambiguous — it may mean the contract is silent, or that extraction failed. Those are very different findings, and a table that does not distinguish them will mislead. Ask for a source citation on every extracted value, and treat values without one as unreviewed rather than absent.
Fallback positions and how to spot the ones that matter
Fallback positions are the terms a party accepted when it did not get its preferred position. In aggregate they are a map of where the target had negotiating leverage and where it did not.
Individually most are unremarkable. What matters is the pattern. A liability cap below the target's standard in one contract is a negotiation outcome. The same concession in every contract with a customer above a certain size is a structural feature of the business, and it changes the risk profile of exactly the revenue that matters most. Read one contract at a time, this is close to invisible. Compared across an extracted table, it is obvious.
The categories that usually repay attention are the ones that interact with the transaction rather than with day-to-day operations: change of control and assignment provisions, which determine whether the acquirer inherits the contract at all; exclusivity and non-compete terms that constrain what the business can do afterwards; most-favoured-nation clauses, which propagate a concession made once to everyone; and any obligation with a survival clause.
Which deviations are material to this particular deal is not something extraction supplies. It depends on the thesis, the price and what the acquirer intends to do with the business. Tooling gets the comparison onto one screen. Someone with the deal in their head decides what it means.
Reading financial tables alongside the contracts that create them
Diligence usually runs as two workstreams that meet at the end: financial review of the model and the accounts, legal review of the agreements. But the contracts are what produce the revenue, and the discrepancies live in the gap between the two.
The questions worth asking sit across that boundary. Is the recurring revenue in the model contracted through the forecast period, or does a material share sit on agreements terminable at thirty days' notice? Do the pricing terms in the largest customer contracts match what is actually being invoiced? Are there volume commitments, rebates or most-favoured-nation obligations that the model treats as fixed price? Does anything in the contract base create an obligation that has not been provisioned for?
Answering those means reading a table from a financial statement against a clause in an agreement — different formats, different systems, often different teams. Automated extraction can put both into a comparable structure, and cross-document table comparison is a well-defined task that tools handle reliably when the source documents are machine-readable rather than poor scans.
What tools do not do is decide what a discrepancy means. A mismatch between contracted and invoiced pricing might be an administrative error, an unfiled side letter, or a deliberate commercial accommodation. Surfacing it quickly is the automatable part. Working out which it is remains the diligence.