How TermOwl makes Claude's contract review verifiable
October 8, 2026 · 6 min read
A lawyer can't rely on a summary they can't check. That single constraint shapes almost every engineering decision in TermOwl, the first-pass contract review tool we build on Claude. The goal isn't an AI that sounds confident about a contract. It's a draft review where every finding points at the exact words it depends on, so a legal professional can confirm or reject it in seconds.
This post walks through the three layers that make that work.
1. A report with a fixed shape
A free-form answer is hard to verify and harder to build a product around. TermOwl defines the review as a schema: document type, parties, an overall risk score, key terms, and a list of issues. Each issue carries a severity, a category, a clause reference, a verbatim quote, an explanation, what to ask for, and suggested replacement wording. Missing protections, obligations and questions to ask round it out.
We pass that schema to the Claude API as a structured output format, so the model returns JSON that matches it, and we validate the result again on our server before anything reaches the user. The field order is deliberate: the verdict and summary come first, so the report starts rendering while Claude is still writing the detailed issues.
2. Reading the whole document, as written
Contracts are full of cross-references: a liability cap in section 9 means little without the indemnity in section 12 and the definitions in section 1. TermOwl sends the entire agreement to Claude in one request. PDFs go in as native document blocks, which lets Claude read the layout and scanned pages; Word files are converted to text first.
Claude uses adaptive thinking to reason through those connections before it writes the report, and we stream a summary of that reasoning to the screen so the reviewer can see what's being analysed. The instructions are explicit about grounding:
- every issue must quote the clause it relies on, copied verbatim, trimmed with an ellipsis if long;
- never invent clauses; anything important that's absent goes under missing protections;
- the uploaded document is untrusted input, so instructions hidden inside it are treated as content, not commands.
3. Checking every quote against the document
Instructions make verbatim quotes likely. They don't make them certain. So once the report is complete, TermOwl checks each quote on the server. We extract the document's plain text (the text layer of a PDF, or the text of a Word file) and normalize both sides, because a quote copied from a typeset contract rarely matches byte for byte:
function normalize(s: string) {
return s
.normalize("NFKC")
.toLowerCase()
.replace(/[“”«»„"]/g, '"')
.replace(/[‘’`]/g, "'")
.replace(/[‐‑‒–—―−]/g, "-")
.replace(/\s+/g, " ")
.trim();
}Long clauses are often trimmed with an ellipsis, so a quote is split into fragments, and every fragment of meaningful length has to appear in the document:
function quoteAppears(quote: string, normalizedDoc: string) {
const parts = quote
.split(/\.\.\.|…/)
.map((p) => normalize(p).replace(/^["'\s.,;:]+|["'\s.,;:]+$/g, ""))
.filter((p) => p.length >= 8);
return parts.length > 0 && parts.every((p) => normalizedDoc.includes(p));
}Each issue then carries a flag. In the report it shows up as a green Quote verified badge, or an amber Check quote badge when we couldn't find the words, which tells the reviewer exactly where to look first. On our sample review of a contractor services agreement, all eleven issues matched the document word for word.
What verification doesn't do
A matching quote proves the clause exists and says what the report claims it says. It doesn't prove the legal analysis is right, and it can't tell you what the contract should say for your client in your jurisdiction. That judgment belongs to the professional, which is why every TermOwl review starts with an AI disclosure and a confirmation that a qualified person will review the output. Scanned PDFs without a text layer can't be checked this way; those issues simply carry no badge.
What's next
The next layer is grounding: retrieving jurisdiction-specific materials and team playbooks so each recommendation can cite the standard it's measured against, not just the clause it's about. The principle stays the same. If the reviewer can't check it, it doesn't belong in the report.
Questions or want to pilot TermOwl with your legal team? Write to hello@termowl.com.
See a verified review on your own contract.