Give your AI a record it can cite.
Your AI's answer is only as defensible as the record behind it. Keeper AI puts vectorised, primary-source patent data in front of the AI your firm already uses: 152M+ full-text patent documents from 83 authorities, retrieved by meaning and delivered inside your own environment. Keep your model. Change what it reads.
Where general AI breaks on patent work
Associates and partners already use AI for prior art searches, office action responses, opinion work, drafting and competitive intelligence for clients. It is fast. What holds it back is the data the model can reach, and whether anyone can follow its trail.
It writes when your opinion needs a citation.
A general model writes the most plausible answer, without the document behind it. In opinion work, every reference needs a document you can open, check and cite.
It reads a partial record.
Only a handful of offices publish in a form the open internet can index. Most of the world's patent record sits in PDF gazettes, national portals and paper archives. The opinion can only be as complete as what the model can reach.
Its sources are hard to trace.
Give a model one document and it does well. Give it a client's portfolio and it skips, merges and stops, and it is hard to see which documents it read. Legal work rests on a chain of authority, so search has to come first, done by a tool built for finding documents, with a trail you can follow.
Keep your model. Change what it reads.
Keeper AI sits between your firm's model and the world's patent record. Your model asks in plain language, and Keeper AI returns the documents it should read, ranked by meaning, from primary-source data collected directly from the offices. Keeper AI finds the documents, and your model reads and reasons.
The record
Bibliographic data for 182M+ patent documents. Legal events from 130 authorities, as each office publishes them. Full text from 83 authorities. PDFs and images. Trademarks and designs in the same format. Every record carries an LHIP ID, so the data sets join without matching work on your side.
The retrieval
Keeper AI runs two retrieval strategies. Coarse: one vector per document, built from the title, abstract and first claim, for landscaping and competitive intelligence. Fine: sentence-level vectors across the full text, descriptions included, for claim-feature search. Filter by classification and jurisdiction. Results come back as ranked document IDs.
The delivery
Vectorised or raw, your choice. Take the raw record in bulk into your own cloud storage, inside your security boundary, with full-document replacement updates, and build your own Boolean search on top. Or use the hosted retrieval API for Keeper AI, in early access, which returns ranked document IDs. A document API handles pulls by number. Agent access over MCP is on the roadmap. Lighthouse IP is ISO 27001 certified.
Start with one matter type
Firms run the same handful of patent workflows for clients. Each one needs a different slice of the record and a different retrieval strategy. Pick one and we scope it with you.
Prior art and patentability opinions
Sentence-level retrieval across full text from 83 authorities. Your model reads the candidates, and the opinion cites document numbers you can open.
Freedom to operate opinions
Claim-feature retrieval filtered by jurisdiction, joined to legal events from 130 authorities, so the opinion rests on the procedural history each office published.
Office action responses
Access to USPTO prosecution data, office actions included, for examiner behaviour and response strategy, joined to the record by LHIP ID.
Litigation and licensing support
Legal events, SEP data for cellular and video codec standards, and ownership resolved to ultimate parents, for exposure, licensing and dispute work.
Client portfolio work
Bibliographic data and legal events, with ownership resolved to groups and ultimate parents through Corporate Tree, so your model reasons over a clean portfolio for each client.
Business development
Who owns what and who is filing where, with assignee name variants resolved to groups and ultimate parents, so the firm knows a prospect's portfolio before the first conversation.
How we start
Trials start with a call. We scope each one with you, so early-access places go to firms with a real matter type and a clear measure of success.
Bring one matter type
Prior art, freedom to operate, office action responses, litigation support or client portfolio work. Tell us what your model needs to read and where it runs.
We scope it in 30 minutes
We agree the data sets, jurisdictions, retrieval strategy and measure of success, and you leave with a clear scope.
You run it in your own environment
Your model, your prompts, your security boundary. We supply the documents and the retrieval. Early-access firms shape the Keeper AI roadmap.
Check the record before you talk to anyone
We publish our coverage. Download it, compare it with what your current sources hold, and bring the gaps to the call.
Built for law firms
The record is the foundation. These are the layers firms ask for most, all joined by the same LHIP ID.
Corporate Tree
Who owns what. Assignee name variants resolved to groups and ultimate parents, checked against business identifiers, with a human in the loop. Query by document ID over an API.
Standard-essential patents
SEP data for cellular and video codec standards, for licensing, exposure and dispute work.
Prosecution data
Access to USPTO prosecution data, office actions included, for examiner behaviour and response strategy.
Trademarks and designs
The same primary-source approach to trademark data from around 200 authorities and to design data, delivered in the same format, for practices that span all three rights.
Common questions
Do you build the AI tool, or replace the one we use?
Your firm keeps its own tool: its model, its prompts and its workflow. Keeper AI supplies the record and the retrieval, so your attorneys' know-how stays at the centre and the model reads the right documents.
Can we cite what the model found?
Yes. Keeper AI returns ranked document IDs from the primary-source record. Every result is a published document with a number you can open, check and cite.
Where does the data live?
Wherever your firm needs it. Bulk data lands in your own cloud storage, inside your security boundary, with full-document replacement updates that keep your ingest simple. Hosted retrieval runs as an API and returns ranked document IDs. Lighthouse IP is ISO 27001 certified.
What does "legal events" mean?
We deliver the procedural events each office publishes, from 130 authorities, with their dates. Your rules, or your model, then derive the status you need from those events.
How current is the record?
Update frequency and latency differ by authority. Both are published per authority in our coverage documents. See the coverage overview.
What does a trial look like?
After a discovery call we scope one matter type: data sets, jurisdictions, retrieval strategy and a measure of success. Bulk data and the document API can start at once. Hosted retrieval is in early access with limited places.
How is it priced?
Bulk data is licensed per data set. Hosted retrieval is priced on usage. We quote once we understand the matter type and the volume.
Which parts of the record are hard to get elsewhere?
Around 11 million records in our collection sit outside DOCDB. We hold historical authorities such as the USSR, East Germany, Czechoslovakia and Yugoslavia, and full text from 83 authorities. The coverage overview lists every authority.
Go deeper
The record behind Keeper AI, in more detail.
Bring a matter type
Tell us the matter type, the model your firm uses and how you access patent data today. In a 30-minute call we map the record and the retrieval to it, and you leave with a scope.
Not ready for a call? Join the early-access list and be first to know when Keeper AI trials open.