Give your AI the whole patent record.
Keeper AI puts vectorised, primary-source patent data in front of the AI your team already uses: 152M+ full-text patent documents from 83 authorities, retrieved by meaning and delivered where your model runs. Keep your model. Change what it reads.
Where general AI breaks on patent work
Your team already uses AI for prior art, invention disclosures, drafting and competitive intelligence. It is fast and it fits how you work. What holds it back is the data the model can reach.
It writes when you need it to find.
A general model writes the most plausible answer, without the document behind it. In IP work, every citation needs a document you can open and check.
It reads a partial record.
Only a handful of offices publish in a form the open internet can index. Most of the world's patent record sits in PDF gazettes, national portals and paper archives. The answer can only be as complete as what the model can reach.
It loses the thread at scale.
Give a model one document and it does well. Give it a portfolio and it skips, merges and stops. Search has to come first, done by a tool built for finding documents.
Keep your model. Change what it reads.
Keeper AI sits between your model and the world's patent record. Your model asks in plain language, and Keeper AI returns the documents it should read, ranked by meaning, from primary-source data collected directly from the offices. Keeper AI finds the documents, and your model reads and reasons.
The record
Bibliographic data for 182M+ patent documents. Legal events from 130 authorities, as each office publishes them. Full text from 83 authorities. PDFs and images. Trademarks and designs in the same format. Every record carries an LHIP ID, so the data sets join without matching work on your side.
The retrieval
Keeper AI runs two retrieval strategies. Coarse: one vector per document, built from the title, abstract and first claim, for landscaping and competitive intelligence. Fine: sentence-level vectors across the full text, descriptions included, for claim-feature search. Filter by classification and jurisdiction. Results come back as ranked document IDs.
The delivery
Vectorised or raw, your choice. Take the raw record in bulk into your own cloud storage, with full-document replacement updates, and build your own Boolean search on top. Or use the hosted retrieval API for Keeper AI, in early access. A document API handles pulls by number. Agent access over MCP is on the roadmap.
Start with one workflow
Corporate IP teams run the same handful of workflows. Each one needs a different slice of the record and a different retrieval strategy. Pick one and we scope it with you.
Prior art and patentability
Sentence-level retrieval across full text from 83 authorities. Your model reads the candidates and cites document numbers it can open.
Freedom to operate
Claim-feature retrieval filtered by jurisdiction, joined to legal events from 130 authorities, so your model works from the procedural history each office published.
Invention disclosure triage
Coarse retrieval to place a disclosure against the landscape, then fine retrieval on the shortlist. Your model summarises what it found, with the record behind it.
Competitive and technology monitoring
Bulk updates into your cloud, coarse retrieval for the watch, and Corporate Tree to roll competitor filings up to the ultimate parent.
Portfolio and ownership review
Bibliographic data and legal events, with ownership resolved to groups and ultimate parents, so your model reasons over a clean portfolio.
Chemistry
Structures, reactions, tables and Markush claims extracted from patent documents, with a containerised deployment that runs inside your own infrastructure.
Try ChemDoc MinerHow we start
Trials start with a call. We scope each one with you, so early-access places go to teams with a real workflow and a clear measure of success.
Bring one workflow
Prior art, freedom to operate, disclosure triage, monitoring or portfolio review. Tell us what your model needs to read and where it runs.
We scope it in 30 minutes
We agree the data sets, jurisdictions, retrieval strategy and measure of success, and you leave with a clear scope.
You run it in your stack
Your model, your prompts, your security boundary. We supply the documents and the retrieval. Early-access teams shape the Keeper AI roadmap.
Check the record before you talk to anyone
We publish our coverage. Download it, compare it with what your current sources hold, and bring the gaps to the call.
Built for corporate IP teams
The record is the foundation. These are the layers corporate teams ask for most, all joined by the same LHIP ID.
Corporate Tree
Who owns what. Assignee name variants resolved to groups and ultimate parents, checked against business identifiers, with a human in the loop. Query by document ID over an API.
Standard-essential patents
SEP data for cellular and video codec standards, for licensing and exposure work.
Prosecution data
Access to USPTO prosecution data, office actions included, for examiner behaviour and response strategy.
Trademarks and designs
The same primary-source approach to trademark data from around 200 authorities and to design data, delivered in the same format.
Common questions
Do you build the AI application?
That stays with you: your model, your prompts and your workflow. Keeper AI supplies the record and the retrieval, so your team's know-how stays at the centre and your model reads the right documents.
Where does the data live?
Wherever you need it. Bulk data lands in your own cloud storage, with full-document replacement updates that keep your ingest simple. Hosted retrieval runs as an API. Chemistry extraction can run as a container inside your own infrastructure.
What does "legal events" mean?
We deliver the procedural events each office publishes, from 130 authorities, with their dates. Your rules, or your model, then derive the status you need from those events.
How current is the record?
Update frequency and latency differ by authority. Both are published per authority in our coverage documents. See the coverage overview.
What does a trial look like?
After a discovery call we scope one workflow: data sets, jurisdictions, retrieval strategy and a measure of success. Bulk data and the document API can start at once. Hosted retrieval is in early access with limited places.
How is it priced?
Bulk data is licensed per data set. Hosted retrieval is priced on usage. We quote once we understand the workflow and the volume.
Which parts of the record are hard to get elsewhere?
Around 11 million records in our collection sit outside DOCDB. We hold historical authorities such as the USSR, East Germany, Czechoslovakia and Yugoslavia, and full text from 83 authorities. The coverage overview lists every authority.
Go deeper
The record behind Keeper AI, in more detail.
Bring a workflow
Tell us the workflow, the model your team uses and how you access patent data today. In a 30-minute call we map the record and the retrieval to it, and you leave with a scope.
Not ready for a call? Join the early-access list and be first to know when Keeper AI trials open.