Keeper AI for corporate IP teams

Give your AI the whole patent record.

Keeper AI puts vectorised, primary-source patent data in front of the AI your team already uses: 152M+ full-text patent documents from 83 authorities, retrieved by meaning and delivered where your model runs. Keep your model. Change what it reads.

152M+Full-text patent documents, vectorised so Keeper AI can retrieve them by meaning.
182M+Patent documents from 170 authorities in the raw record, for teams that would rather build their own search.
1782 to todayRecords back to the eighteenth century, so your prior art search reaches well beyond what public sources have digitised.
83Authorities with full text, claims and descriptions included, with more being added.

Where general AI breaks on patent work

Your team already uses AI for prior art, invention disclosures, drafting and competitive intelligence. It is fast and it fits how you work. What holds it back is the data the model can reach.

It writes when you need it to find.

A general model writes the most plausible answer, without the document behind it. In IP work, every citation needs a document you can open and check.

It reads a partial record.

Only a handful of offices publish in a form the open internet can index. Most of the world's patent record sits in PDF gazettes, national portals and paper archives. The answer can only be as complete as what the model can reach.

It loses the thread at scale.

Give a model one document and it does well. Give it a portfolio and it skips, merges and stops. Search has to come first, done by a tool built for finding documents.

Keep your model. Change what it reads.

Keeper AI sits between your model and the world's patent record. Your model asks in plain language, and Keeper AI returns the documents it should read, ranked by meaning, from primary-source data collected directly from the offices. Keeper AI finds the documents, and your model reads and reasons.

Your model or agent ChatGPT, Claude, Copilot or your own stack. It asks in plain language, then reads and reasons.
Keeper AI Retrieval by meaning over the vectorised full text. Returns ranked documents with their numbers.
The Lighthouse record 152M+ full-text patent documents from 83 authorities, sourced from the offices and kept current.
Available now

The record

Bibliographic data for 182M+ patent documents. Legal events from 130 authorities, as each office publishes them. Full text from 83 authorities. PDFs and images. Trademarks and designs in the same format. Every record carries an LHIP ID, so the data sets join without matching work on your side.

Early access

The retrieval

Keeper AI runs two retrieval strategies. Coarse: one vector per document, built from the title, abstract and first claim, for landscaping and competitive intelligence. Fine: sentence-level vectors across the full text, descriptions included, for claim-feature search. Filter by classification and jurisdiction. Results come back as ranked document IDs.

Available now

The delivery

Vectorised or raw, your choice. Take the raw record in bulk into your own cloud storage, with full-document replacement updates, and build your own Boolean search on top. Or use the hosted retrieval API for Keeper AI, in early access. A document API handles pulls by number. Agent access over MCP is on the roadmap.

Start with one workflow

Corporate IP teams run the same handful of workflows. Each one needs a different slice of the record and a different retrieval strategy. Pick one and we scope it with you.

Prior art and patentability

Sentence-level retrieval across full text from 83 authorities. Your model reads the candidates and cites document numbers it can open.

Freedom to operate

Claim-feature retrieval filtered by jurisdiction, joined to legal events from 130 authorities, so your model works from the procedural history each office published.

Invention disclosure triage

Coarse retrieval to place a disclosure against the landscape, then fine retrieval on the shortlist. Your model summarises what it found, with the record behind it.

Competitive and technology monitoring

Bulk updates into your cloud, coarse retrieval for the watch, and Corporate Tree to roll competitor filings up to the ultimate parent.

Portfolio and ownership review

Bibliographic data and legal events, with ownership resolved to groups and ultimate parents, so your model reasons over a clean portfolio.

Chemistry

Structures, reactions, tables and Markush claims extracted from patent documents, with a containerised deployment that runs inside your own infrastructure.

Try ChemDoc Miner

How we start

Trials start with a call. We scope each one with you, so early-access places go to teams with a real workflow and a clear measure of success.

1

Bring one workflow

Prior art, freedom to operate, disclosure triage, monitoring or portfolio review. Tell us what your model needs to read and where it runs.

2

We scope it in 30 minutes

We agree the data sets, jurisdictions, retrieval strategy and measure of success, and you leave with a clear scope.

3

You run it in your stack

Your model, your prompts, your security boundary. We supply the documents and the retrieval. Early-access teams shape the Keeper AI roadmap.

Bulk data and the document API are available today. Hosted retrieval is in early access with limited places while Keeper AI matures. Book a discovery call to reserve one.

Check the record before you talk to anyone

We publish our coverage. Download it, compare it with what your current sources hold, and bring the gaps to the call.

Estimated 70%Of the IP analytics market runs on Lighthouse data, by our estimate.
~11MRecords in our collection that sit outside DOCDB, the EPO family database many sources rely on.
152M+Documents with claims and descriptions in the record.
300+IP professionals registered for our IPWatchdog webinar on reducing AI hallucinations with better data.

Built for corporate IP teams

The record is the foundation. These are the layers corporate teams ask for most, all joined by the same LHIP ID.

Corporate Tree

Who owns what. Assignee name variants resolved to groups and ultimate parents, checked against business identifiers, with a human in the loop. Query by document ID over an API.

Standard-essential patents

SEP data for cellular and video codec standards, for licensing and exposure work.

Prosecution data

Access to USPTO prosecution data, office actions included, for examiner behaviour and response strategy.

Trademarks and designs

The same primary-source approach to trademark data from around 200 authorities and to design data, delivered in the same format.

Common questions

Do you build the AI application?

That stays with you: your model, your prompts and your workflow. Keeper AI supplies the record and the retrieval, so your team's know-how stays at the centre and your model reads the right documents.

Where does the data live?

Wherever you need it. Bulk data lands in your own cloud storage, with full-document replacement updates that keep your ingest simple. Hosted retrieval runs as an API. Chemistry extraction can run as a container inside your own infrastructure.

What does "legal events" mean?

We deliver the procedural events each office publishes, from 130 authorities, with their dates. Your rules, or your model, then derive the status you need from those events.

How current is the record?

Update frequency and latency differ by authority. Both are published per authority in our coverage documents. See the coverage overview.

What does a trial look like?

After a discovery call we scope one workflow: data sets, jurisdictions, retrieval strategy and a measure of success. Bulk data and the document API can start at once. Hosted retrieval is in early access with limited places.

How is it priced?

Bulk data is licensed per data set. Hosted retrieval is priced on usage. We quote once we understand the workflow and the volume.

Which parts of the record are hard to get elsewhere?

Around 11 million records in our collection sit outside DOCDB. We hold historical authorities such as the USSR, East Germany, Czechoslovakia and Yugoslavia, and full text from 83 authorities. The coverage overview lists every authority.

Bring a workflow

Tell us the workflow, the model your team uses and how you access patent data today. In a 30-minute call we map the record and the retrieval to it, and you leave with a scope.

Not ready for a call? Join the early-access list and be first to know when Keeper AI trials open.