Keeper AI for law firms

Give your AI a record it can cite.

Your AI's answer is only as defensible as the record behind it. Keeper AI puts vectorised, primary-source patent data in front of the AI your firm already uses: 152M+ full-text patent documents from 83 authorities, retrieved by meaning and delivered inside your own environment. Keep your model. Change what it reads.

152M+Full-text patent documents, vectorised so Keeper AI can retrieve them by meaning.
182M+Patent documents from 170 authorities in the raw record, for teams that would rather build their own search.
1782 to todayRecords back to the eighteenth century, so a prior art search reaches well beyond what public sources have digitised.
83Authorities with full text, claims and descriptions included, with more being added.

Where general AI breaks on patent work

Associates and partners already use AI for prior art searches, office action responses, opinion work, drafting and competitive intelligence for clients. It is fast. What holds it back is the data the model can reach, and whether anyone can follow its trail.

It writes when your opinion needs a citation.

A general model writes the most plausible answer, without the document behind it. In opinion work, every reference needs a document you can open, check and cite.

It reads a partial record.

Only a handful of offices publish in a form the open internet can index. Most of the world's patent record sits in PDF gazettes, national portals and paper archives. The opinion can only be as complete as what the model can reach.

Its sources are hard to trace.

Give a model one document and it does well. Give it a client's portfolio and it skips, merges and stops, and it is hard to see which documents it read. Legal work rests on a chain of authority, so search has to come first, done by a tool built for finding documents, with a trail you can follow.

Keep your model. Change what it reads.

Keeper AI sits between your firm's model and the world's patent record. Your model asks in plain language, and Keeper AI returns the documents it should read, ranked by meaning, from primary-source data collected directly from the offices. Keeper AI finds the documents, and your model reads and reasons.

Your firm's model or agent ChatGPT, Claude, Copilot or the legal AI tool your firm has adopted. It asks in plain language, then reads and reasons.
Keeper AI Retrieval by meaning over the vectorised full text. Returns ranked documents with their numbers.
The Lighthouse record 152M+ full-text patent documents from 83 authorities, sourced from the offices and kept current.
Available now

The record

Bibliographic data for 182M+ patent documents. Legal events from 130 authorities, as each office publishes them. Full text from 83 authorities. PDFs and images. Trademarks and designs in the same format. Every record carries an LHIP ID, so the data sets join without matching work on your side.

Early access

The retrieval

Keeper AI runs two retrieval strategies. Coarse: one vector per document, built from the title, abstract and first claim, for landscaping and competitive intelligence. Fine: sentence-level vectors across the full text, descriptions included, for claim-feature search. Filter by classification and jurisdiction. Results come back as ranked document IDs.

Available now

The delivery

Vectorised or raw, your choice. Take the raw record in bulk into your own cloud storage, inside your security boundary, with full-document replacement updates, and build your own Boolean search on top. Or use the hosted retrieval API for Keeper AI, in early access, which returns ranked document IDs. A document API handles pulls by number. Agent access over MCP is on the roadmap. Lighthouse IP is ISO 27001 certified.

Start with one matter type

Firms run the same handful of patent workflows for clients. Each one needs a different slice of the record and a different retrieval strategy. Pick one and we scope it with you.

Prior art and patentability opinions

Sentence-level retrieval across full text from 83 authorities. Your model reads the candidates, and the opinion cites document numbers you can open.

Freedom to operate opinions

Claim-feature retrieval filtered by jurisdiction, joined to legal events from 130 authorities, so the opinion rests on the procedural history each office published.

Office action responses

Access to USPTO prosecution data, office actions included, for examiner behaviour and response strategy, joined to the record by LHIP ID.

Litigation and licensing support

Legal events, SEP data for cellular and video codec standards, and ownership resolved to ultimate parents, for exposure, licensing and dispute work.

Client portfolio work

Bibliographic data and legal events, with ownership resolved to groups and ultimate parents through Corporate Tree, so your model reasons over a clean portfolio for each client.

Business development

Who owns what and who is filing where, with assignee name variants resolved to groups and ultimate parents, so the firm knows a prospect's portfolio before the first conversation.

How we start

Trials start with a call. We scope each one with you, so early-access places go to firms with a real matter type and a clear measure of success.

1

Bring one matter type

Prior art, freedom to operate, office action responses, litigation support or client portfolio work. Tell us what your model needs to read and where it runs.

2

We scope it in 30 minutes

We agree the data sets, jurisdictions, retrieval strategy and measure of success, and you leave with a clear scope.

3

You run it in your own environment

Your model, your prompts, your security boundary. We supply the documents and the retrieval. Early-access firms shape the Keeper AI roadmap.

Bulk data and the document API are available today. Hosted retrieval is in early access with limited places while Keeper AI matures. Book a discovery call to reserve one.

Check the record before you talk to anyone

We publish our coverage. Download it, compare it with what your current sources hold, and bring the gaps to the call.

Estimated 70%Of the IP analytics market runs on Lighthouse data, by our estimate.
~11MRecords in our collection that sit outside DOCDB, the EPO family database many sources rely on.
152M+Documents with claims and descriptions in the record.
300+IP professionals registered for our IPWatchdog webinar on reducing AI hallucinations with better data, with law firm heads of AI on the panel.

Built for law firms

The record is the foundation. These are the layers firms ask for most, all joined by the same LHIP ID.

Corporate Tree

Who owns what. Assignee name variants resolved to groups and ultimate parents, checked against business identifiers, with a human in the loop. Query by document ID over an API.

Standard-essential patents

SEP data for cellular and video codec standards, for licensing, exposure and dispute work.

Prosecution data

Access to USPTO prosecution data, office actions included, for examiner behaviour and response strategy.

Trademarks and designs

The same primary-source approach to trademark data from around 200 authorities and to design data, delivered in the same format, for practices that span all three rights.

Common questions

Do you build the AI tool, or replace the one we use?

Your firm keeps its own tool: its model, its prompts and its workflow. Keeper AI supplies the record and the retrieval, so your attorneys' know-how stays at the centre and the model reads the right documents.

Can we cite what the model found?

Yes. Keeper AI returns ranked document IDs from the primary-source record. Every result is a published document with a number you can open, check and cite.

Where does the data live?

Wherever your firm needs it. Bulk data lands in your own cloud storage, inside your security boundary, with full-document replacement updates that keep your ingest simple. Hosted retrieval runs as an API and returns ranked document IDs. Lighthouse IP is ISO 27001 certified.

What does "legal events" mean?

We deliver the procedural events each office publishes, from 130 authorities, with their dates. Your rules, or your model, then derive the status you need from those events.

How current is the record?

Update frequency and latency differ by authority. Both are published per authority in our coverage documents. See the coverage overview.

What does a trial look like?

After a discovery call we scope one matter type: data sets, jurisdictions, retrieval strategy and a measure of success. Bulk data and the document API can start at once. Hosted retrieval is in early access with limited places.

How is it priced?

Bulk data is licensed per data set. Hosted retrieval is priced on usage. We quote once we understand the matter type and the volume.

Which parts of the record are hard to get elsewhere?

Around 11 million records in our collection sit outside DOCDB. We hold historical authorities such as the USSR, East Germany, Czechoslovakia and Yugoslavia, and full text from 83 authorities. The coverage overview lists every authority.

Bring a matter type

Tell us the matter type, the model your firm uses and how you access patent data today. In a 30-minute call we map the record and the retrieval to it, and you leave with a scope.

Not ready for a call? Join the early-access list and be first to know when Keeper AI trials open.