For IP analytics platforms, search tools and software vendors

Stop maintaining scrapers. Build the product instead.

Patent data for analytics platforms from 170 authorities, and trademark data from around 200. Sourced directly from the offices, normalised into one format, and delivered to you in bulk to provide the data layer that your platform sits on.

From published to usable

Every IP office publishes its data, but only some make it easy to collect. Building your own collection looks affordable for the first thirty or forty offices. Beyond that, the cost climbs quickly.

The coverage cliff

A scraper can reach the thirty to forty patent offices that publish usable digital files. The rest publish PDF gazettes, sit behind portals that block bulk access, or still publish on paper, and some release their archives only to organisations with a local presence in the country.

Downloading is the easy part

Once you have the documents, the entities still need resolving. Our database holds more than 23 million distinct assignee names before consolidation, and every authority formats names, dates and classifications its own way. Making all of it consistent is the work that takes years.

The target keeps moving

Offices change file formats, retire gazettes, move to new platforms and raise their prices. Keeping up is a permanent engineering cost that takes time away from your product. We handle format changes, past and future, as part of the feed.

The cost curve runs the wrong way

Retrieval built on agents and model tokens costs more as token prices rise. Good retrieval surfaces the documents you need before you know which ones they are, and whatever you build on top is only as good as the data underneath it.

How Lighthouse IP produces patent data for analytics platforms: sources, gazettes and datafeeds arrive, then standardisation and cleanup, OCR, keying and conversion, human name translation, data loading and validation, and finally delivery through AWS.

What sits between a gazette and a usable record. Sources arrive as datafeeds, gazettes, PDFs and paper, then go through standardisation, OCR, keying, human name translation and validation before delivery. See how our patent data is produced.

What you get

Six collections in one standardised format. Take one of them, take all of them, or take a single authority. You can build the package that matches the product you are creating.

Diamond bibliographic

The front page of the patent: title, abstract, CPC codes, inventors, assignees and dates, from 170 authorities. This is where the coverage gap against the industry standard datasets is widest.

Explore patent data

Diamond legal

Legal events from more than 130 authorities. Filings, grants, renewals, lapses and oppositions, as the office published them. You get the events and the identifiers, so you can apply your own status rules.

Check authority coverage

Full text

Searchable claims and descriptions from 83 authorities, with machine translation into English. Sourced direct from the offices, including the ones where the only source is an image or a paper file that has to be scanned and read by OCR.

See full text coverage

Documents and images

Multipage PDFs of the original publication, stored as individual files you can address one by one, or pulled on demand through the PDF API when you only need them for specific cases.

More on document delivery

Trademarks

Around 200 authorities, with updates from more than 180, in WIPO ST.66. Goods and services machine translated, owner names translated for China, Japan and Korea. The back file includes authorities that no longer exist, such as the USSR, which are still cited as priority.

Explore trademark data

Designs

Industrial design records alongside the patent and trademark collections, including the offices that publish designs within their patent feeds.

Explore design data
Every record carries an LHIP ID. Each collection works on its own and joins cleanly to the others, so adding full text or legal events later is a simple data load.

How we differ from the industry standard datasets

Most platforms in this market are built, directly or indirectly, on a central aggregation of what national offices report upward. That model limits both coverage and latency, and those limits pass straight on to your product.

Collected directly from the offices

A central aggregator receives a legal event only when the national office reports it, which can take a long time. We collect from each office directly and publish as soon as we have the data, so we control the timing.

Around 11 million additional records

These are records that pipelines built on the standard bibliographic dataset miss. In fast-growing and developing jurisdictions, that extra coverage gives your product a real edge.

Compare like for like

When a coverage sheet looks similar to ours, compare the metadata level and the record count per authority, which tell you more than the country count. We list an authority as full text only when we hold a real proportion of its output and history.

Our published coverage is broken down authority by authority, so you can run that comparison yourself. See global IP data coverage.

The patent and trademark offices behind the data

We collect directly from each office named below, from the USPTO, EPO, WIPO, JPO, CNIPA and KIPO to the national offices that still publish on paper, then normalise all of it into one format. Every entry links to the official office site. The markers show the collections we hold for that authority. Record counts, first publication year and update frequency per authority are on the global IP data coverage page.

PPatent bibliographic dataFTPatent full textTMTrademarksDIndustrial designs
Global and regional systems 8 offices
Europe 52 offices
Americas 47 offices
Asia-Pacific 41 offices
Middle East and Africa 50 offices
Historical authorities (back file only) 7 authorities
  • CzechoslovakiaP FT
  • East GermanyP FT
  • Federal Republic of YugoslaviaP FT
  • Serbia and MontenegroP FT
  • Socialist Federal Republic of YugoslaviaP FT
  • USSR (Soviet Union)P FT
  • Yemen (Aden)TM
Depth differs by office. A few small authorities appear with trademarks only. The historical authorities are back file only, still cited as priority and still delivered. Before you build on a small authority, check its record count and first publication year on the coverage page, or ask us for the per authority publication schedule.

Speed is a product feature, so we treat it as one

Trademark watching runs against a ninety day opposition window. Your users feel every day of latency you inherit from your data supplier.

Under 24 hoursTypical time from office publication to the record being live in our dataset, for the offices that publish frequently.
5 working daysFor authorities that publish only three or four times a year, measured from their publication date.
Near real timeWith S3 bucket replication, our updates land in your bucket as soon as we publish them.
WeeklyA fixed weekly load instead, if that suits your pipeline better. Both options work for us.
Offices publish on their own schedules, from daily to twice a year, and our updates follow their releases. We can share a per authority publication schedule so you can plan ingest around each office's real cadence.

How bulk delivery works

Four simple steps from first look to a live feed. Every update is a full record, so your ingest stays simple and runs on standard tools.

1

Sample bucket

Live production data: a back file plus four weekly updates, with a sample of the Diamond files included. It simulates exactly what switching to us looks like, so your team can test the ingest before anyone talks about price.

2

Back file

The full historical collection on S3, as zipped XML to WIPO ST.36 for patents and ST.66 for trademarks. JSON and other formats are available, and we can also send the back file on a hard drive. Original documents come as multipage PDFs.

3

Updates as full replacements

Updates load in sequence, and every document is a complete replacement: take the record out and put the new one in its place. Your index stays consistent after every update.

4

Bucket replication

Point replication at your own S3 and new records land in your environment as we produce them. Setup is about a day of work on both sides. We have taken a customer from signed contract to a full data set in their own bucket inside two days.

For a corpus of more than 182 million patent documents, bulk delivery is the right tool. Full text and PDFs are also available through an API for targeted, lower volume retrieval, and bibliographic and legal event data are delivered as bulk datasets. Licensing runs per authority, so what you pay follows the authorities you take. If you want AI-ready vectors instead of raw documents, we offer a separate service built on the same corpus.

Go deeper on any of it

If you are evaluating patent data for analytics platforms, these pages carry the detail behind each collection: coverage tables, formats and the whitepapers our data team writes.

Benchmark us against what you already hold

Tell us the authorities and the record types that matter to your roadmap, and we will set up a sample bucket with live production data. Most teams run it against their current provider before they talk commercials, which is exactly how we would want it.

Rather look before you speak to anyone? Browse the collections on our data marketplace and pull a small sample yourself, no conversation required.

Browse the data marketplace