Lighthouse IP and MolGenie

The molecule you need is in that patent. It is a drawing, so no search will ever find it.

We turn the structures drawn, named, tabled and claimed in patent documents into validated, machine-readable chemistry. Every structure traces back to the exact image it came from.

182M+Patent documents held at source by Lighthouse IP.
170Patent authorities, with coverage from 1782 to today.
49,050Claimed compounds enumerated from the claim tables of a single patent.
Every structurePosition-anchored to the image and document it came from.

The chemistry in a patent is the part everyone needs and nobody can search

Molecules are drawn as pictures, written as systematic names, listed in tables, and claimed as generic structures that stand for thousands of compounds that are never drawn at all. Four things block a reader today.

Structures are pictures

A patent draws its compounds. Keyword search cannot see them, and neither can a language model reading the text layer.

Markush claims hide the real scope

A generic scaffold plus tables of R-group definitions can stand for tens of thousands of compounds that are never drawn. The answer to “what does this patent actually cover” sits inside those tables.

Official structure records can mislead

In one US patent we tested, 173 of 198 office-supplied structure files did not match the drawing on the page. Reading the image means you do not inherit that error.

Public chemistry databases skim the surface

On the same herbicide patent, a leading public chemistry database captured 4 claimed compounds. The claim tables define 49,050.1

1 EP2550280B1. Comparison run on seven test patents; figures for this patent shown.

Three ways in

Try it on one document, deploy the engine inside your own network, or subscribe to the chemistry of the global patent corpus as data.

Try

ChemDocMiner

Free web tool. Upload a PDF or image, or fetch a patent by number. Get the structures back with SMILES, InChI, InChIKey, molfile, molecular weight and a confidence value. Correct anything by hand in the editor. Download a CSV that opens straight in DataWarrior.

Try ChemDocMiner free

Hosted by MolGenie. Fetch-by-number runs on Lighthouse IP’s patent document warehouse.

Deploy

Enterprise extraction

The full engine inside your own network, as a container. Reactions, tables and Markush enumeration, working from patent XML and clipped images as well as PDFs. REST API for automation. Nothing leaves your environment.

Talk to us
Subscribe

Chemistry-annotated patent data

In development. Chemistry from the global patent corpus, delivered as data. Weekly forward flows and bulk history: in-line annotated patent text in JSON, stand-off structures in JSON or CSV, SDF files per chemistry patent. Structure search across the corpus to follow.

Register interest

Upload it, the ensemble reads it, you check it

Three steps for the chemist. Six stages under the hood.

1

Upload or fetch

Drop in a PDF or image, or type a publication number.

2

The ensemble reads it

Each image is classified, then routed to the recognition engine that handles that kind of image best. Chemical names in the text are converted too.

3

Check, correct, download

See the original drawing beside the recognised structure. Fix anything by hand. Export.

Ingestion

Patent PDFs, XML and clipped image packages, one document or the whole corpus.

Classification

Each image is sorted: single molecule, multi-molecule grid, table, reaction, or no chemistry.

Ensemble recognition

The image goes to the engine that handles that class best. Outputs are cross-checked against each other.

Text-to-structure

Chemical names in the text are converted to structures, including German and half-trivial names.

Chemical validation

Chemistry logic removes impossible atoms and isotopes. Every structure is standardised to a canonical SMILES and InChIKey.

Export

JSON, CSV or SDF, with file, section, position, source image and confidence on every entity.

We do not bet your chemistry on one model. The engine behind all three products is ChemPatExtractor, built by MolGenie and run on Lighthouse IP’s primary-source patent data.

Built to be checked

In a field full of confident machine output, the argument is not an accuracy number. It is that every result can be traced and verified.

Other tools show you the scaffold. We give you the compounds the claim actually covers.

EP2550280B1 claims its herbicides through a scaffold drawing and substituent tables. The pipeline reads the scaffold from the image, the substituents from the tables, converts the names to structures, and enumerates 49,050 valid, unique compounds.

Visual to come from MolGenie: the EP2550280B1 claim table on the left, the enumerated compound list on the right.

Every structure traces back to the picture it came from

Each extracted entity carries its file, section, position, source image and a confidence value. Verification takes seconds, not a re-read of the patent.

A structure string is not the same as a real molecule

Every recognised structure passes chemistry logic that removes impossible atoms and isotopes, then standardisation to a canonical SMILES and InChIKey. What reaches you is chemically valid, not merely plausible.

The team that wrote the benchmark built the engine

MolGenie authored the independent comparison of open structure-recognition tools in Digital Discovery (2024) and maintains a validated patent-image benchmark of 1,361 images. On it, the best engine in the pipeline reaches 94.2% accuracy on single structures. The pipeline is also tested on the public JPO, CLEF, USPTO and UOB sets, with the error analysis published as part of the method.

Built for the world’s patents, not just the US

WIPO publishes its tables as images; the US and EP offices as XML. The pipeline reads both. Behind it sits Lighthouse IP’s primary-source corpus: 182M+ documents from 170 authorities, including translated and image-only documents and around 11 million records outside DOCDB.

Real patents, real output

Public patents, run through the pipeline. Fetch any of them yourself and compare.

US20110021525A1

Kinase inhibitors, 172 pages

1,425 chemical entities extracted from one document.

Fetch this patent in ChemDocMiner
EP2550280B1

Herbicidal pyridothiazines

49,050 claimed compounds enumerated from the claim tables.

Fetch this patent in ChemDocMiner
WO2021090855A1

Macrocyclic peptides

A large macrocycle recognised from the drawing. 7 of 7 compounds recovered from a table image.

Fetch this patent in ChemDocMiner
EP1678168B1

Kinase inhibitors

A reaction scheme converted to reaction SMILES with reactants and product.

Fetch this patent in ChemDocMiner

Who is behind it

Lighthouse IP

Primary-source patent, trademark and design data from 170+ authorities, supplied as data to the teams that build on it. The patent documents behind every extraction on this page come from the Lighthouse IP corpus.

Explore patent data

MolGenie

Cheminformatics specialists in structure recognition, chemical ontologies and patent chemistry. Their Organic Chemistry Ontology is published in PubChem and covers 4.78 million compounds.

Visit MolGenie

(Quote from Willem Lagemaat, Lighthouse IP, to come.)

(Quote from Lutz Weber, MolGenie, to come.)

Questions we get asked

Is ChemDocMiner really free?

Yes. The free and registered tiers have page limits per run. A Pro tier exists for heavier use.

Which patents can I fetch by number?

Any publication in Lighthouse IP’s document warehouse, which covers 170 authorities, plus patent office sources such as the EPO.

What formats do I get?

CSV from the free tool. JSON, CSV, SDF and RDF from the enterprise engine and the data feeds.

How accurate is it?

Benchmark-tested on public sets and on MolGenie’s own 1,361-image patent benchmark, with published error analysis. Every structure is shown beside its source image so you can check it yourself.

Can I correct errors?

Yes. The editor lets you fix a structure before you export.

Does it handle Markush claims and reactions?

The enterprise engine does, including enumeration of claimed compounds from substituent tables. The free tool handles single structures from PDFs and images.

Is my document used for anything else?

Only if you leave the sharing option ticked when you upload. Untick it and your results stay private. The enterprise container runs inside your own network.

Can I automate it?

Yes. A REST API with OpenAPI documentation is available on request.

When is the searchable database coming?

It is in development. Register interest below and we will tell you first.

Try it on your hardest patent

Upload a document and look at the output. That is the whole pitch.