The molecule you need is in that patent. It is a drawing, so no search will ever find it.
We turn the structures drawn, named, tabled and claimed in patent documents into validated, machine-readable chemistry. Every structure traces back to the exact image it came from.
The chemistry in a patent is the part everyone needs and nobody can search
Molecules are drawn as pictures, written as systematic names, listed in tables, and claimed as generic structures that stand for thousands of compounds that are never drawn at all. Four things block a reader today.
Structures are pictures
A patent draws its compounds. Keyword search cannot see them, and neither can a language model reading the text layer.
Markush claims hide the real scope
A generic scaffold plus tables of R-group definitions can stand for tens of thousands of compounds that are never drawn. The answer to “what does this patent actually cover” sits inside those tables.
Official structure records can mislead
In one US patent we tested, 173 of 198 office-supplied structure files did not match the drawing on the page. Reading the image means you do not inherit that error.
Public chemistry databases skim the surface
On the same herbicide patent, a leading public chemistry database captured 4 claimed compounds. The claim tables define 49,050.1
1 EP2550280B1. Comparison run on seven test patents; figures for this patent shown.
Three ways in
Try it on one document, deploy the engine inside your own network, or subscribe to the chemistry of the global patent corpus as data.
ChemDocMiner
Free web tool. Upload a PDF or image, or fetch a patent by number. Get the structures back with SMILES, InChI, InChIKey, molfile, molecular weight and a confidence value. Correct anything by hand in the editor. Download a CSV that opens straight in DataWarrior.
Try ChemDocMiner freeHosted by MolGenie. Fetch-by-number runs on Lighthouse IP’s patent document warehouse.
Enterprise extraction
The full engine inside your own network, as a container. Reactions, tables and Markush enumeration, working from patent XML and clipped images as well as PDFs. REST API for automation. Nothing leaves your environment.
Talk to usChemistry-annotated patent data
In development. Chemistry from the global patent corpus, delivered as data. Weekly forward flows and bulk history: in-line annotated patent text in JSON, stand-off structures in JSON or CSV, SDF files per chemistry patent. Structure search across the corpus to follow.
Register interestUpload it, the ensemble reads it, you check it
Three steps for the chemist. Six stages under the hood.
Upload or fetch
Drop in a PDF or image, or type a publication number.
The ensemble reads it
Each image is classified, then routed to the recognition engine that handles that kind of image best. Chemical names in the text are converted too.
Check, correct, download
See the original drawing beside the recognised structure. Fix anything by hand. Export.
Ingestion
Patent PDFs, XML and clipped image packages, one document or the whole corpus.
Classification
Each image is sorted: single molecule, multi-molecule grid, table, reaction, or no chemistry.
Ensemble recognition
The image goes to the engine that handles that class best. Outputs are cross-checked against each other.
Text-to-structure
Chemical names in the text are converted to structures, including German and half-trivial names.
Chemical validation
Chemistry logic removes impossible atoms and isotopes. Every structure is standardised to a canonical SMILES and InChIKey.
Export
JSON, CSV or SDF, with file, section, position, source image and confidence on every entity.
Built to be checked
In a field full of confident machine output, the argument is not an accuracy number. It is that every result can be traced and verified.
Other tools show you the scaffold. We give you the compounds the claim actually covers.
EP2550280B1 claims its herbicides through a scaffold drawing and substituent tables. The pipeline reads the scaffold from the image, the substituents from the tables, converts the names to structures, and enumerates 49,050 valid, unique compounds.
Every structure traces back to the picture it came from
Each extracted entity carries its file, section, position, source image and a confidence value. Verification takes seconds, not a re-read of the patent.
A structure string is not the same as a real molecule
Every recognised structure passes chemistry logic that removes impossible atoms and isotopes, then standardisation to a canonical SMILES and InChIKey. What reaches you is chemically valid, not merely plausible.
The team that wrote the benchmark built the engine
MolGenie authored the independent comparison of open structure-recognition tools in Digital Discovery (2024) and maintains a validated patent-image benchmark of 1,361 images. On it, the best engine in the pipeline reaches 94.2% accuracy on single structures. The pipeline is also tested on the public JPO, CLEF, USPTO and UOB sets, with the error analysis published as part of the method.
Built for the world’s patents, not just the US
WIPO publishes its tables as images; the US and EP offices as XML. The pipeline reads both. Behind it sits Lighthouse IP’s primary-source corpus: 182M+ documents from 170 authorities, including translated and image-only documents and around 11 million records outside DOCDB.
Real patents, real output
Public patents, run through the pipeline. Fetch any of them yourself and compare.
Kinase inhibitors, 172 pages
1,425 chemical entities extracted from one document.
Fetch this patent in ChemDocMinerHerbicidal pyridothiazines
49,050 claimed compounds enumerated from the claim tables.
Fetch this patent in ChemDocMinerMacrocyclic peptides
A large macrocycle recognised from the drawing. 7 of 7 compounds recovered from a table image.
Fetch this patent in ChemDocMinerKinase inhibitors
A reaction scheme converted to reaction SMILES with reactants and product.
Fetch this patent in ChemDocMinerWho is behind it
Lighthouse IP
Primary-source patent, trademark and design data from 170+ authorities, supplied as data to the teams that build on it. The patent documents behind every extraction on this page come from the Lighthouse IP corpus.
Explore patent dataMolGenie
Cheminformatics specialists in structure recognition, chemical ontologies and patent chemistry. Their Organic Chemistry Ontology is published in PubChem and covers 4.78 million compounds.
Visit MolGenie(Quote from Willem Lagemaat, Lighthouse IP, to come.)
(Quote from Lutz Weber, MolGenie, to come.)
Questions we get asked
Is ChemDocMiner really free?
Yes. The free and registered tiers have page limits per run. A Pro tier exists for heavier use.
Which patents can I fetch by number?
Any publication in Lighthouse IP’s document warehouse, which covers 170 authorities, plus patent office sources such as the EPO.
What formats do I get?
CSV from the free tool. JSON, CSV, SDF and RDF from the enterprise engine and the data feeds.
How accurate is it?
Benchmark-tested on public sets and on MolGenie’s own 1,361-image patent benchmark, with published error analysis. Every structure is shown beside its source image so you can check it yourself.
Can I correct errors?
Yes. The editor lets you fix a structure before you export.
Does it handle Markush claims and reactions?
The enterprise engine does, including enumeration of claimed compounds from substituent tables. The free tool handles single structures from PDFs and images.
Is my document used for anything else?
Only if you leave the sharing option ticked when you upload. Untick it and your results stay private. The enterprise container runs inside your own network.
Can I automate it?
Yes. A REST API with OpenAPI documentation is available on request.
When is the searchable database coming?
It is in development. Register interest below and we will tell you first.
Try it on your hardest patent
Upload a document and look at the output. That is the whole pitch.