Settings → OCR
Scanned PDFs and images have no text layer, so without a reader they stay outside your memory. This page installs one. The rail entry is OCR; the page is titled OCR — scanned documents and its subtitle reads Read text from scanned PDFs and images. Pick an engine and install it once — it runs locally.
Every control on this page acts as soon as you click it. There is no Save button.
For what happens to a scanned document once an engine exists, see OCR. This page is about the settings.
OCR engine
The section opens with Loading engines…. If the list cannot be read: Could not load the OCR engines. and a Retry button.
Two engines exist. Their names are in English in every language.
| Engine | The app's own note | GPU | Gate |
|---|---|---|---|
| RapidOCR (light) | Light — fast, ~200 MB, no GPU. Best for everyday PDFs. | no | free; the active engine by default |
| MinerU (power, GPU) | Power — best accuracy (layout, tables), GPU, large download. | yes | Engramm license |
Each row is a radio: clicking it makes that engine the active one, but only once it is installed. The row shows a CPU or ⚡ icon, the name, and · Active on the engine in use.
On the GPU engine, a licence line follows the note once the app has an answer (nothing is shown before): · ⚡ Engramm license detected — full power unlocked or · 🔒 Engramm license required — unlocks high-precision OCR: layout, tables and formulas come back intact.
- Not installed: an Install button. While it runs, the button counts Installing 40% with a progress bar and the current stage: Preparing environment, Installing dependencies, Downloading GPU runtime, Installing OCR engine, Downloading OCR models, Finalizing. The install runs in the background and survives leaving the page. A successful install makes the engine active.
- Installed: ✓ Ready, a Test button and a Remove button. Test (Testing…) opens a file dialog (PDF, PNG, JPG, TIFF, WebP, BMP) and prints the raw text under the row in a box titled OCR result, capped at 5 000 characters, or No text detected on this document. Remove uninstalls the engine with no confirmation.
| The row says | Why |
|---|---|
| Another install is already running — wait for it to finish. | Only one install at a time. |
| Requires an Engramm license. | The GPU engine was asked for without a licence. |
| It failed without saying why. | The install failed with no message. |
Any other failure shows its message as is.
MinerU is part of the Engramm license. The padlock on the row is a preview; the refusal happens in the engine itself, which will neither install nor read a page without the licence. The light engine is free and is enough for everyday PDFs.
MinerU installs the GPU runtime first (the Downloading GPU runtime stage), and its models only come down at the first scan it reads. Both engines need Python 3.10 or newer on the machine.
Read what a scan hides
Without a licence, and because a GPU engine exists, a card sits under the engines: Read what a scan hides. An Engramm license unlocks the GPU engine: your scanned books, notarial deeds and archives come back as text your memory can actually search — layout, tables and formulas included. One licence, no subscription. The buy button under it is the real one from the wallet; when nothing is for sale, the machine is offline, or there is no wallet, it renders nothing rather than a dead button.
Scanned PDFs waiting
When the vaults hold scanned PDFs that nobody can read, an hourglass line counts them: n scanned PDF(s) in your vaults are NOT indexed — they have no text layer and no engine to read them. They will be read automatically as soon as one is installed. It disappears once they are read.
The line at the bottom
Text-based PDFs already work without OCR. Install an engine only to read scanned or image PDFs. · Needs Python 3.10+ installed.
Traps
- The GPU engine refuses to install and to read without an Engramm licence; the padlock in the list only announces it.
- Installing an engine needs a system Python 3.10 or newer.
- Remove asks nothing. The counter of pending PDFs is how you see what is silently unindexed.