This document provides instructions on how to add additional language packs for the OCR tab in Stirling-PDF, both inside and outside of Docker.
Please update your tesseract docker volume path version from 4.00 to 5
Stirling-PDF uses OCRmyPDF which in turn uses tesseract for its text recognition. All credit goes to them for this awesome work!
Tesseract OCR supports a variety of languages. You can find additional language packs in the Tesseract GitHub repositories:
- tessdata_fast: These language packs are smaller and faster to load, but may provide lower recognition accuracy.
- tessdata: These language packs are larger and provide better recognition accuracy, but may take longer to load.
Depending on your requirements, you can choose the appropriate language pack for your use case. By default Stirling-PDF uses the tessdata_fast eng but this can be replaced.
- Download the desired language pack(s) by selecting the
.traineddata
file(s) for the language(s) you need. - Place the
.traineddata
files in the Tesseract tessdata directory:/usr/share/tesseract-ocr/5/tessdata
(Debian) or/usr/share/tesseract/tessdata
(Fedora)
If you are using Docker, you need to expose the Tesseract tessdata directory as a volume in order to use the additional language packs.
Modify your docker-compose.yml
file to include the following volume configuration:
services:
your_service_name:
image: your_docker_image_name
volumes:
- /location/of/trainingData:/usr/share/tesseract-ocr/5/tessdata
Add the following to your existing docker run command
-v /location/of/trainingData:/usr/share/tesseract-ocr/5/tessdata
If you are not using Docker, you need to install the OCR components, including the ocrmypdf app. You can see OCRmyPDF install guide
Debian based systems, install languages with this command:
sudo apt update &&\
# All languages
# sudo apt install -y 'tesseract-ocr-*'
# Find languages:
apt search tesseract-ocr-
# View installed languages:
dpkg-query -W tesseract-ocr- | sed 's/tesseract-ocr-//g'
Fedora:
# All languages
# sudo dnf install -y tesseract-langpack-*
# Find languages:
dnf search -C tesseract-langpack-
# View installed languages:
rpm -qa | grep tesseract-langpack | sed 's/tesseract-langpack-//g'