-
Ensure you have nvidia-container-toolkit installed. Refer here for install instructions: https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html
-
Ensure you have docker buildx installed. Refer here for install instructions: https://github.com/docker/buildx?tab=readme-ov-file#installing (Should be pre installed if you have docker desktop)
-
Configure docker to allow GPU pass through using nvidia-container-toolkit:
sudo nvidia-ctk runtime configure --runtime=docker sudo systemctl restart docker -
Figure out what gpu architecture to target: run the following
nvidia-smi --query-gpu=compute_cap --format=csv,noheader
The command returns a number eg: 8.6 remove the . and change CUDA_TARGET in the docker compose file to this number.
Or Look for the GPU name in the following table, then cross reference:
GPU Compute Cap RTX 40 series (4090, 4080 etc) 89 RTX 30 series (3090, 3080, 3070 etc) 86 RTX 20 series (2080, 2070 etc) 75 GTX 16 series (1660 etc) 75 GTX 10 series (1080, 1070 etc) 61 A100 80 H100 90 -
Build the container:
docker compose build
-
Create a new folder at projet root called data. Create two folders inside: images, output. The pdfs to ocr should be in /data, not /data/pdf etc. you should have the following file structure:
Softwerk-ocr-hackathon ├── .gitignore ├── .dockerignore ├── cargo.lock ├── cargo.toml ├── docker-compose-yml ├── dockerfile ├── README.md ├── src └── data ├── images ├── output ├── <place pdfs to be ocr here> └── murderer.pdf -
Run the container:
docker compose up
-
Output of the pipeline can be found in the output folder with each page being its own markdown document. Please use a mardown reader like obsidian or VsCode to read this.
NOTE Please clean the images directory before running again. This should happen automatically but some files might remain.