Searching through a bulk of pdf files

Darkassassin07@lemmy.ca · 3 days ago

Searching through a bulk of pdf files

tofu@lemmy.nocturnal.garden · 3 days ago

The OCR thing is it’s own task but for just searching a string in PDFs, pdfgrep is very good.

pdfgrep -ri CoolNumber69 /path/to/folder

Darkassassin07@lemmy.ca · 3 days ago

That works magnificently. I added -l so it spits out a list of files instead of listing each matching line in each file, then set it up with an alias. Now I can ssh in from my phone and search the whole collection for any string with a single command.

Thanks again!

tofu@lemmy.nocturnal.garden · 2 days ago

Glad to hear that!

Darkassassin07@lemmy.ca · 3 days ago

Interesting; that would be much simpler. I’ll give that a shot in the morning, thanks!

hoppolito@mander.xyz · 3 days ago

In case you are already using ripgrep (rg) instead of grep, there is also ripgrep-all (rga) which lets you search through a whole bunch of files like PDFs quickly. And it’s cached, so while the first indexing takes a moment any further search is lightning fast.

It supports a whole truckload of file types (pdf, odt, xlsx, tar.gz, mp4, and so on) but i mostly used it to quickly search through thousands of research papers. Takes around 5 minutes to index everything for my 4000 PDFs on the first run, then it’s smooth sailing for any further searches from there.

lIlIllIlIIIllIlIlII@lemmy.zip · 3 days ago

Try paperless-ngx. It can do OCR and has search.

hoppolito@mander.xyz · 3 days ago

For the OCR process you can probably wrangle up a simple bash pipeline with ocrmypdf and just let it run in the background once until all your PDFs have a text layer.

With that tool it should be doable with something like a simple while loop:

find . -type f -name '*.pdf' -print0 |
    while IFS= read -r -d '' file; do
        echo "Processing $file ..."
        ocrmypdf "$file" "$file"
        # ocrmypdf "$file" "${file%.pdf}_ocr.pdf"   # if you want a new file instead of overwriting the old
    done

If you need additional languages or other options you’ll have to delve a little deeper into the ocrmypdf documentation but this should be enough duct tape to just whip up a full OCR cycle.

Darkassassin07@lemmy.ca · 2 days ago

That’s a neat little tool that seems to work pretty well. Turns out the files I thought I’d need it for already have embedded OCR data, so I didn’t end up needing it. Definitely one I’ll keep in mind for the future though.

MysteriousSophon21@lemmy.world · 3 days ago

You might want to check out Docspell - it’s lighter than paperless-ngx but still handles PDF indexing and searching realy well, plus it can do basic OCR on those image-based PDFs without much setup.

Brkdncr@lemmy.world · 3 days ago

In windows you may need to add an ifilter. Adobe’s is pretty good. Then windows search will be able to search contents.

o/1MS\o ⌨️🐧 | #WeAreNatenom@norden.social · 3 days ago

@Darkassassin07 Did you already considered https://pdfgrep.org/?

With pdfgrep --ignore-case --recursive “text” **/*.pdf for example you can search a directory hierarchy of pdf files for “text”.