Moving from print to ebook - tweaking downloaded pdfs

Aug 16, 2026 Last reply: 17 minutes ago 4 Replies

I have a few hundered feet of bookshelves. I've been going through the books and downloading scans of them, if I can find them somewhere.



These are mostly in PDF format (or sometimes djvu). If they are readable and complete I'll dispose of the physical copy.



But the PDF's in particular are often very slow to nagivate because of the colouring and texture of thepage background, which can take quite a while for the PDF program to generate.



I've tried converting the background colour to white, and the scan postprocess tools available in the likes of Acrobat, but these always interfere with the text (make it blurry) and mahe the PDF's unpleasant to read.



Has anyone been through this and found a decent workflow to lighten PDF backrounds without touching the text?


Gee, isn't this fun? (not) I have about 30 ft of shelfspace that remains to be digitized (before I move on to "paper notes and receipts") after having processed some 80 "photocopier paper cartoons" of books. <frown>

Those that remain have to be "unbound" (I have a guillotine cutter that allows me to chop the bindings off without introducing the "cutting skew" that a paper cutter introduces) before scanning.

I only use PDF for those things where presentation is important as it limits the devices I can use to read the files.

The PDF *render* or the *composer*? I.e., is this a one-time activity or one that occurs every time you try to read the document?

For one that you "found online", do you have a URL to demonstrate?

Are you trying to preserve color? Greyscale? I.e., have you tried converting color to greyscale; greyscale to monochrome?

I used to *cling* to "books" as if they were sacred. No marks, no dog-ears, etc.

But, after a lifetime, you end up with *so* many that they start to make demands on you -- where they are stored, how htey are stored, how they are organized, etc.

When I got my first Nook, I quickly realized that *all* of my novels were just taking up space -- instead of putting them on microSD cards and carrying the whole library with me (sadly, ereaders don't have an effective way of organizing large collections -- they seem to be tailored to folks reading the most recent crop of "Best Sellers")

Text books are more difficult as they often have illustrations and other presentation aspects that are key to understanding the content. Ditto research papers, data books, etc.

Databooks can often be DLed so that avoids that issue (except for older items that have particular interest).

Research papers are almost always available in PDFs -- even if they happen to be in 8.5x11 format.

Text books are the hardest to find, create and review.

I now use an 18" tabletPC for the items that can't easily be "reflowed" like an epub. It's a bit large/heavy -- but no worse than a similarly sized textbook.

As with every "file" that I have, I have a database that tracks its location (which medium, which folder/container, etc). Along with any duplicates that may be replicated elsewhere (e.g., projects often reference the same datasheets so those naturally appear in each "project folder")

My database lets me annotate each entry with a freeform note (to myself). Sorting based on the *presence* of such a note is often enough to find what I am interested in as its presence means the item was of sufficient interest for me to have made that annotation.

Or, look at the enclosing folder hierarchy ("Ah, this was part of Project X" and rely on personal memory to augment that structure)

PDFs are just containers. You can put images, text, binaries, multimedia, etc. into a PDF. How it *appears* is up to the person who crafts the PDF.

There are a lot of online tools to process PDFs. But, that becomes tedious when you have gigabytes of individual files that have to be "uploaded", processed and downloaded. And, you likely need to review each to ensure something hasn't been excised from the original that may have value.

I use databases too, but there's no way I could use a database to track the location of a book because the database would never be up to date with the book's current location.

I trust original paper copies of books to survive.

AI companies are buying up secondhand books at a prodigious rate and destroying them to train AI ready for the singularity. Although what use "How to pass your driving test in Wales" is to something intending world domination is a mystery to me.

Image to text can work but you need to carefully proofread every page. many books are already available in online digital libraries.

The trick is to use a greedy histogram equalisation down to 4 colours which leaves enough margin to smooth the text and also lose the foxing. If they are full colour scans then split to RGB first and use only the red channel (which is usually pristine).

Smart histogram equalisation is the only way to do it. Possibly with a bit of unsharp masking first to sharpen up the edges.

Join the Discussion

Have something to add? Share your thoughts — no account required.

Didn't find your answer?

Ask the community — no account required