I have a few hundered feet of bookshelves. I've been going through the books and downloading scans of them, if I can find them somewhere.
These are mostly in PDF format (or sometimes djvu). If they are readable and complete I'll dispose of the physical copy.
But the PDF's in particular are often very slow to nagivate because of the colouring and texture of thepage background, which can take quite a while for the PDF program to generate.
I've tried converting the background colour to white, and the scan postprocess tools available in the likes of Acrobat, but these always interfere with the text (make it blurry) and mahe the PDF's unpleasant to read.
Has anyone been through this and found a decent workflow to lighten PDF backrounds without touching the text?
Didn't find your answer? Ask the community — no account required.
II think the PDF, or DJVU, EPUB, or ______, will pretty much last forever. Maybe...
D
Don Y
Gee, isn't this fun? (not) I have about 30 ft of shelfspace that remains to be digitized (before I move on to "paper notes and receipts") after having processed some 80 "photocopier paper cartoons" of books. <frown>
Those that remain have to be "unbound" (I have a guillotine cutter that allows me to chop the bindings off without introducing the "cutting skew" that a paper cutter introduces) before scanning.
I only use PDF for those things where presentation is important as it limits the devices I can use to read the files.
The PDF *render* or the *composer*? I.e., is this a one-time activity or one that occurs every time you try to read the document?
For one that you "found online", do you have a URL to demonstrate?
Are you trying to preserve color? Greyscale? I.e., have you tried converting color to greyscale; greyscale to monochrome?
D
Don Y
I used to *cling* to "books" as if they were sacred. No marks, no dog-ears, etc.
But, after a lifetime, you end up with *so* many that they start to make demands on you -- where they are stored, how htey are stored, how they are organized, etc.
When I got my first Nook, I quickly realized that *all* of my novels were just taking up space -- instead of putting them on microSD cards and carrying the whole library with me (sadly, ereaders don't have an effective way of organizing large collections -- they seem to be tailored to folks reading the most recent crop of "Best Sellers")
Text books are more difficult as they often have illustrations and other presentation aspects that are key to understanding the content. Ditto research papers, data books, etc.
Databooks can often be DLed so that avoids that issue (except for older items that have particular interest).
Research papers are almost always available in PDFs -- even if they happen to be in 8.5x11 format.
Text books are the hardest to find, create and review.
I now use an 18" tabletPC for the items that can't easily be "reflowed" like an epub. It's a bit large/heavy -- but no worse than a similarly sized textbook.
As with every "file" that I have, I have a database that tracks its location (which medium, which folder/container, etc). Along with any duplicates that may be replicated elsewhere (e.g., projects often reference the same datasheets so those naturally appear in each "project folder")
My database lets me annotate each entry with a freeform note (to myself). Sorting based on the *presence* of such a note is often enough to find what I am interested in as its presence means the item was of sufficient interest for me to have made that annotation.
Or, look at the enclosing folder hierarchy ("Ah, this was part of Project X" and rely on personal memory to augment that structure)
PDFs are just containers. You can put images, text, binaries, multimedia, etc. into a PDF. How it *appears* is up to the person who crafts the PDF.
There are a lot of online tools to process PDFs. But, that becomes tedious when you have gigabytes of individual files that have to be "uploaded", processed and downloaded. And, you likely need to review each to ensure something hasn't been excised from the original that may have value.
E
Edward Rawde
I use databases too, but there's no way I could use a database to track the location of a book because the database would never be up to date with the book's current location.
M
Martin Brown
I trust original paper copies of books to survive.
AI companies are buying up secondhand books at a prodigious rate and destroying them to train AI ready for the singularity. Although what use "How to pass your driving test in Wales" is to something intending world domination is a mystery to me.
Image to text can work but you need to carefully proofread every page. many books are already available in online digital libraries.
The trick is to use a greedy histogram equalisation down to 4 colours which leaves enough margin to smooth the text and also lose the foxing. If they are full colour scans then split to RGB first and use only the red channel (which is usually pristine).
Smart histogram equalisation is the only way to do it. Possibly with a bit of unsharp masking first to sharpen up the edges.
D
Don Y
Physical books tend to "wander" more than files do. I.e., if "here" was a good place for a file, then it is likely *still* a good place.
You can always make a *copy* of it, elsewhere, but no need to eliminate the original (esp if you've established what amounts to a "library" to contain it)
D
Don Y
I'm not that confident. I have many paper books that have pages that have become very brittle. Others where the pages "snapped off" at the binding edge.
Just more "patterns of words" to train against.
You can store the image (monochrome, greyscale, color, etc.) AND the OCR'd text "on top of it" (invisibly) in a PDF. This increases the size of the PDF beyond what a "pure" PDF would require. But, is invaluable if you suspect the OCR process of being faulty.
You can also repeat the OCR with the imagery in the PDF as technology improves. (But, this means you want to scan -- and preserve -- at high enough resolution to enable that "post processing")
C
Cóilín Nioclásín Glostéir
Don Y snipped-for-privacy@foo.invalid wrote: |---------------------------------------------------------------------| |"[. . .] | |> I trust original paper copies of books to survive. | | | |I'm not that confident. I have many paper books that have pages that| |have become very brittle. Others where the pages "snapped off" at | |the binding edge." | |---------------------------------------------------------------------|
Dear Don Y:
Thanks for these data. Do you have any conclusions or hints for predictions of which books shall have problems? E.g. are acid-free pages better than acidic pages? Are cheap paperbacks worse than expensive first runs? (I dislike that novels' first editions are big (i.e. heavy) whereas if I wait many months I can get comfortable small paperback editions.) Are attics and garages and sheds bad storage locations? One of our libraries uses a dehumidifer - my own first dehumidifier is of almost the same model but I do not use it specifically for papers. Are plastic folder polypockets really effective at protecting papers? One end thereof is open and humidity is microscopic.
Old pages can become yellow. (S.
formatting link
fuer Kontaktdaten!)
D
Don Y
E.g., this is one of my most precious books:
formatting link
I have taken EXTREMELY good care of it.
But, the pages are of a very heavy weight -- "quality" -- to support the color plates rendered thereon. The pages are almost like index cars in terms of weight. It has a sewn binding (instead of a "perfect binding") and the pages now "snap" off at those perforations.
I will, eventually, let each page snap off and scan them on my flatbed (much better resolution and color integrity than the autofeed scanners). But, I'd like to delay the stress associated with that event...
P
Phil Hobbs
Physical media rocks, for many reasons.
Phil Hobbs
M
Martin Brown
I guess you have a much bigger variation in humidity than we do.
Even my ancient reference books that are in the attic where temperatures fluctuate a lot are still in pretty good condition half a century later. The paper yellows and there is some foxing on the oldest books.
A few of which are centuries old (but they are well cared for).
I think I have only ever had one of my own books fall apart on me.
It has come on a long way since the early days when any tables of data would end up as complete nonsense. These days AI closes the loop so that you get words of the right type (noun/adjective/verb etc) in the right places but not always the words that were on the printed page.
D
Don Y
Likely. Most of the year, it's around 10%RH. *This* time of year, closer to 70%. During storms, it obviously increases from there.
The local library ended up losing a stash of donated books -- they left them in boxes on the floor. A crack in the cement allowed
*termites* to attack them.
We don't store anything "important" in non-living spaces as the temperature and humidity swings are too large.
The one mentioned has very thick pages -- almost like glossy cardstock. When you "flip" the pages, they don't bend -- they carry high quality illustrations (essentially an up-scale comic book). If you tried to dog-ear a page, you'd likely BREAK that corner off.
That's why you "pay the price" to keep the image with the OCR'd text -- so your eyes can be the final judge.
J
JM
Just like the Morecambe and Wise routine with André Previn ...
C
chrisq
I bought a second hand document scanner. Any format up to 24 bit colour, duplex, adf. All the uk / us sizes including A3 and A4. Will get through a 1" thick wedge of A4 in a couple of minutes. Scans to various formats, one file per page, then outputs composite as pdf, or other formats. Also define custom page sizes.
File size (and load speed) does depend on resolution and format. For textbooks, grey scale or sometimes even black and white, is good enough. 600dpi if needed, but usually
300dpi.
Use Foxit reader for display, which is free and pretty fast, though they do have paid for tools as well. Ime, web browsers like Firefox make poor pdf readers and are quite slow.
Join the Discussion
Have something to add? Share your thoughts — no account required.
Didn't find your answer?
Ask the community — no account required
Report Content
You are reporting this content to the moderators. They will look at it
ASAP.