Re: SSD TBW

Aug 21, 2026 Last reply: 1 month ago 72 Replies

Nope. You are mixing up ms windows (where the file extension is believed to always be right and descriptive of the contents) with Unix (which actually looks at the files contents and pays no attention to any part of the name).

One configures applications (CVS in this case) to use the extension of the file as an indication of how it should be treated.

The contents still don't tell you what the file is or how it is intended to be used. Because no OS can infer the purpose of a "stream of bytes" without having been previously informed of that.

When a file (or, the metadata for the file) claims it is a "FIGGLEBOB", then the OS can throw up its hands and claim not to know how to handle it -- even if it looks like a shell script, text file, TIFF, etc. Because the explicit type declaration has told it that it is none of those things, despite the resemblance.

Is a .AI file a "text" file -- because the contents APPEAR to be text -- even though they aren't really? Is myprogram.q source code in C, just because peeking inside LOOKS like that's the case?

No *nix-type system works that way.

Correct. It's a shortcoming of UNIX.

Search "magic numbers"

I looked at a few hits, this page is better than most, the first paragraph should be enough.

formatting link

Again, it only works if you tell the OS about your particular file type and every other machine that will ever see files of that type. That's what Apple did, ages ago.

In UNIX, you only have the file name to rely on -- there is no other metadata available to an application/system that wants to ascertain the nature of the file/object that it is being asked to process.

In CVS, "wrappers" lets you associate "file types" with specific handling options. So, I can (effectively) say:

* assume binary, as default *.c C sources *.h C headers *.BMP bitmaps *.bz BZIP archive *.ppt PowerPoint presentation *.o object, ELF *.xwd xwd(1) screen capture etc.

Then, indicate how each "type" should be handled. E.g., to check if a "bitmap" coincides with another, do a binary comparison. To show the differences between them, create a two-valued map where one value represents "agree" and another represents "differ". For C sources, create a context diff. For powerpoint presentations...

If the creator of the files (e.g., a developer) is consistent in his naming conventions, then creating such rules is (relatively) easy and ensures the file isn't manglede by the VCS (e.g., replacing CRLFs with LFs, keyword substitutions that weren't intended, etc.)

You can *manually* tag individual files -- but, then you may be faced with thousands of such tagging operations (especially when importing someone else's sources)

For example, I have bitmap called M.A.C.C. -- how should it be treated?

And, "ReadMe" matches the "default" wildcard so should I treat its contents as "binary"? Or, add rules:

ReadMe special notes Makefile make(1) rules makefile make(1) rules

This mess because UNIX doesn't let an application (that created a particular file!) *tag* that file with a specific "type".

Fine. Store that info as part of the metadata for the backup snapshot.

Yes, but again, that’s what backup snapshots are for. VCSes serve a somewhat different purpose.

Here’s a photo of someone else’s back yard. It started out similar to yours, but has undergone its own renovations. You like some of what they’ve done, and would like to selectively apply those changes to your own back yard, which has in its turn evolved somewhat in the meantime. How would you do that?

VCSes are best precisely at dealing with textual source code, and with anything that can be represented as that.

Apple had the old “FourCC” or “OSType” type/creator system decades ago, but abandoned all that in favour of file extensions.

These days, we have MIME types. And common Linux filesystems allow those to be attached to files in a similar way to those type/creator codes on the old Apple Mac systems. Except MIME types are a somewhat more extensible system, using more descriptive strings instead of cryptic four-byte codes.

Another "bolt on solution" instead of addressing the real issue.

Your VCS can't handle anything but text -- so, pretend anything that is NOT text doesn't count.

Sounds like software developers deciding what the world should be like instead of adapting to the world that exists. And, using "what is" in place of "what should be".

No. You're just limiting them to that.

P4 seems to be able to address these issues.

formatting link
"Ah, but that's PAID software; we're too cheap to use such things!"

Can he even DESCRIBE what he did? Ask Da Vinci how you could paint a Mona Lisa. I'm sure his explanation would all be textual, right? And, sufficient to enable you to reproduce a Mona Barbara...

No. Just the VCSs that you've used and just with the limitations you've set upon your work style.

Because the rest of THEIR world went to file view extensions as having significance. That doesn't mean it's the right answer.

Why does "foo.c" have to be a C source file? Why can't "foo" serve to name that thing? If you insist on knowing what sort of thing "foo" represents, then have something augment its name with that information. You don't hesitate to augment its name with its size, owner, access permissions -- who decided that THOSE were the most important bits of metadata??

Just because they are convenient to access doesn't make them important. If a file type was important and convenient to access (by design), you'd display it, too, right?

The limitation of 4 byte codes is unrealistic. 4 billion different

*types* of files? That's not enough so we have to resort to mime types expressed as text (so humans can glance at them and understand their meanings?).

Do we have 4 billion mime types? How did mime solve THAT limitation??

Don Y snipped-for-privacy@foo.invalid wrote: |----------------------------------------------------------------------------| |"Again, it only works if you tell the OS about your particular file type and| |every other machine that will ever see files of that type. That's what | |Apple did, ages ago." | |----------------------------------------------------------------------------|

Commodore-Amiga owners used to similarly promote the original operating system for Commodore Amigas when they used to belittle Microsoft DOS. (S.

formatting link
fuer Kontaktdaten!)

Treating everything as TEXT or NOT_TEXT is arrogant and self-serving.

/etc/ethers, /etc/bootptab, /etc/hosts and scads of other files would be assessed as "text files".

Yet, if I took the *contents* of those files and randomly reassigned them to the 'wrong" names, they would all fail spectacularly in their intended function.

Because they are text *representation* of radically different THINGS.

Worse, you wouldn't KNOW they were "corrupted" until they were referenced. So, you'd have a latent bug in your system, because there is nothing ensuring that /etc/ethers complies with the definition of an "ethers(5) file", etc.

And, because they are text files, one would assume they could be manipulated by a text editor (!). Thus allowing anyone to (un)intentionally corrupt them with the same delayed realization AB:CD:EF:GH:IJ:KL is not a valid MAC *(&^jsdf no such host etc.

With *typed* objects, you can create tools that preserve and enforce those type constraints -- in addition to MEANINGFULLY showing you the difference between arbitrary versions thereof.

So, X = 2 and #define TWO (2) X = TWO and X=1+1 are identical programs. They are just EXPRESSED differently.

You can ignore differences in whitespace -- because you rationalize that it has no meaning (unless quoted). So, why can't you ignore differences in expression -- of identical concepts?

If I globally replace "identifier1" with "identifier2", how many silly hits are you going to come up with to highlight a difference that doesn't exist?

diff -picknits A B

How does (*ptr).member differ from ptr->member, other than lexical form? If I systematically went through my sources and made that change, you'd show me (possibly) hundreds of differences -- that aren't REALLY differences. Simply because you fixate on ASCII symbols and their corresponding glyphs.

I.e., your tool is crippled by "cheap" assumptions -- likely put in place in case someone wants to drag out a PDP-11 with 16K of core to run the tool!

But, let's keep our feet firmly tied in the past so when we have graphical programming languages, everyone will insist on some way of representing them as unambiguous text with precise 1:1 mappings (so the spatial relationships are preserved, etc.). Anyone wanting to use things other than text can reinvent the wheel -- likely with backwards support for text for those folks still tied to that representation.

The great thing about "evolution" is it takes forever to make real progress!

That is one of the many values of version control.

Another interesting thing is that if I make a change to a file and at the same time you make a change to a file, we should be able to consolidate both changes when the files are checked in. (Not all CS do this, and there are some disadvantages, but it can be better when you have coarse granularity because there's a lot of stuff in a file.)

--scott

"The current build is failing, I get syntax errors in function X." -- programmer

"Oh, don't use function X, I haven't finished writing it." -- so-called project director

You can do this, but you can also let it use the magic numbers instead. Which is appropriate depends on your environment.

That's why we have magic numbers. Not every file has a valid magic number but if it doesn't, it will default to "data" type which is the default "I don't know." Use the "file" command on the command line and see how effective it is.

--scott

Here's something to add to your dossier, since I know you are computer guys without girlfriends:

Fed up with a girl I was working on, I basically told her to f*ck off, something I had never quite said to her before. She gave a short quip chiding me, and later I returned from an errand at AT&T to find standing water on my keyboard.

There were three categorical alerts raised, one of which was a hard drive fail.

The good news is that, even though I'd heard of such, everything came back to normal, but it surprised me because the timeline was much longer than I'd expected.

And, its going to magically be able to compare JPGs, executables, .so's, etc?

Play with something like Beyond Compare (Windows) and see how helpful file comparison can be.

See my post regarding file equivalence and how "text" -- so common in UNIX systems (esp for configuration "databases") -- means so little. Yet, folks want to bias the VCS towards supporting that over all other "file types".

So diff and merge become even more important.

"But, lets limit them to just text files... and someone else can create a tool that performs the same functionality for NON-test files (?)"

She should have used coca-cola. I worked years ago at a hospital where someone had spilled cola into an HP2626 terminal and left it over the weekend. There were large sections of the board where the traces were completely dissolved.

--scott

No, but it (at least SCCS) will know to treat different kinds of file differently.

There's no reason you couldn't have a jpeg-comparing function these days. That's the marvel of the Unix philosophy... modularity.

Well, yes, because in the Unix world most useful files can be treated that way. Remember that "version control" started out initially as "source code control."

--scott

Yet you don't see a "-jpeg" option to diff...

"Text" does nothing for you -- except let you use a text editor to maintain it.

Note that vipw(8) adds value -- it ensures all of the passwd-associated files remain in a consistent format. BECAUSE IT IS AWARE OF THAT FORMAT, even though the files still identify as "text" (no magic numbers, etc.)

Tools should augment your abilities, enhance your productivity, etc. Not constrain it or coerce it into a specific form to suit THEIR limitations.

There's no reason for the absence of a viethers tool. Or, vihosts. Or, viboottab. And just imagine all of the tools you could create to keep X configured properly!! The existence of such tools would free the implementation to use a form other than "text files".

[I have all of the configuration details for my machines in a series of relational databases. I had initially thought of backporting this into the BSD distros that I use but realized that effort would likely offend "historical purists". So, instead, I have scripts that query the RDBMS and know how to "write" the appropriate files that are then copied into each host. My *current* project doesn't rely on such a kludge and tools directly access the records in the appropriate tables. Constrraints on the tables ensure that the tools need not validate their input -- the RDBMS won't accept bad data and is "guaranteed" to maintain the integrity of the data given to it. My, what a novel idea...]

Instead, you have to "manually" enforce the rules for these with discipline and awareness *or* deliberately invoke soomething else that will check them for you.

"Use leading spaces instead of tabs" "Ensure file is terminated with a newline"

Because that's the way it has always been (yet we can freely make sweeping changes to OTHER things that are more consequential). But, we want to keep these things for just the use of The Elite, right? <rolls eyes>

As a kid, we used to use coke to clean the rust off our bicycles. Imagine what it was doing to our *teeth*!

[Apparently, kids, nowadays, have their teeth "sealed" and don't experience the sort of "repairs" of ages past.]

Regarding water...

I volunteer at a place that recycles various types of kit (electronic, medical, etc.). One of the guys brings batches of keyboards home and cleans them in his swimming pool (!).

I've rescued kit that has been left outdoors (at that facility) during rainstorms and have been surprised as to how much of it still works. Conceptually, rain water should be pure -- but, in reality, it picks up lots of crap as it passes through the atmosphere, changing its "content" and pH.

No. Read the man page, diff calls itself a tool for comparing text files.

If you want to compare jpeg files you use odiff. And you can call odiff from your source code control system if it's modular.

And it makes it much easier for human beings to maintain. If you'd like, you can dispose of the compiler and assembler and just create executables with cat > a.out but you're not going to do that because you are a human being and human beings work well with text.

Yes, and emacs has context-sensitive modes for C and for English. If you like format-specific editors you can use them. Personally they drive me up the wall and I don't like them, but it's your call as a user.

This is very much contrary to the Unix philosophy. If you like it, that is fine but it's not Unixlike. But I am not sure why you are taking this thread so far afield.

--scott

Join the Discussion

Have something to add? Share your thoughts — no account required.

Didn't find your answer?

Ask the community — no account required