Re: SSD TBW

Aug 21, 2026 Last reply: 1 month ago 72 Replies
Jump to replies (72)

Lawrence D’Oliveiro snipped-for-privacy@NZ.invalid wrote: |----------------------------------------------------------------------| |"I found out that a common measure of the projected life of an SSD is | |being given in units of “terabytes written” (TBW). This will typically| |be much larger than the capacity of the drive itself, because it | |includes the sum total of all write operations on the drive, including| |deletion of data. | | | |For example, a 1-terabyte drive could have an expected figure of | |merit, by this measure, of 1000TBW -- that means it should be able to | |endure a total of a petabyte (1000 terabytes) written to it before | |showing signs of failure. | | | |But it seems to me, a 2-terabyte drive made of parts with the same | |quality, meaning each storage cell has the same expected endurance as | |before, should have a proportionately greater figure of merit, namely | |2000TBW. | | | |But the bigger drive will likely not last longer than the smaller one | |-- unless you don’t actually make use of the extra space. | | | |So why not divide the TBW by the actual capacity of the drive? Then | |you end up with a ratio of how much can be written in total, to the | |drive capacity -- call it, say, the “cumulative write ratio”. Both | |those drives would have a cumulative write ratio of 1000. | | | |That number, it seems to me, correlates better to the overall quality | |of the unit than TBW does. E.g. if you see a bigger drive with a | |smaller cumulative write ratio, you can suspect that they are cutting | |corners somewhere to keep the cost down." | |----------------------------------------------------------------------|



I do not yet have a first-hand experience of an SSD failure. Don Y complains in news:sci.electronics.design that an SSD suddenly completely failed, instead of gradual partial degradations of a hard disk. What does Lawrence D’Oliveiro mean by "before showing signs of failure."? This quotation gave me the impression that Lawrence D’Oliveiro does not expect the dramatic all-or-nothing scenario that Don Y reported.



This is a crosspost to news:sci.electronics.design (S.

formatting link
fuer Kontaktdaten!)


With a conventional hard drive, first you might see some error numbers creeping up on your SMART tool, as individual blocks start failing. There is a bad block replacement process that is transparent to the OS so all you see are bad numbers on SMART. Then you start getting media errors and then it's all over.

You can also have dramatic all-at-once errors when the interface fails. I have seen many USB SSDs where the SSD memory remained fine but the USB interface could not get to it.

I have seen lots of ways that SSDs fail. There are likely lots more than I haven't seen yet.

--scott

The firmware in spinning rust controllers has matured over decades. This is not the case with SSDs (especially early entrants to the market).

Note that the media fails in different ways and the controller has to understand these and adequately address them, in the wild.

Many "Thumb drives" have had problematic firmware (some Phison and Hynix controllers). The fix requires a firmware upgrade -- WHILE the drive is still willing to talk to its host.

Or, it suddenly becomes R/O -- better than inaccessible but only if you're not in the process of trying to write to it!

You can likely run a magnetic disk for many years (I have drives with 80K PoH) that still haven't encountered a remapped sector. But, you can exhaust the TBW limit for an SSD in a short time -- if you are ignorant of this limitation. At 200+MB/s, you can scribble 12GB in a minute -- almost a TB in an hour! The type of FLASH used, extent of overprovisioning, level of smarts in the controller, etc. all have a big impact on real-world numbers.

In the early days of WAROM (e.g., ER3400), naive implementations that had previously used BBSRAM for that nonvolatile function would wear out in minutes ("No, you DON'T want to write the changed settings back to the store each time an individual setting is changed! Keep a shadow copy in RAM and use an "impending power fail" signal to quickly stash them to the medium *IFF* (!) SOMETHING HAS CHANGED. It doesn't take long to go through

10^4 erase/write cycles!

First of all, rust isn’t magnetic.

Secondly, flash-storage firmware is indeed very complex -- most of that complexity seems to go into making the device behave like a magnetic disk drive. Details are proprietary, but certain filesystem experts have voiced the opinion that the firmware implements something resembling a log-structured filesystem.

(You know how filesystems nowadays commonly have journals? The journal is usually treated as a temporary holding place, to store transactions in progress, to allow clean recovery from crashes. But once you have a journal, you can consider the regular part of the filesystem to be redundant; what if the filesystem itself was all journal? That’s a “log-structured filesystem”.)

You could do away with most of this complexity just by using an OS-level filesystem that has wear-levelling built into its allocation algorithms. Several of these exist for the Linux kernel, and I think are deployed in embedded applications. Unfortunately I don’t think you can get regular consumer products designed for such usage, since the lowest common denominator, namely Microsoft Windows, has no capability to take advantage of them.

I had an EeePC with *two* "drives" -- a tiny one and a slightly larger one. I set the tiny one to boot a BSD kernel which then mounted the rest of the filesystem (e.g., /usr) that resided on the "larger" one.

For a while, I used this to DL the most recent "distfiles" archive (see pkgsrc) as I could attach an external magnetic disk and then note the differences (additions) present on the public web site vs. what I had accumulated to date. And, store the EeePC in a desk drawer!

[Often, older sources would disappear so rsync(1) was a bad choice]

But, the small size and the fact that it had a built-in TINY display and keyboard eventually grew tiresome. So, I replaced it with a full-size laptop -- where I could install a larger magnetic disk and not be bothered with mounting the (necessary!) filesystem present on the secondary drive. Now it resides on a shelf in the closet instead of inside the desk...

I think EeePCs aren't intended for much more than email, etc. And, running a UNIX on it (BSD in my case) still does writes to the / partition (think /var/log). Perhaps retooling the system might cut down on that (mount /var as a tmpfs?)

I presently use small (16G) thumb drives as the "boot disk" in several of my appliances -- so I can keep the drive bays free and "pure" for regular media. I've not had any problems, yet, but have played the tmpfs trick to reduce the risk.

HDDs tend to either have a problem spinning up or a problem with the head actuator assembly, when they "die dramatically". Before that time, they just accrete bad sectors on the GDL.

[I've not explored whether this can be "reset" like on a SCSI drive -- though suspect that would be A Bad Idea]

Dunno as I don't run Linux. Maybe start here:

formatting link

I rescue "recovery media" for Windows machines (typically 8GB thumb drives that have been factory marked as R/O). These tools let me convert them to usable 8G drives -- a nice size to emulate a DVD-DL. Especially as optical drives are becoming scarce in machines (or, physically incompatible with the sizes of those machines -- e.g., NUCs)

Did you forget the smiley face?

There are also significant differences between consumer, commercial and enterprise grade devices. The technology used in the FLASH (SLC, MLC, TLC, etc.), amount of over-provisioning, size of SRAM buffer, interface speed/family, etc. (E.g., I've not yet encountered a SAS or SCA SSD, not to say that anything -- besides market demand -- precludes their offerings). Picking a "sweet spot" that fits a perceived market demand then becomes the challenge.

It would require a different kind of drive as the role of its controller would be much different. Similar to installing FLASH media ("chips") in a device and taking on the task of managing that memory -- which might often be "write once" (or write rarely)

Yes it is. What you want is gamma ferric oxide. It's very fancy rust. (Admittedly rotating disks today all use plated media and not the ferric oxide of the 1970s).

--scott

There’s a way around that: you can use the “--link-dest” option to create multigenerational backups, with deduping of unchanged files. Or, if the files are too small to bother with deduping, just do the rsync to a new destination directory each time.

Would you believe, it was small enough to fit in a pants pocket -- at least on one pair of pants I had at the time ;).

I remember using it heavily at a client’s place, to SSH into the Asterisk server to watch its logs while debugging the integration of a predictive dialler with my call-management system and coordinating with the operator making the test calls.

I usually use BeyondCompare to compare/DL source and target directories. It lets me see both and get a feel for how much work needs to be done.

Also highlights differences (size/timestamp) between the two in those cases when both have a file but they obviously differ I have lots of cases of foo being renamed foo.old on my copy so the remote foo doesn't overwrite it; then foo.older, foo.oldest, etc. -- a place where versioning could be of benefit (but, some other repo might have a different notion of foo that I will have to reconcile with my instances.

You're either a "big person" or where HUGE pants! :>

A checkbook is about the largest item that I can put into a pants pocket. Phone sticks out (need a smaller phone!)

I inherited this one when a buddy's wife died. He asked me to remove all of her email and personal information and then recycle it (I do volunteer work at a place that recycles various types of kit).

Once wiped, I figured it would make a nice LITTLE "self-contained" machine to save me the hassle of dragging out a monitor, "CPU" and keyboard. And, it did that job well (doesn't take much horsepower to connect to an FTP service and transfer files to an external drive).

But, I had *seven* laptops -- 17 inch displays, etc. -- just collecting dust in the closet. (did I mention that I volunteer at a place that recycles kit??) "So, do I discard the EeePC, or one of the laptops, especially given that the laptops are more capable (and usable!) machines?"

Now, if I don't want to drag out the laptop, I've configured a NUC with the same software and use the TV in the living room as the monitor.

[This windows machine would be a bad choice to perform the mirror as differences in filesystem support would eat my lunch]

If these are text files, then that is the reason why software developers invented version control. Git is the premier VCS these days; its particular strength is reconciling different versions of a file that have forked off from a common ancestor, which sounds like your use case.

You can use it with binary files, too, if you are careful. P4 is my goto tool to control executables and other binaries.

Yes, I use CVS. I have "imported" many legacy codebases that were built with that as the VCS (some even use RCS or SCCS!). So, it's easier to recover the history as well as reconstitute different branches directly. (moving to git, svn, p4, etc. would mean I would lose the ability to easily move back through the revision history to reconstitute particular historical branches as that new VCS would not have "seen" those versions as they existed at that time.)

The fact that their authors had opted to reuse the same filename (in this case) for different versions was something I was stuck with.

Giving them ".old", ".older", etc. let me see which "version"

*this* repository considered to be authoritative ("Ah, foo has the same size as my foo.oldest so no need to download it, even though it obviously differs from *my* foo")

Add to this the fact that pkgsrc will try to verify the size and hash of ONE of those foo's that *it* thinks to be authoritative -- though you don't know which the current version of pkgsrc thinks is the real one until you try to build whatever references it.

[IIRC, there are tens of thousands of files in the repository and I don't build everything so no idea when/if I will ever try to resolve the "foo ambiguity"] <shrug> You can control YOUR system but can't do much about how someone else wants to control theirs.
[]

Ironicall the one I use the keyboard has failed and I prefer a "proper" mouse; so I do have to lug the extras - a proper keyboard is so much nicer to use.

Bit late to this thread, but as ssd started to become more affordable, started replacing all the mech drives with ssd. Had a couple of failures, but never buy new. Run zfs file system, which allows recovery from one or even two single disk failures, depending on initial setup. For root drives, run zfs mirrored root, which is an install option with FreeBSD. Much faster than spinning drives, lower power consumption, and worth it, just for those reasons alone.

Some systems here have been upgraded with sas interface controllers. That allows the use of ex corporate ssds, which tend to have a much better spec than some of the consumer quality drives. Some of the Samsung sas drives, for example, guarantee a full disk write and read every day for five years or more, without failure. That's effectively 100% reliable, under the lightly loaded conditions here.

Ebay is a great source, and purchases here include a batch of 17 sas 3840 Gb drives, marked as end of life, but smartcontrol shows 9% wear, which is irrelevant here. Several in a zfs pool for a few years, without a single failure. Another was a disk array with 24 x 800Gb Samsumg ssd. 520 byte sector size, but trivial to reformat to

512. 120 ukp delivered. More storage than i'll ever need, most likely.

Sata drives not too bad either, a couple of failures over the years,, but nothing serious. The real problem is when such drives use a sata to usb converter, where all bets are off worst case, and can even end up completely bricking the drive. Still, they get cheaper all the time.

The worst culprits are ssd memory sticks, some of which wont even take a large file, ~5Gb write in one go, without failure.

Chris

ChrisQ snipped-for-privacy@GfSys.co.UK> wrote: |-------------------------------------------------------| |"Bit late to this thread[. . .] | |[. . .] | | | |[. . .]" | |-------------------------------------------------------|

Dear ChrisQ,

Welcome anyway and thanks for those many good data! I often follow up after many years.

|-------------------------------------------------------| |"[. . .] | |Several in a zfs pool for a few years, without a single| |failure. [. . .] | |[. . .] | |-------------------------------------------------------|

OpenSolaris with ZFS on a hard disk failed on me.

|-------------------------------------------------------| |"Sata drives not too bad either, a couple of failures | |over the years,, but nothing serious. The real problem | |is when such drives use a sata to usb converter, where | |all bets are off worst case, and can even end up | |completely bricking the drive. [. . .] | |[. . .] | | | |The worst culprits are ssd memory sticks[. . .] | |[. . .]" | |-------------------------------------------------------|

USB hard disks and USB flash sticks are bad.

|-------------------------------------------------------| |"Still, they get cheaper | |all the time." | |-------------------------------------------------------|

Cheap crap continues to be crap. (S.

formatting link
fuer Kontaktdaten!)

It’s not about being “careful”, it’s about the usefulness of applying version control to such files.

With text files, you can use the “diff” command to narrow down exactly the parts that have changed. And the output of “diff” can be passed to “patch” to apply those changes to another copy of the original file.

And here’s the fun part: you can use “patch” to apply *diffs from multiple sources to the same file*. Yes, there are occasions when this will lead to conflicts where different patches affect the same text lines, but the rest of the time, it works fine. This is the key to open-source collaborative development, being able to merge change submissions from multiple contributors.

With binary files, none of this really works.

The PostgreSQL folks went through exactly this issue. That didn’t stop them migrating to Git

formatting link
. It’s all a matter of planning.

Yes my Linux install runs in a tmpfs in RAM. The start of the SSD had been worn out by the time I got the EeePC second-hand, probably running the factory WinXP install. User files are saved of course, but excluding things like the Firefox disk cache, so writes are mainly for OS upgrades, which hopefully won't be enough to wear the rest of the SSD out too soon.

I keep a USB CD/DVD drive for machines like the EeePC that don't come with one - the BIOSs all seem to support them fine. I guess the firmware upgrades aren't worth digging into for me - most of my dead USB drives are just dead. There's only one that started getting filesystem corruption, but didn't go r/o.

The value of version control is to be able to recreate a particular point-in-time of <whatever> you are controlling.

You don't always have the ability to "create from source" (or from a human-readable form) an item that you want to snapshot. E.g., a PURCHASED tool or component that you are using.

I commit typefaces, executables, images, etc. to the repository so I can retrieve them in the state they existed when I was using them in a version-controlled product.

Exactly. But NOT being able to snapshot their state at a point in time is even worse. Or, having to invent another mechanism to track their state "in parallel" with the other components ("Do I have the right versions of the non-text components here to be able to restore the conditions in effect at said point in time/history?")

E.g., I have hand-drawn the icons for the "HiFi replacement" I'm making for my other half. Should I export those bitmaps as text files just to be able to preserve them AS text, instead of bitmaps? Or, just make a note that it doesn't make sense to "diff" bitmaps (though I have tools that will do that)?

Sadly, UNIX assumes file type is completely indicated by a suffix on a filename -- said suffix being under the control of the person who selected the filename.

It's a matter of human resources. Do I want to replace one VCS with another (at some amount of time, effort and risk), just to say I did so? Will I be THAT much "better off" after having made the switch than if I had continued using the system that had been in place when those things were created?

Or, would I rather spend that time working on the project at hand?

My "project-based solution" is to simply snapshot the complete system and save that as a VMDK. Then, if I want to recreate the state of the development effort at some future date, just copy the appropriate VMDK to another "working" VMDK (so the preserved VMDK is not altered by your new efforts) and move forward from that point.

It is expensive in terms of disk space, but disk space is cheap.

[I've got 64T on my ESXi server and another few hundred TB on the SAN if I am willing to access the store over the wire]

I built a lab for homeless kids to access the web, do their homework, etc. Windows would boot into a jail, allow them to access and store THEIR (personal) files on external media. Then, the jail would be discarded when they logged off (so, any changes they had made to the system were discarded)

The machines that boot from USB (thumb) drives, here, have such small systems (e.g., they easily fit on a 16G thumb drive) that I imagine anything that is pulled in off the USB drive eventually sits in the disk cache (machines have more than

100MB of RAM) so it effectively acts like a "demand loaded" tmpfs -- instead of explicitly building an MFS at boot.

I use the tools on the referenced site to convert R/O media into R/W media. No idea what version Y does that version X didn't -- I just want the drive wiped and made R/W.

Most of my "workstations/servers" have room for one or two optical drives. But, I am finding keeping optical media around is annoying -- they seem to end up scattered around instead of finding their way back to their "storage place".

Using 8G thumb drives, I can keep 50 of them in a coffee mug and just sort through them as needed -- tossing them back into the mug, when done.

[E.g., I am preparing to image a disk with CZ -- locate the thumb drive, boot, image, replace thumb drive. By contrast, there are probably half a dozen optical media on my desk -- few of them MARKED with their contents (burned from ISOs, as needed, then discarded).]

I'll likely be removing the second optical drive from my machines and replacing with a SATA/SAS dock -- so I can plug in a bare drive and not have to deal with a USB dock. I am trying to get rid of "little boxes" -- like USB docks -- as they end up looking like clutter (and, having to put them "away" when not in use is just an inconvenience).

Any backup/restore system can do that, surely.

The specific value add of version control is being able to keep track of what has changed. And more than that, to be able to pursue multiple parallel lines of development, and be able to reconcile them later. This is particularly important for collaborative development, which is how a lot of open-source projects operate these days.

Why would you need to keep track of what has changed? Imagine getting a report of a bug that has appeared in a new release, that wasn’t noticed during development; there is a command called git-bisect, that comes with Git, that helps you track down the precise commit that introduced the bug.

You should have drawn them as SVG. That gives you resolution-independent scaling, and it’s also a text-based format (actually XML).

Not quite how it works

formatting link
.

If you want to encourage others to contribute to your project, then Git is the way to go.

That is time based. You want to declare a "version" -- branch -- of the development and capture THAT.

Do you happen to recall what DATE version 27.3.1204 was in effect? VCS lets you deal with "time" as a history of a product/component instead of as a mark on a calendar.

It also lets you address *pieces* of the product/component instead of having to move everything to that point in "time".

E.g., I had to revise a PCB layout for a client -- who had upgraded the layout tool (newer is better, right?). Only to discover that the new version couldn't read files that the previous one had produced.

Should he uninstall the new version, find the previous version and reinstall that? Or, just retrieve it from the VCS (along with required support files)?

The advantage of saving entire development system states is you don't have to wonder what OTHER things affect the operation of the tool that you need to access, again.

And, as they are so easy to call up -- and discard -- the impact on your *current* system's state is minimalized; once done with this retro activity, you can return to whatever you were doing before that need arose.

Here's a photo of my back yard. Here's another. What has changed (please indicate that in a human-friendly form).

Not everything is "source code". Not every commit can be traced back to specific actions: first I took the photo. then I enhanced the contrast by 23%. then I applied a filter that made all green tints a bit darker. then I cropped it to the final size.

Far easier to show before and after especially as there is no guarantee that the "user" can determine which of these steps needs to be changed -- all except the cropping are subjective.

I'm dealing with a *small* bitmapped display. So, glyphs are on-the-order of 5x7 arrays. I don't want some tool to decide how to map the antialiased pels to on/off states (which is all I can get from the display). *I* want to look at how each presents so I can evaluate the quality of the display instead of leaving that to some tool to decide.

file(1) is almost WORSE than using file extensions. As with a Mac, I (or an application) should be definitively declaring the form and content of a file, instead of leaving that to something else to deduce or declare (possibly incorrectly)

The people having access to my sources (contributing) have no qualms using the tools (compilers, preprocessors, VCS, etc.) that I've put in place. "It all works". They are free to import the portions of the codebase of interest to them, personally, and manage it with whatever tools *they* choose. I.e., ADOPT the codebase as their own. They don't have to accept the constraints that I've put on the runtime, can recode the algorithms in CLU, rewrite the comments in Esperanto, etc. It's THEIRS to do with as they want.

OTOH, if they want to keep abreast of MY efforts, then they have to adapt to the tools that *I* choose. No one is holding a gun to their heads.

Nope. You are mixing up ms windows (where the file extension is believed to always be right and descriptive of the contents) with Unix (which actually looks at the files contents and pays no attention to any part of the name).

Join the Discussion

Have something to add? Share your thoughts — no account required.

Didn't find your answer?

Ask the community — no account required