#HD's go bad
39 messages in this thread
I've had 3 Seagate Barracuda's (2 x 2.1gb fast and wide, 1 x 4.3gb fast) fail
this week. They are in different machines but the configuration is the same.
Asus PCI/EISA dual Pentium (1 cpu installed), tower case (2 fans) with lots of
space around the drives. The 4.3gb barracuda was attached to the PVR.
Any ideas on why these drives could fail? 3 in a week is disturbing. We do have
air conditioning for the office (8am – 6pm) and it hasn't been exceptionally
hot this week.
puzzled,
Martin
Hi Martin!
We've had the same problems with the larger (>1 MB) drives too. A mutual friend
of ours has also been having the same problems. (I mention this so's you don't
feel all alone. <G>)
I called engineers everywhere. The symantec people say that DOS isn't ready for
drives with capacities greater than 1 GIG. They suggested NT or Novell's
Server. (Which won't work for our purposes.) They also suggested smaller
partitions which is a pain I know.
Best thing I can say is contact the controller people and make sure you've got
the very latest bios upgrades for your controllers. (Personally I think they
should put the larger hard drives and controllers back in the oven until
they're done!)
…Earl
Earl,
we have quite a few 9gb drives on our SGI file server. They seem to run without
problems and are external units with their own power supply.
The ones that failed seem to be making grinding noises like the heads are
scraping on the platter (ouch!). I'm suspecting heat and/or power supply
problems.
Martin
Martin:
I had two drives in my server fail on the same day, a couple weeks back (both
Micropolis: a 1.7 Gb and a 4Gb) — both with horrible whining jet-aircraft like
noises. The 4Gb drive was a new one that I put in April to replace the
previous 4Gb Micropolis that failed in the same way. The room is plenty cool,
but I figured there wasn't enough air-flow in the case.
So, I got one of those ISA cards that has two fans on it, and also removed the
blanker plates covering a bunch of the accessable bays and fitted a
foamcore-board panel with a muffin fan mounted in it.
Hopefully this'll keep the 4 drives cool enough that I won't see this any more
— I'm getting tired of spending my weekends and nights rebuilding my
network!<g>
Dave
PS, a)Congrats on Virtuosity — saw it yesterday, and your (& Ralph's)
stuff looked great!
b)Once again we managed to miss meeting at SIGGRAPH. I saw you across
the crowded room Sunday night, but as I had just spent 15 minutes
wading back and forth from the crammed back-table to bar and back
again, I was less than keen to sally forth until once more until I
had done my beer some damage!<g> Then, we were just so tired from
our previous night's adventures, that we left fairly early (eager to
get back to our pleasant and comfortable new room. And I don't even
want to TALK about trying to get around at the Autodesk party!
David,
wow! It seems that hardrive problems are a lot less rare than I thought.
Bummer. I may replace the power supplies with PC Power and Cooling units with
some kind of alarms for heat build up. Although, I don't think heat was the
problem. We've had some really hot, sauna-like week-ends with no a/c and no
failures. This was mid-week with a/c on and office temps around 72f/20c.
Martin
Maybe the drive manufacturers ought to start naming these drives after _desert_
creatures, like the Micropolis Iguana, or the WD Cactus or thee Quantum
Bleached-Cow-Skull-In-A-Dry-River-Bed. <g>
Dave
David,
thanks for the comments on "Virtuosity". What fun, eh?
Yeah, sorry we didn't meet 🙁 I was kindof overwhelmed actually with all the
schmoozing – my vocal chords were suffering after Siggraph <g>.
Dave, you had EXACTLY the same thing happen to you as happened to me!
Same drive (was yours the AV?), same noise…and it happened to you TWICE?
Arrrgh!
I just got off the phone w/Mic tech support and their gonna FedEx me a new,
non-AV drive today. None of the tech people I talked to would admit to any
known problems with the 3243's. They sounded pretty on-the-level, too, if that
makes any difference…
So do you have one that works now?
-Alan
Alan:
Actually, no — the 3243's aren't the AV model. Yeah, I've got it back up, and
I'm duplexing the two 3243's, so I'm pretty well covered at the moment,
provided that they both don't go down at once.
I was pretty surprised when my 3 month old one went down, as I understood that
Micropolis was pretty much "the Man", as far as large drives go. I'm hoping
that my doing everything but bringing in a eunich to fan the drives w/
palm-fronds will cover it…. We'll see.
Dave
<g>
Yup, I'm gettin' my new one today.
Stay tuned…
-Alan
The problems with Micropolis drives, especially, the 4G, are legend. Micropolis
knows they have problems, but, of course, can't admit it. It stems from their
difficulty in securing sources of reliable media and heads. They're small fry
compared to many of the other drive manufactuers and wind up having to take the
left-overs in many cases.
On the other hand, their 9G drive is the best in the industry. Unlike the
smaller capacity drives, they are 5 1/4" media, don't push the technical
envelope like the small media and are consequenty much more reliable.
Regards,
BP
PMJI,
>>DOS isn't ready for drives with capacities greater than 1 GIG
) They also suggested smaller partitions which is a pain I know.
Are you serious? I have an external 4gig Micropolis that has worked fine up to
now (knocking on wood)…is this large HD problem limited to Seagates (please
say yes…)… As I think back, I do recall a couple times when it made those
"grinding noises" on boot up (double ouch) but 90% time cool and quiet. I
only have 2 (2gig) partitions.
Lydell
P.s. don't have PAR yet so no continuous read…
Yes that's what they said. I don't know how large your animations are. Ours
tend to be thousands of files and hundreds of megs. The last project we
finished had about 14,000 targa files which was re-rendered maybe 6 or 7 times.
This is why my hair is falling out, hands are swollen and eyes are red. <G>
…Earl
Earl,
>>14,000 targa files<<
What are you rendering, the 10 commandments or the Civil War?
My files are nowhere near that size…rank beginner comparatively.
But your word on the Micropolis AV 4.3gig drive has me looking for a better
backup system….just in case. 🙂
Lydell
>> >>14,000 targa files<<
What are you rendering, the 10 commandments or the Civil War? <<
FYI, by my quick calculation that's only about 6 minutes of animation.
– Dave
>>FYI, by my quick calculation that's only about 6 minutes of animation.
Not at 30 frames per _minute_ <BG>
I guess it is more like 7.78 minutes of video. See how green I am and what a
little sleep deprivation can do? <sheepishly adjusting dunce cap>
Lydell
Lydell:
>> Not at 30 frames per _minute_ <<
Talk about sssssssslllllllllooooooooowwwwww mmmmmmotttttioooooooonnnnn. How
about 30 frames per second? Just busting on you <G> :-0
– Bob Ritger
>> >>Not at 30 frames per _minute_ <<
Talk about sssssssslllllllllooooooooowwwwww mmmmmmotttttioooooooonnnnn. How
about 30 frames per second?<<
It's the latest version of CUSeeMe 🙂 (J.K.)
Lydell
I'm puzzled too, Martin.
I just installed a new Mic 4G SCSI AV 'cause the original (new in late June)
one died with a horrible scream Monday.
Micropolis replaced it free and shipped it FedEx immediately, no questions
asked, which leads me to suspect that this wasn't the first one.
Maybe some of this new HD stuff just needs to settle down a little…
-Alan
Alan,
>> I just installed a new Mic 4G SCSI AV 'cause the original (new in late June)
one died with a horrible scream Monday. <<
ouch! You too, huh? 3 in one week is really disturbing …. that's why I'm
suspicious of stuff like heat and power supplies. But nothing "seems" out of
wack ….
Martin
Alan,
I have not seen any dramatic failure of Micropolis 4.3 gig AV drives.
The turnaround from Micropolis is their standard replacement policy – not, as
far as I know, due to any known difficulties.
Kevin Krell – Computer Support Associates
Excellent to hear, Kevin. I've always been very happy with Micropolis
equipment.
I'll get the non-AV version tomorrow and I'm sure all my problems will simply
go away! <g> (knocking on wood…)
Thanks,
-Alan
Most drives (the new ones, at least) are meant to work flawlessly for at least
250,000 continuos hours (about 28 years – I guess we have to wait that long to
prove it). I have a [ridiculous] number of [large] drives running, some for as
long as 8 years, and very rarely have a problem. The two times I remember, the
problems were rather induced (shock and vibration).
Most problems I see are software problems and not physical problems. This is
the vast majority of zombie disks out there (they aren't really dead). I mean
by software problems those caused by bogus data recorded in important tracks.
Physical problems are actual media defects, head alignments, etc. These are,
for the most part, much rarer.
Media defect comes to you from factory. It's pretty damn difficult to cause a
media defect (unless you have the custom of opening disk drives with craw
bars). Head alignments are the most common but also the easiest to avoid.
Simply don't ever, ever move a hard disk while spinning. A simple knock on a
computer chassis could cause problems if you get it while the head is tracking
a disk turning 7,000 rpm. Dropping a spinning disk from as little as one inch
will most certainly trash a disk beyond repair. Overheating "might" cause
problems. If it gets too hot, the expansion might be enough to get the head too
close to the media. If that happens, you end up with bogus data recorded (not a
"dead" disk, just zombie). In the worse case, you could end up with a
misaligned head assembly (simple to fix with proper tools – certainly not at
home).
Low level formatting fixes a great majority of zombie disks. These new, large
format, SCSI disks need special tools for low level format. I would not suggest
people afraid of the DOS prompt trying it (at least until Mr.
Look-it's-me-Norton-the-ego comes up with a more pseudo-friendly user
interface, which I doubt it as they would lose their data restoring business).
By restoring just vital disk portions such as partition and boot areas, you
will, most of the time, be able to restore all data. The problem is only
hopeless when the entire disk has bogus data recorded (rare).
By the way, someone said "DOS" isn't ready for large drives. Bull Shit. For
that matter Irix ain't either as I have had more crashes with it than I had
with DOS. DOS' problem is that it has no memory protection which would allow
programs to trash system memory and therefore, could cause bogus data be
written to disk. This is indeed a problem but for 1G disks and 160 kb, single
sided floppies just as well. It has nothing to do with the size of the disk. It
just happens that you get a lot more pissed losing 1G of data than a floppy.
If a important track goes to wallah wallah land, such as track 0, there are no
tools a mortal would have that will fix it. What you might try is to check the
manufacturer BBS or FTP site and look for low level tools. Some, like NEC, DEC,
and Seagate have them available. Sometimes it is necessary to reprogram the
disk itself (download a new track information to track 0). These programs are,
for the most part, Greg pardon me, Greek. But face it, the disk is trashed,
ain't it? Unless you have data you want restored. In that case, stay away from
those tools and seek help. Just ask Greg how it is to low level format a 9gig
Seagate monster (ducking and running away…)
Gus;
I have had a problem recently with a Micropolis 2217A 1.6 Gig HD. On boot up,
it would be doing a chkdsk (DOS 5.0), and say stuff like "cannot cd to (some
directory)", and it would do this over and over and over. Then autoexec was set
up to do an "image" of the disk, and it would I guess overwrite the FAT, and my
FAT got all messed up. I used Norton a few times to fix the FAT, but finally
had to format the disk (I had backed it up luckily).
I thought it was one particular disk, but then it began happening on a second
one (replaced the first) in the same system. Now I have it sitting outside the
computer so I can change it quickly.
Could this be caused by 1. A bad controller cable? 2. A bad connection to the
controller card?
It seems to be one of these, and I purchased a new cable, but have not put it
in yet.
Thanks,
Dennis
I'm almost certain it has nothing to do with cables and/or connections. More
likely, you have some program trashing system memory. Specially if you're using
a disk cache. As far as the controller goes, if it has its own cache, there is
a chance it could be getting screwed up but it's by far a less likely scenario.
Run the system with a clean setup for a while and see if that continues to
happen.
Gus;
This will be a long one, but I believe it may be beneficial to all if
you would consider responding to this further. Maybe some lessons can
be learned and save others from a possible similar fate.
>> plus the fact you run mirror (as far as I can remember), you had
>> (note the past tense – too late now) a good chance of restoring the FAT.
While I had a good booting system, I stopped updating the mirror
(it was called image actually). So is it possible that I can retrieve
something?
The issue with the missing jumper may have been part of the cause, but
certainly not all. I now have no (usable) access to either drive. I was
using one as a backup for the other during this time.
One of them (the original) refuses to be recognized by the BIOS. The
other will not currently boot. It simply states HD failure. Press F1
to resume.
The recent history follows. I am on a separate system
(old faithful) writing this. Without it, I would be cut off from the
world because of no access to my HD's.
I needed to restore some data (.3DS file) from tape to the HD. I attempted to
put it onto the "E" drive. The tape software said "operation successful".
I did a dir, and there were a large number of strange files on the
list, but the one I wanted was at the end. I believe I changed to the
C drive, and back to E, and when I looked again, all of the files were
gone except my one I had restored, and it was 0 bytes.
So I did it a second time. Same results.
So I did it a third time but installed to the C Drive. Did a dir, and
the file was there. I needed to install some additional ram for the
rendering I was going to do. So I turned off the computer. Installed
the RAM. Upon boot up, it failed to find the hard disk, and said
system files are missing. At this point, I put in a boot floppy and booted
from it. I can read the C drive by doing a DIR, but cannot access
anything on the drive. If I attempt to run anything, it gives a fail
message Abort, Retry, XXXXX, Quit.
It also says the system files are missing. From the A drive, I typed
"sys c:" and got an error message of "cannot put system on this type
of drive or something to that effect.
I did an Fdisk, and found that instead of reporting a DOS disk, it had
some strange characters in that column. I think it was 4 little
square boxes and four hearts alternating. I think I jabbed it in the
heart or something. 😉
Can I use the "image" to get back a week in time? If so, how?
>> I'm almost certain it has nothing to do with cables and/or
>> connections. More likely, you have some program trashing system memory.
I definitely agree with you. I had not had a problem with this system
until I had NT on it. I don't recall if I had problems with NT, or
just wanted to uninstall it, but either way, I believe NT has
something to do with this.
In that direction, I attempted to use PCSHELL to go in and view and
possible undeleted some NT hidden file(s), and see what I could about my
SYSTEM files that were reported missing.
PCSHELL cannot get started. Probably because of the Report By FDISK
that it is not a DOS disk. I was able to make a copy of the
autoexec.bat, and config.sys file to a floppy.
Is there any hope of getting back to where I was without formatting the
HD at this point?
I will await any possible response. Thanks!, …….Dennis
It seems like your partition table and boot sectors are gone. The partition
table defines what type of file system is installed on a particular partition.
That's why you get FDISK telling you that's not a DOS disk. The boot sector
defines what type of hard disk you have (how many sectors per cylinder, sector
sizes, root directory size, etc.) Mirror (or image) won't do you much good by
now. It saves the state of the FAT (which I can't tell if it's bad or not).
You need to first rebuild the partition record. It has to be done by hand using
some disk editor. You will then need a new boot record. If you installed NT, it
made a copy of it for you. You should have a file in your root directory called
BOOTSECT.DOS which you would just copy to sector 0. You can use Norton stuff
for all this making sure you only copy a single sector ("from" and "to" must be
the same sector number).
If that's all that got screwed, you should be up and running. If the FAT is
also ruined, then use the FAT image to restore it. If the damage went beyond
the FAT (root directory for instance), then you will need to put the puzzle
together by hand a piece at a time.
Note that running FDISK, deleting the partitions, remaking a new DOS partition
and reformatting the disk will also fix the disk but it obviously also deletes
everything.
Gus;
Thanks for your suggestions. I could probably format the disk and start over,
but I wouldn't learn anything that way. I grabbed my Norton Disk Explorer
manual and away we go.
Thanks for your continued help. I may request clarification if I don't
understand something if you don't mind.
Dennis
Dennis,
One other thing to check – you can't have a directory name that's the
same as the device driver name, or DOS won't let you CD to it.
Kevin Krell – Computer Support Associates
Kevin;
>> you can't have a directory name that's the same as the device driver
>> name, or DOS won't let you CD to it.
Thanks for your help.
Gus' comment about
>>"More likely, you have some program trashing system memory."
messing things up made me think about this further, and……..
I hate to admit this, but I think it was my own st.. st.. stu.. stupidity.
It's not hard to say, but hard to admit in front of the forum. I had been
swapping out two identical hard disks between two different systems. One had
two hard disks, and the other only had one. I didn't have a jumper on either
one that said "I am the Master, the other is the slave", but it had no jumpers
at all which by what my documentation says is "I am the Master", and who know
what it thinks about the slave. The slave was jumpered as a slave.
So I am guessing that, as Gus said, "something was trashing memory" and I
believe that something was a missing jumper causing it to try to read both
Drives at once.
Does that sound logical/possible?
Anyway, I appreciate your help and Gus' help on this.
Now that I have the jumper in place, it seems to be working without a problem.
Thanks again,
Dennis
Thanks for the cool HD info, Gus!
-Alan
Gus,
I think in this case the drives have gone bad. They make excruciating noises
and won't pass the SCSI initialization phases, so I doubt they could be
low-level formatted.
Martin
I'm not discarding the possibility of a physical problem, but I still think
there is a good chance it's just weird data recorded in important tracks. One
of the sides of the disk, inaccessible to the user, has all drive information,
sector positioning, etc. If that goes bad, specially pointing a sector to some
place outside the disk space, the behavior is a loud knocking, rattling noise.
That's the sound of the head assembly trying to stick its face outside of the
disk case. Yes, the SCSI controller will not be able to access the disk using
"normal" methods as the information it gets back from the disk is all screwed
up (the information the disk itself tried to get from the bad sector). To fix
this it is necessary to reprogram the disk. Sending back to the manufacture
should suffice. Sometimes, it is even possible to fix by simply reprogramming
the disk to reformat itself (you can do that if you get the tools I was talking
about).
The bottom line question is: do you need the data you have there? If not, it's
probably cheaper, time wise, to just dump it. Anything longer than 3 hours to
fix will more likely be more expensive than buying a new disk…
Gus,
Obviously, you know more about HD's than me – that would be pretty easy <g>.
I don't think the data on the disks is critical. Having workstations out of
action is more critical, though. Seagate is sending out replacements, although
they claimed to be backordered on some of them.
Martin
Gus,
Thanks for the excellent HD seminar!
Alec Jason
Gus:
>> If a important track goes to wallah wallah land, such as track 0,
>> there are no tools a mortal would have that will fix it.
But you have them, right? <bg>
>> These programs are, for the most part, Greg pardon me, Greek.
I find Greek much easier than Portuguese! :^)
>> Just ask Greg how it is to low level format a 9gig Seagate monster…
Ugly! I was doing this on a Sun Workstation, and the low level format command
kept timing out before the disk was finished many hours later! I ended up
having to find the place in the source code in the SunOS UNIX kernel,
lengthening the default time, and then recompiling and praying! And all this
without Gus to hold my hand and tell me it would be all right! (Because he was
laughing too hard to stand still, as I recall!)
Greg Pyros
Gus:
I've had a total of 3 Micropolis SCSI drives go belly up in my server — one (a
4Gb model 3243) in April and two (a brand new 3243 that I installed in April
and a 2217 that's about a year and a half old) just a couple weeks ago. All
made horrendous noise that could be heard clear across the office. The
controller is an Adapptec 2742T (twin channel) EISA adapter. I beleive
failures have occurred on both channels.
With the first failure, Netware (3.12) deactivated the drive due to failures
reading or writing or something. After shutting down the server and blowing a
fan on the thing for a couple hours, I was able to get the drive back up (still
screaming) and get a new backup run (so as not to lose the 3/4 day's
activities). The second failure, was much the same, but I had just removed the
bad new 4Gb drive and replaced it, and started it formatting whilst I went home
to eat. I came back, and noticed on the tracker screen that a bunch of blocks
were being re-directed to the hot-fix zone on one of the other drives (which
was making a bit of noise, but I didn't think it sounded serious). I
dismounted the drive and tried running VREPAIR, which failed (I don't remember
exactly why). I took the whole thing down, cooled it, and tried bring it back
up — but it wouldn't even boot (this last drive held the DOS partition and the
SYS volume, of course.
The server is on an APC Smart-UPS 600, so I don't think power is the problem
(unless it's a faulty power supply), and so either it's heat, just rotten luck
w/ less-than-average-quality drives, or the controller. I'm hoping it's heat,
as the steps I've taken have been to try and eliminate that. I guess if I have
more problems, I'll have to look at the other things. Any comments?
Dave
Martin,
Are the drives staying dead? That's sort of the test. If setup or
partition info got blown away, that's one thing. If there is physical damage,
then suspect either heat or design problems. They will heat up, and ideally
should be in external cases with a fan.
Kevin Krell – Computer Support Associates
Kevin,
>> Are the drives staying dead? That's sort of the test. <<
yeah, I have a disk that has been dying and then coming back to life
sporadically for months. It's an old 760mb full height Siemens drive from when
that was considered huge.
>> suspect either heat or design problems. They will heat up, and ideally
should be in external cases with a fan. <<
I am considering putting them in external cases, but I bought big towers and
now there's little space left to put the external cases. Most of the Indigo2's
have that arrangement.
Martin