#Suicide as an option <g>
27 messages in this thread
I'm thinking of changing careers. I guess I could put all my stuff up
for sale, but after all the messages I've been leaving nobody's going
to want to mess with this setup. It's the SPAWN OF
SATAN!!!!!!!!!!!!!!!!
Seriously… I did what you recommended, and took out ALL
peripherals, TSR's, etc., untill I was down to 1 Orchid Farenheit
(slow but compatible with everything), my Adaptec 2742 SCSI
controller (don't have a substitute), and a mouse with Microsoft
drivers. I took out all memory management, and booted up 3ds. I set
up a VP sequence to render, because this thing causes MEGA traffic on
the network and was crashing it at every turn, and BINGO! No dice,
crashed the same as always. I checked my 3ds.set, 3dsnet.set on all
machines, checked Lantastic startup parameters, you name it. WHAT
GIVES????!!!!?????!!!!
I really, really, really need to work this problem out. I can't
afford anymore down time as I'm holding on by the skin of my teeth as
it is. The demo won't be ready untill December at this rate, and I'm
running out of money, options. I guess the next thing for me to try
would be a different SCSI controller.
Ok. You've eliminated a good portion of possible sources. You are now
down to network, hard disk controller, and motherboard. It's got to
be somewhere around here. You've said that you are now having
problems with the machine rendering by itself (in the message to
Gary). It is not finding the MAP path somewhere down the road. You
said it will render 20 to 50 frames and then it "forgets" where the
MAP path is.
The one thing I can think of is that the number of file handles are
being exhausted by somebody not releasing them. That's the "FILES="
statement in your config.sys. I'm not sure about this because if that
was the case, 3DS should return an error message saying that too many
files are open when trying to load the map rather than saying it
couldn't find it. Gary should be able to fill us in as to how that
sort of error is handled. It could be that the function opening the
MAP file just returns a generic "error" and 3DS assumes it was
because the file couldn't be found.
Assuming this is true, and Gary can't figure this one by tomorrow, I
can write a simple IXP file (so it would be part of your Video Post
list) that keeps track of the number of files being opened and who
owns all handles. This should point out the culprit.
The Main workstation never dies when doing a rendering to its local
path, just when there's network queue and some machines start to
fail. Then it's @ 30% chance of failure.
I spent the ENTIRE day with QEMM and Artisoft tech support, and have
optimized both to the point where there are no conflicts (none they
and I could find, anyway), so I'm going to have to rule that one out.
Maybe I'll be able to switch the 2742 controller for traditional 1742
for a couple of days to see if that's the problem.
I'm really suspecting r3/Lantastic just can't keep up with the
traffic. Not so much that Lantastic can't do it, or that R3 can't do
it, but that put together, in the Video post situation (5-12 seconds
per frame, .5-1.5 megs per frame loading and writing, 3nodes) it's
just too much.
Meshes render fine, and non net rendering is rock solid, now that I've
ironed out every kink I could find.
If you have further thoughts please let me know, and I'd be very happy
to try the IXP if you find time to knock it out. Sure would like to
know EXACTLY what's going on.
Thanks mega mucho again,
John Tissavary (La Luna cie)
Have you looked at the performance of your network yet as a possible
cause of the problem? Do you have access to any type of a network
analyzer, hardware or software?
My thought is that your heavy duty rendering may be exceeding the
available bandwidth of your network cable. The result will be that
stations on the network will be unable to access the cable and will
time out. You will have lost your mapping, login, etc.
This may seem unlikely but I run a network here that is normal running
5 to 10% of bandwidth. When I put 2 users on who were doing very
heavy network trafficing, (GIS work) they brought the whole system
down by tieing up the available bandwidth. I had to put them on a
seperate segment in order to keep everyone else running.
Your network dealer may have some type of an analyzer you can borrow.
As an alternative, the folks at ADesk may be able to give you a
ballpark estimate on how much traffic your rendering is generating.
Then your network dealer could tell you if you are pushing the
bandwidth of your cable to the point where stations will time out.
John – This *is* possible… Though it's not all that easy to
overwhelm a 10MB/sec Ethernet. But if a CPU is bogging down enough,
you can have time-outs happening behind the scenes, which *might* be
returned in a Lantastic environment as "file not found" variations. I
should be working OK, from what I've gleaned from your messages to
everyone, but if Lantastic is not robust enough you might try to pick
up a copy of NetWare 2.2, which does not require a dedicated server.
We had 2.2 here and were running some fairly heavy graphics traffic
on it. It bogged down, but did not die… FWIW
I'll keep Novel 2.2 in mind if I can't get things squared away on this end.
Thanks,
John Tissavary (La Luna cie)
David T. How many systems are you using with Netware 2.2 and how much
rendering? Just curious, this network thing I've been watching with
interest because I am in the process of integrating my systems for R3
but have not taken the plunge, as I have worked with networks before
and I know its not exactly a walk in the park, especially when you
start tieing in more than two platforms. Thanks -JE
John – Don't let me lead you astray, here. We used NW 2.2 where I
work, where we do not have 3DS at all, and probably never will. We
*did* have numerous (20+/-) Mac and PC workstations doing
multi-megabyte graphics work and printing to Postscript devices,
however. For a while we made the mistake of storing all our fonts out
on the server, where every Mac was accessing them during printing.
This generated monuMENtal traffic, along the lines described by John
Tissavary. The excessive traffic slowed the server *way* down,
sometimes stopping the entire network in its tracks for seconds on
end, but we *never* got corrupted data and the system *never* crashed
from the load. We did upgrade to 3.x and it's better than 2.2… I
suggested v.2.2 only because it'd be cheaper than 3.x or 4.x, does
not require a dedicated server and could do the job. 3.11 would be
better, but it ain't cheap, even for a 5-user version *and* requires
a dedicated server. FWIW
David, I can see how the configuration you described could cause alot
of traffic and setting up a network for rendering appears to need a
dedicated server, if one is using more than two platforms. I am
curious to know what the consensus on this is going to be once the
dust has settled. Thanks for the clarification. -JE
I do believe the network performance is part of the problem. I just don't know
what to do about it short of getting Novell and a dedicated server, which I
definitely can't afford right now.
Thanks,
John Tissavary (La Luna cie)
PMJI David – your explanation is intruiging and makes alot of sense. I
am wondering if it is the bandwidth of the cable being exceeded or
the remote systems access to the server's CPU. For instance if the
rendering is taking lets say 15 mins. or more on each system the
systems have to be able to check in from time to time, to signal that
the path is open, correct? And if all the CPU's are busy processing
(as in rendering) in a 3 platform LAN do they time-out to check that
the path is open? And if they are sharing "texture maps" couldn't
there be a conflict, or bottle neck. This could be a significant
aspect to the design of an integrated system. In other words doing it
peer to peer or using a hub with a dedicated server. Any insight in
to this would be most appreciated. Thanks -JE
>> And if they are sharing "texture maps" couldn't there be a conflict, or
>> bottle neck.
Remember, texture maps only get loaded once during a typical
rendering (except of course for animated texture maps, which no one
has mentioned yet.)
We have stored all of our texture maps on the server for years, and
haven't seen any major network trafice due to that after the first
frame was rendered. JohnT's problem with video post is more serious,
compositing will cause more of a problem.
Half of our systems are on thin coax, the other half are on twisted
pair. We have never noticed a difference between the two for our 16
machines, so my guess is that it has to do with the server's I/O
speed. (Can't help you there, we run Sun workstations for servers.)
Never had a network rendering problem after we got 3DS3 net rendering
figured out.
Greg – Compositing in Video Post is a better example of how a system might
bottleneck for sure. I used texture mapping as an example because like on your
system, it is a shared resource. So essentially you're saying that one needs a
high-speed server when rendering on multiple platforms, even though the server
is not going to be doing any rendering. Or do you have a recommendation for a
three platform system? I've got 3 66's and have contemplated using a 386-40 for
a server. Thanks -JE
People will probably find that the requirements for networking will
be a dedicated server with lots of buffers and a fast hard disk.
I couldn't agree more, Greg. I'm experiencing that under moderate
traffic conditions (as in normal rendering) Lantastic can hold up
nicely, but when the sh*t starts to fly it just gets clogged. I'm
still haven't switched out the risc co-processed SCSI controller for
a traditional one, so untill then I'm going on the concept that my
problems are hardware related.
John – Are your network cards 8-bit or 16-bit? That might make a
difference…
They're all 16bit with more on board memory than regular NE2000's,
although I don't know how much off hand. I think Lantastic just
can't deal with the amount of simultaneous file requests in the Video
Post test (rescaling of bitmaps) I've been using. All the error
messages in the net log indicate "file xxxx*.tga not found in map
path(s) or image path(s)". Looks like share problems to me, but I'm
really no expert.
Here's an idea – change your parameters in 3DSNET.SET to only poll
every 10 seconds or so. Maybe the constant squalking is what is
overloading your peer-to-peer network.
Oooohhhh! Sounds good! I'll try it now.
Thanks,
John Tissavary (La Luna cie)
John:
>> Thanks for your and Burt's help,
You're welcome!
Greg Pyros
Greg – It would seem that there would be some logic to setting a
system up, ie: what is the most efficient method. For instance would
the server be the fastest platform? Would modeling be done on it and
render out to the other systems? As you can imagine I am trying to
get a functional grasp of the situation. I can see alot of options
with different pluses and minuses. Your approach seems to be a really
fast server and several 66's for rendering. When you say "a dedicated
server with lots of buffers…" are you refering to System RAM or
cache's or "buffer statement in the config.sys". Thanks for your
input. -JE
UNIX workstations are set up differently than DOS in the way their
buffers work, but the concept is the same. You need enough file I/O
buffers and also large enough ones to hold your data so as not to
slow down your network. In Sun's UNIX, the easiest way is to just
set your MAXUSERS value higher than normal and re-compile the
kernel. This automatically increases them for you.
You would have to know how to do the equivalent to this in your DOS
networking scheme to achieve the same result.
Greg – Your setup intrigues me. Are you running 3ds on a Unix
workstation, or have you figured out a way to render on that platform
or are you rendering to it? I'm running out of questions here, help
me out <g> There is a "files" and "buffers" statement in DOS which
can be increased, and the software you purchase to network with has
recommendations on what it should be set at. Apparently, alot of RAM
helps also as well as a fast processor. So tell me, I'm really
curious to know how far you've been able to stretch the boundaries of
rendering with 3DS 🙂 -JE
>> Are you running 3ds on a Unix workstation
No, 3DS only runs on 386/DOS PCs (and Windoze, but I haven't worked
with it there). All of our PCs are networked through PC-NFS to three
Sun SparkStations, which from the PCs appear as only additional
drives. However, with the true multitasking available on the Suns,
we also use them for AutoCAD, Modems, print/plot spoolers, mail,
desktop publishing, etc., etc., etc. at the same time.
Yes, we spoiled ourselves, but we made the committment to have a
serious 3DS animation studio years ago, and have been seriously
building towards that goal for a while. I can't imagine any other
way we could have gone that would have worked out so well for us. As
hard as it is to have a good crystal ball in the computer world, ours
worked pretty well this time!
Greg – I mentioned in a thread to a guy interested in becoming a
dealer that it was very competitive and that only a few had really
made it because it required being a good businessman and visionary. I
was thinking of you among a handful of others with those qualities.
Your system confirms that visionary quality. -JE
Greg – Exactly my recommendations. For a server that mostly acts as a
staging area for TGAs to be composited and/or animated maps, that CPU
won't really be doing all that much – it's just a file server – but a
FAST disk and plenty of buffers, maybe even disk caches, would be the
needed items. There will be the same requirements as for a multiuser
database server. The only difference would be in the size of the
files being transmitted, so the network cards should be 16-bit,
rather than 8-bit. 8-bit would probably work OK, but 16-bit cards are
only about $100 or so these days, so why take the chance? FWIW
I don't have a Lantastic LAN here, but I do run a Novell 3.11 Ethernet
LAN with about 45 users. Under normal circumstances the 30 or so
users that are on the system at once use maybe 20 – 30 percent of the
bandwidth. This covers uses like EMAIL, Word Processing, some light
database work, etc.
We are getting ready to run Geographic Information software over the
network. The file movement will be very heavy. As a test I put a
few users up on it and watched the network traffic and I was using 50
to 60 percent of my bandwidth, more than enough to cause very
noticable slowdowns. When it peaked higher than that many systems
timed out. With Novell they were told they had lost the Network
connection. I don't know how Lantastic would report it.
On an Ethernet network each card is watching the state of the cable.
If it needs to access the cable it waits for a "clear moment" and
then sends out packets. If it misjudged due to a software or
hardware problem you get a collision and then the two colliding units
backoff for a random amount of time and then try again. The more
traffic there is on the network the more likelyhood there is of
collisions. If a station encounters too many collisions it will
abort the attempt to send and your station might report either the
network connection has been lost or file not found, depending on how
network aware the software is. I don't know how Lantastic handles
collisions. If the problem is chronic the solution is to split your
cable into several segments with a bridge or a router. I wouldn't do
this without using a protocal analyzer first to see if this was
really the problem. the analyzer can help you determine if the
problem is too much traffic, a faulty driver, a faulty Network card
or a combination.
If you suspect bandwidth utilization and collisions may be a problem
but don't have an analyzer you might try taking one of the 3 machines
in question off line and see if the problem still occurs with just 2
machines. As I said earlier…3 machines pumping GIS data brought me
to 60% of bandwidth. ADESK could tell you if a hard core render with
3 machines could result in similar traffic. If you suspect a bad
card or driver is causing excessive collisions try removing 1 of the
3 machines, work awhile, then put it on line and try again with
another machine removed, then again with the third. Removing the
machine with the bad card or driver should result in a more
pronounced improvement than removing a fully functioning machine. A
quick check of cables and connectors might also be in order.
My last suggestion is based on my experience with Novell. I don't
know if this applies to Lantastic and other Peer to Peer's. If a
Novell Network doesn't have enough RAM in the server to handle the
number of users, processes, and the number of directories and the
size of the disk it complains mightily by turning into a molasses
network. Acess times slow way down and can cause stations to report
file not found or network connection lost. Lantastic should be able
to tell you if you have adequate RAM for your system resources.
Hope this helps…..Holler if you need more! Let me know if you want
to talk too. I think your initial message said you had severe time
limits?