CompuServe Thread

#Suicide as an option <g>

27 messages in this thread
#62478From: John TissavaryOct 20, 1993 3:41 AM
I'm thinking of changing careers. I guess I could put all my stuff up for sale, but after all the messages I've been leaving nobody's going to want to mess with this setup. It's the SPAWN OF SATAN!!!!!!!!!!!!!!!! Seriously… I did what you recommended, and took out ALL peripherals, TSR's, etc., untill I was down to 1 Orchid Farenheit (slow but compatible with everything), my Adaptec 2742 SCSI controller (don't have a substitute), and a mouse with Microsoft drivers. I took out all memory management, and booted up 3ds. I set up a VP sequence to render, because this thing causes MEGA traffic on the network and was crashing it at every turn, and BINGO! No dice, crashed the same as always. I checked my 3ds.set, 3dsnet.set on all machines, checked Lantastic startup parameters, you name it. WHAT GIVES????!!!!?????!!!! I really, really, really need to work this problem out. I can't afford anymore down time as I'm holding on by the skin of my teeth as it is. The demo won't be ready untill December at this rate, and I'm running out of money, options. I guess the next thing for me to try would be a different SCSI controller.
#62490From: Gus GrubbaOct 20, 1993 7:36 AM
Ok. You've eliminated a good portion of possible sources. You are now down to network, hard disk controller, and motherboard. It's got to be somewhere around here. You've said that you are now having problems with the machine rendering by itself (in the message to Gary). It is not finding the MAP path somewhere down the road. You said it will render 20 to 50 frames and then it "forgets" where the MAP path is. The one thing I can think of is that the number of file handles are being exhausted by somebody not releasing them. That's the "FILES=" statement in your config.sys. I'm not sure about this because if that was the case, 3DS should return an error message saying that too many files are open when trying to load the map rather than saying it couldn't find it. Gary should be able to fill us in as to how that sort of error is handled. It could be that the function opening the MAP file just returns a generic "error" and 3DS assumes it was because the file couldn't be found. Assuming this is true, and Gary can't figure this one by tomorrow, I can write a simple IXP file (so it would be part of your Video Post list) that keeps track of the number of files being opened and who owns all handles. This should point out the culprit.
#62606From: John TissavaryOct 20, 1993 7:52 PM
The Main workstation never dies when doing a rendering to its local path, just when there's network queue and some machines start to fail. Then it's @ 30% chance of failure. I spent the ENTIRE day with QEMM and Artisoft tech support, and have optimized both to the point where there are no conflicts (none they and I could find, anyway), so I'm going to have to rule that one out. Maybe I'll be able to switch the 2742 controller for traditional 1742 for a couple of days to see if that's the problem. I'm really suspecting r3/Lantastic just can't keep up with the traffic. Not so much that Lantastic can't do it, or that R3 can't do it, but that put together, in the Video post situation (5-12 seconds per frame, .5-1.5 megs per frame loading and writing, 3nodes) it's just too much. Meshes render fine, and non net rendering is rock solid, now that I've ironed out every kink I could find. If you have further thoughts please let me know, and I'd be very happy to try the IXP if you find time to knock it out. Sure would like to know EXACTLY what's going on. Thanks mega mucho again, John Tissavary (La Luna cie)
#62511From: David SomersOct 20, 1993 10:36 AM
Have you looked at the performance of your network yet as a possible cause of the problem? Do you have access to any type of a network analyzer, hardware or software? My thought is that your heavy duty rendering may be exceeding the available bandwidth of your network cable. The result will be that stations on the network will be unable to access the cable and will time out. You will have lost your mapping, login, etc. This may seem unlikely but I run a network here that is normal running 5 to 10% of bandwidth. When I put 2 users on who were doing very heavy network trafficing, (GIS work) they brought the whole system down by tieing up the available bandwidth. I had to put them on a seperate segment in order to keep everyone else running. Your network dealer may have some type of an analyzer you can borrow. As an alternative, the folks at ADesk may be able to give you a ballpark estimate on how much traffic your rendering is generating. Then your network dealer could tell you if you are pushing the bandwidth of your cable to the point where stations will time out.
#62548From: David TaffetOct 20, 1993 1:42 PM
John – This *is* possible… Though it's not all that easy to overwhelm a 10MB/sec Ethernet. But if a CPU is bogging down enough, you can have time-outs happening behind the scenes, which *might* be returned in a Lantastic environment as "file not found" variations. I should be working OK, from what I've gleaned from your messages to everyone, but if Lantastic is not robust enough you might try to pick up a copy of NetWare 2.2, which does not require a dedicated server. We had 2.2 here and were running some fairly heavy graphics traffic on it. It bogged down, but did not die… FWIW
#62608From: John TissavaryOct 20, 1993 7:52 PM
I'll keep Novel 2.2 in mind if I can't get things squared away on this end. Thanks, John Tissavary (La Luna cie)
#62666From: John EllisOct 21, 1993 2:08 AM
David T. How many systems are you using with Netware 2.2 and how much rendering? Just curious, this network thing I've been watching with interest because I am in the process of integrating my systems for R3 but have not taken the plunge, as I have worked with networks before and I know its not exactly a walk in the park, especially when you start tieing in more than two platforms. Thanks -JE
#62696From: David TaffetOct 21, 1993 9:48 AM
John – Don't let me lead you astray, here. We used NW 2.2 where I work, where we do not have 3DS at all, and probably never will. We *did* have numerous (20+/-) Mac and PC workstations doing multi-megabyte graphics work and printing to Postscript devices, however. For a while we made the mistake of storing all our fonts out on the server, where every Mac was accessing them during printing. This generated monuMENtal traffic, along the lines described by John Tissavary. The excessive traffic slowed the server *way* down, sometimes stopping the entire network in its tracks for seconds on end, but we *never* got corrupted data and the system *never* crashed from the load. We did upgrade to 3.x and it's better than 2.2… I suggested v.2.2 only because it'd be cheaper than 3.x or 4.x, does not require a dedicated server and could do the job. 3.11 would be better, but it ain't cheap, even for a 5-user version *and* requires a dedicated server. FWIW
#62714From: John EllisOct 21, 1993 11:30 AM
David, I can see how the configuration you described could cause alot of traffic and setting up a network for rendering appears to need a dedicated server, if one is using more than two platforms. I am curious to know what the consensus on this is going to be once the dust has settled. Thanks for the clarification. -JE
#62607From: John TissavaryOct 20, 1993 7:52 PM
I do believe the network performance is part of the problem. I just don't know what to do about it short of getting Novell and a dedicated server, which I definitely can't afford right now. Thanks, John Tissavary (La Luna cie)
#62665From: John EllisOct 21, 1993 2:08 AM
PMJI David – your explanation is intruiging and makes alot of sense. I am wondering if it is the bandwidth of the cable being exceeded or the remote systems access to the server's CPU. For instance if the rendering is taking lets say 15 mins. or more on each system the systems have to be able to check in from time to time, to signal that the path is open, correct? And if all the CPU's are busy processing (as in rendering) in a 3 platform LAN do they time-out to check that the path is open? And if they are sharing "texture maps" couldn't there be a conflict, or bottle neck. This could be a significant aspect to the design of an integrated system. In other words doing it peer to peer or using a hub with a dedicated server. Any insight in to this would be most appreciated. Thanks -JE
#62707From: Oct 21, 1993 11:05 AM
>> And if they are sharing "texture maps" couldn't there be a conflict, or >> bottle neck. Remember, texture maps only get loaded once during a typical rendering (except of course for animated texture maps, which no one has mentioned yet.) We have stored all of our texture maps on the server for years, and haven't seen any major network trafice due to that after the first frame was rendered. JohnT's problem with video post is more serious, compositing will cause more of a problem. Half of our systems are on thin coax, the other half are on twisted pair. We have never noticed a difference between the two for our 16 machines, so my guess is that it has to do with the server's I/O speed. (Can't help you there, we run Sun workstations for servers.) Never had a network rendering problem after we got 3DS3 net rendering figured out.
#62716From: John EllisOct 21, 1993 11:47 AM
Greg – Compositing in Video Post is a better example of how a system might bottleneck for sure. I used texture mapping as an example because like on your system, it is a shared resource. So essentially you're saying that one needs a high-speed server when rendering on multiple platforms, even though the server is not going to be doing any rendering. Or do you have a recommendation for a three platform system? I've got 3 66's and have contemplated using a 386-40 for a server. Thanks -JE
#62790From: Oct 21, 1993 10:22 PM
People will probably find that the requirements for networking will be a dedicated server with lots of buffers and a fast hard disk.
#62799From: John TissavaryOct 21, 1993 11:35 PM
I couldn't agree more, Greg. I'm experiencing that under moderate traffic conditions (as in normal rendering) Lantastic can hold up nicely, but when the sh*t starts to fly it just gets clogged. I'm still haven't switched out the risc co-processed SCSI controller for a traditional one, so untill then I'm going on the concept that my problems are hardware related.
#62870From: David TaffetOct 22, 1993 1:05 PM
John – Are your network cards 8-bit or 16-bit? That might make a difference…
#62953From: John TissavaryOct 23, 1993 12:17 AM
They're all 16bit with more on board memory than regular NE2000's, although I don't know how much off hand. I think Lantastic just can't deal with the amount of simultaneous file requests in the Video Post test (rescaling of bitmaps) I've been using. All the error messages in the net log indicate "file xxxx*.tga not found in map path(s) or image path(s)". Looks like share problems to me, but I'm really no expert.
#62996From: Oct 23, 1993 1:24 PM
Here's an idea – change your parameters in 3DSNET.SET to only poll every 10 seconds or so. Maybe the constant squalking is what is overloading your peer-to-peer network.
#63056From: John TissavaryOct 24, 1993 12:04 AM
Oooohhhh! Sounds good! I'll try it now. Thanks, John Tissavary (La Luna cie)
#62949From: Oct 23, 1993 12:09 AM
John: >> Thanks for your and Burt's help, You're welcome! Greg Pyros
#62809From: John EllisOct 22, 1993 12:34 AM
Greg – It would seem that there would be some logic to setting a system up, ie: what is the most efficient method. For instance would the server be the fastest platform? Would modeling be done on it and render out to the other systems? As you can imagine I am trying to get a functional grasp of the situation. I can see alot of options with different pluses and minuses. Your approach seems to be a really fast server and several 66's for rendering. When you say "a dedicated server with lots of buffers…" are you refering to System RAM or cache's or "buffer statement in the config.sys". Thanks for your input. -JE
#62950From: Oct 23, 1993 12:09 AM
UNIX workstations are set up differently than DOS in the way their buffers work, but the concept is the same. You need enough file I/O buffers and also large enough ones to hold your data so as not to slow down your network. In Sun's UNIX, the easiest way is to just set your MAXUSERS value higher than normal and re-compile the kernel. This automatically increases them for you. You would have to know how to do the equivalent to this in your DOS networking scheme to achieve the same result.
#62963From: John EllisOct 23, 1993 12:44 AM
Greg – Your setup intrigues me. Are you running 3ds on a Unix workstation, or have you figured out a way to render on that platform or are you rendering to it? I'm running out of questions here, help me out <g> There is a "files" and "buffers" statement in DOS which can be increased, and the software you purchase to network with has recommendations on what it should be set at. Apparently, alot of RAM helps also as well as a fast processor. So tell me, I'm really curious to know how far you've been able to stretch the boundaries of rendering with 3DS 🙂 -JE
#62997From: Oct 23, 1993 1:24 PM
>> Are you running 3ds on a Unix workstation No, 3DS only runs on 386/DOS PCs (and Windoze, but I haven't worked with it there). All of our PCs are networked through PC-NFS to three Sun SparkStations, which from the PCs appear as only additional drives. However, with the true multitasking available on the Suns, we also use them for AutoCAD, Modems, print/plot spoolers, mail, desktop publishing, etc., etc., etc. at the same time. Yes, we spoiled ourselves, but we made the committment to have a serious 3DS animation studio years ago, and have been seriously building towards that goal for a while. I can't imagine any other way we could have gone that would have worked out so well for us. As hard as it is to have a good crystal ball in the computer world, ours worked pretty well this time!
#63028From: John EllisOct 23, 1993 7:24 PM
Greg – I mentioned in a thread to a guy interested in becoming a dealer that it was very competitive and that only a few had really made it because it required being a good businessman and visionary. I was thinking of you among a handful of others with those qualities. Your system confirms that visionary quality. -JE
#62869From: David TaffetOct 22, 1993 1:05 PM
Greg – Exactly my recommendations. For a server that mostly acts as a staging area for TGAs to be composited and/or animated maps, that CPU won't really be doing all that much – it's just a file server – but a FAST disk and plenty of buffers, maybe even disk caches, would be the needed items. There will be the same requirements as for a multiuser database server. The only difference would be in the size of the files being transmitted, so the network cards should be 16-bit, rather than 8-bit. 8-bit would probably work OK, but 16-bit cards are only about $100 or so these days, so why take the chance? FWIW
#62712From: David SomersOct 21, 1993 11:25 AM
I don't have a Lantastic LAN here, but I do run a Novell 3.11 Ethernet LAN with about 45 users. Under normal circumstances the 30 or so users that are on the system at once use maybe 20 – 30 percent of the bandwidth. This covers uses like EMAIL, Word Processing, some light database work, etc. We are getting ready to run Geographic Information software over the network. The file movement will be very heavy. As a test I put a few users up on it and watched the network traffic and I was using 50 to 60 percent of my bandwidth, more than enough to cause very noticable slowdowns. When it peaked higher than that many systems timed out. With Novell they were told they had lost the Network connection. I don't know how Lantastic would report it. On an Ethernet network each card is watching the state of the cable. If it needs to access the cable it waits for a "clear moment" and then sends out packets. If it misjudged due to a software or hardware problem you get a collision and then the two colliding units backoff for a random amount of time and then try again. The more traffic there is on the network the more likelyhood there is of collisions. If a station encounters too many collisions it will abort the attempt to send and your station might report either the network connection has been lost or file not found, depending on how network aware the software is. I don't know how Lantastic handles collisions. If the problem is chronic the solution is to split your cable into several segments with a bridge or a router. I wouldn't do this without using a protocal analyzer first to see if this was really the problem. the analyzer can help you determine if the problem is too much traffic, a faulty driver, a faulty Network card or a combination. If you suspect bandwidth utilization and collisions may be a problem but don't have an analyzer you might try taking one of the 3 machines in question off line and see if the problem still occurs with just 2 machines. As I said earlier…3 machines pumping GIS data brought me to 60% of bandwidth. ADESK could tell you if a hard core render with 3 machines could result in similar traffic. If you suspect a bad card or driver is causing excessive collisions try removing 1 of the 3 machines, work awhile, then put it on line and try again with another machine removed, then again with the third. Removing the machine with the bad card or driver should result in a more pronounced improvement than removing a fully functioning machine. A quick check of cables and connectors might also be in order. My last suggestion is based on my experience with Novell. I don't know if this applies to Lantastic and other Peer to Peer's. If a Novell Network doesn't have enough RAM in the server to handle the number of users, processes, and the number of directories and the size of the disk it complains mightily by turning into a molasses network. Acess times slow way down and can cause stations to report file not found or network connection lost. Lantastic should be able to tell you if you have adequate RAM for your system resources. Hope this helps…..Holler if you need more! Let me know if you want to talk too. I think your initial message said you had severe time limits?