CompuServe Messages

#sound sync.

    28-Aug-94 17:32:53
Sb: #120748-#sound sync.
Fm: M. G. BATCHELOR 71532,1214
To: Digimation – David Avgik 72320,2042
David / ALL, I've been very loosely following this thread, but since an IPAS developer such as yourself has requested ideas (unusual in itself), I guess it's time to speak up. I'm trying to type this as I think about it (in a hurry), but when I'm done – if you or anyone else here would like to go into more specific details – I'd be glad to provide them. I think the basic idea can be made more do-able, but also needs to be expanded in other related ways: First of all, there's obviously nothing new about any of this, it's been done before, & the techniques go back about 3/4 century or so – using a sound editor, mag tracks, wax pencils, stopwatches, metrenomes, bar sheets, & exposure or "dope" sheets……and these techniques & materials work pretty well. And of course anyone having created sound-synced animation without native tools available, has done all of this manually with some of the above materials, pencil & paper. The fundemental process of creating/adjusting animation to the beat and/or various accents of a music track, or for lip-syncing is well-established. Beats/accents is pretty quick & easy. Analyzing & adjusting for speech-track syncing isn't so easy. Plus, almost no notable feature animation is precisely matched to an audio sync track – ON PURPOSE. A long time ago, folks at Disney determined that, for whatever reason, the visual accent should lead the sound event, but varies with context. The fundamental GUI layout therfore is pretty easy & should mimic those materials & techniques historically used in feature cel animation. To try to attemp to support all the audio h/w AND make a phonetic breakdown from a composite audio signal is bound to be de-railing in nature wrt the project…and really, unnecessary. Unless a discrete analog or MIDI track is fed to the sampler, I imagine just extracting a reliable beat & accents would be difficult…although I know it is being done. What's REALLY needed is an all-encompassing interface template for visually relating ALL elements of an animation to one another, with the ability to easily adjust any parameter – INCLUDING a reference to audio tracks. This ALONE, would greatly enhance 3DS's animation capabilities. This should consist of a near full-screen dialog with additional pop-up dialogs, comprising an adaptation of vertically/horizontally scrolling phonetic bar sheet & beat/accent tracks at the top….followed by spline-graph controlled channels for EVERY animatable parameter, including VP stuff if possible…..along with text notation areas where appropriate & even thumbnail w/f's or renderings of the scene/item at that frame coulmn entry. As I said, each animatable parameter should be displayed as & controlled by a spline curve. I suppose Bezier would be good as it's consistent with 3DS, the curve passes through the control points, & it can be easily shaped interactively with familiar Bezier adjustments. NOW, this *could* have the first (top) channel read & display a sampled composite or discrete waveform over time, which it would read from the output of the sampling h/w & s/w. – if that would be helpful to some. Then followed by the phonetic breakdown bar sheet channel which you can analyze & create manually & plug in the data (text). Followed by beat/accent track that again, you analyze manually and perhaps can be computer-assisted here with consistent rhythms & accents. Of course, if it's not too much to ask some h/w / s/w combo to analyze speech/music tracks for breakdown and insertion here – GREAT – incorporate it. Or design hooks for reading this data at some future point in time with an internal or external program. Augmenting the interactive dialog, should be a full complement of timeline/graph editing functions for inserting, moving, sliding, cutting, pasting, adding, deleting single or ranges of frames….as well as controls relating to the control curves. Also, the main dialog should contain a preview viewport with at least a wireframe display capable of being driven manually, or locked to a set frame/field rate. And there should be pointers to dynamically point to the various channel displays over time – as the preview is running under computer or manual control. To top it all off, (and I know little about modern sound cards), if possible – let the s/w sync with & trigger the audio playback in real-time or proportionally-scaled near-real-time to accompany the preview. The PV can be pre-calculated of course, if necessary. If this *IS* possible, don't worry about supporting all of the sound hardware….pick a good one & the user will just have to get it to use the system. If you or anyone can pull off this control/display system, I don't think too many folks will squabble over having to replace their card. And when considering the above description, especially excluding the options, I really don't think this would be terribly difficult to code. That's about it, without getting into all the fine details & specifics. Regards, BILL