Showing posts with label mixing. Show all posts
Showing posts with label mixing. Show all posts

Thursday, September 12, 2013

Just Because it's Sound Doesn't Mean it has to be Mixed

Mixing is like driving—everybody does it, it gets you from here to there, and it seems like it’s been part of the culture forever.

For recording or broadcast requirements with a limited channel count, a stereo or mono mix will usually fit the bill, but for live events, perhaps we can do better.

As a case in point, consider a talker at a lectern in a large meeting room. Conventional practice would dictate routing the talker’s microphone to two loudspeakers at the front of the room via the left and right masters, and then feeding the signal with appropriate delays to additional loudspeakers throughout the audience area. A mono mix with the lectern midway between the loudspeakers will allow people sitting on or near the center line of the room to localize the talker more or less correctly by creating a phantom center image, but for everyone else, the talker will be localized incorrectly toward the front-of-house loudspeaker nearest them.

In contrast to a left-right loudspeaker system, natural sound in space does not take two paths to each of our ears. Discounting early reflections, which are not perceived as discrete sound sources, direct sound naturally takes only a single path to each ear. A bird singing in a tree, a speaking voice, a car driving past—all these sounds emanate from single sources. It is the localization of these single sources amid innumerable other individually localized sounds, each taking a single path to each of our two ears, that makes up the three-dimensional sound field in which we live. All the sounds we hear naturally, a complex series of pressure waves, are essentially “mixed” in the air acoustically with their individual localization cues intact.

Our binaural hearing mechanism employs inter-aural differences in the time-of-arrival and intensity of different sounds to localize them in three-dimensional space—left-right, front-back, up-down. This is something we’ve been doing automatically since birth, and it leaves no confusion about who is speaking or singing; the eyes easily follow the ears. By presenting us with direct sound from two points in space via two paths to each ear, however, conventional L-R sound reinforcement techniques subvert these differential inter-aural localization cues.

On this basis, we could take an alternative approach in our meeting room and feed the talker’s mic signal to a single nearby loudspeaker, perhaps one built into the front of the lectern, thus permitting pinpoint localization of the source. A number of loudspeakers with fairly narrow horizontal dispersion, hung over the audience area and in line with the direct sound so that each covers a fairly small portion of the audience, will subtly reinforce the direct sound as long as each loudspeaker is individually delayed so that its output is indistinguishable from early reflections in the target seats.

Such a system can achieve up to 8 dB of gain throughout the audience without the delay loudspeakers being perceived as discrete sources of sound, thanks to the well known Haas- or precedence-effect. A talker or singer with strong vocal projection may not even need a single “anchor” loudspeaker at the front at all.

As an added benefit to achieving intelligibility at a more natural level, the audience will tend to be unaware that there is a sound system in operation, an important step in reaching the elusive system design goal of transparency—people simply hear the talker clearly and intelligibly at a more or less normal level. This approach, which has been dubbed “source-oriented reinforcement,” precludes the sound system from acting as a barrier separating the performer from the audience, because it merely replicates what happens naturally, and does not disembody the voice through the removal of localization cues.

Traditional amplitude-based panning, which, as noted above, works only for those seated in the sweet spot along the centre axis of the venue, is replaced in this approach by time-based localization, which has been shown to work for better than 90 per cent of the audience, no matter where they are seated. Free from constraints related to phasing and comb-filtering that are imposed by a requirement for mono-compatibility or potential down-mixing—and that are largely irrelevant to live sound reinforcement—operators are empowered to manipulate delays to achieve pin-point localization of each performer for virtually every seat in the house.

Source-oriented reinforcement has been used successfully by a growing number of theatre sound designers, event producers and even DJs over the past 15 years or so, and this is where a large matrix comes into its own. Happily, many of today’s live sound boards are suitably equipped, with delay and EQ on the matrix outputs.

The situation becomes more complex when there is more than one talker, a wandering preacher, or a stage full of actors, but fortunately, such cases can be readily addressed as long as correct delays are established from each source zone to each and every loudspeaker on a one-to-one basis.

This requires more than a console level matrix with just output delays, or even assigning variable input delays to individual mics, since it necessitates a true delay-matrix allowing multiple independent time-alignments between each individual source zone and the distributed speaker system.

One such delay matrix that I have used successfully is the TiMax2 Soundhub, which offers control of both level and delay at each crosspoint in matrixes ranging from 16 x 16 up to 64 x 64 to define unique image definitions anywhere on the stage or field of play.

The Soundhub is easily added to a house system via analog, AES digital, and any of the various audio networks currently available, with the matrix typically being fed by input-channel direct outputs, or by a combination of console sends and/or output groups, as is the practice of the Royal Shakespeare Company, among others.

A familiar looking software interface allows for easy programming as well as real-time level control and 8-band parametric EQ on the outputs. A PanSpace graphical object-based pan programming screen allows the operator to drag input icons around a set of image definitions superimposed onto a jpg of the stage, a novel and intuitive way of localizing performers or manually panning sound effects.


 
The TiMax PanSpace graphical object-based pan programming screen


For complex productions involving up to 24 performers, designers can add the TiMax Tracker, a radar-based performer-tracking system that interpolates softly between image definitions as performers move around the stage, thus affording a degree of automation that is otherwise unattainable.

Where very high SPLs are not required, reinforcement of live events may best be achieved not by mixing voices and other sounds together, but by distributing them throughout the house with the location cues that maintain their separateness, which is, after all, a fundamental contributor to intelligibility, as anyone familiar with the “cocktail party” effect will attest.

As veteran West End sound designer Gareth Fry says, “I’m quite sure that in the coming years, source-oriented reinforcement will be the most common way to do vocal reinforcement in drama.”

While mixing a large number of individual audio signals together into a few channels may be a very real requirement for radio, television, cinema, and other channel-restricted media such as consumer audio playback systems, this is certainly not the case for corporate events, houses of worship, theatre and similar staged entertainment.

It may sound like heresy, but just because it’s sound doesn’t mean it has to be mixed. With the proliferation of matrix consoles, adequate DSP, and sound design devices such as the TiMax2 Soundhub and TiMax Tracker available to the sound system designer, mixing is no longer the only way to work with live sound—let alone the best way for every occasion.

Thursday, June 20, 2013

Putting "Excellence" Into Perspective

One of the hard lessons learned while working on an episodic TV series is that there is neither the time nor the budget to make everything "perfect." It's been said that a show is never finished, just abandoned when you run out of time or money or both.

In a recent blog entry on the ProTools Expert site, editor Russ Hughes wrote, "Perfection is said to be as much a curse as it is a blessing, especially for creative types. We record, edit, mix audio or shoot, cut and grade video and often we just can’t leave it alone, or indeed be satisfied with the end results.

"Our clients often never know the lengths we go to when working on their projects; they certainly won’t pay for half the work we did in the name of perfection. It’s a difficult balance being a creative professional with a budget on the one hand and a personal desire to do the best we can on the other. It’s the little things that take a project from good to great, as I’ve already alluded to most of us seldom feel we have done that, despite our best efforts," he said. (http://bit.ly/11IioLh)

One of the areas that causes us the most grief is all those annoying, unwanted noises that plague our recordings, especially production tracks for film and TV shows.

As supervising sound editor on 66 episodes of Relic Hunter a few years back, I calibrated our edit rooms to 82 dBA SPL and worked on eliminating unwanted noise perceptible at that level, and not at a higher level, since 82 was the level set in the mix theatre at Deluxe. Of course you would hear every little glitch and tick if you were to boost the monitor pot, but the average audience won't, and we're not going for absolute perfection, just doing the appropriate job for our client, the production company, within the available time and budgetary constraints without killing ourselves.

One takeaway from this is that for picture work, calibrate your monitor level appropriate for the program type and then DON'T TOUCH IT AGAIN. That's how it's done on the mix stage. In fact, I've seen film consoles with the monitor pot removed. Seasoned mixers know when dialog is at the right level just using their ears and never looking at a meter. If you mix consistently to, say, feature film level of 85 dB SPL each and every day for as little as 4 weeks without ever changing the level, you'll soon train yourself to accurately gauge level, and you'll love the freedom this brings to the work, along with a concomitant lack of stress over little things that will never be heard in the intended listening environment.

This is also the antidote to level creep in music mixes, where the monitor level goes up as the hours stretch on, and ultimately changes the track's spectral content due to the way we perceive the amount of bass and treble at different listening levels (due to the equal-loudness contours). Try to avoid this temptation at all costs and leave the monitor pot alone. Failing that, calibrate it to something like 85 or 90 dBA SPL and remove the knob. Try it. You may like it.

Saturday, April 13, 2013

The Passing of Online

The essential distinction between offline and online is that an offline process is one of construction; an online process, one of execution. In media production, online usually follows offline, as in the case of video editing, where a product that has been laboriously constructed in an offline edit suite—perhaps over the course of days or weeks—is executed by machinery following an edit decision list (EDL) in minutes or hours in an online suite.

Since the hourly rate of a well appointed online suite is typically several orders of magnitude higher than that of a small offline studio—often equipped with not much more than a desktop computer running editing software—the distinction between online and offline has long been etched into the steely heart of many a production manager.

Applying this distinction to the field of music, you might say that playing an instrument is generally an online process, and requires the talent to perform. Constructing a musical performance using MIDI step input, for example, is an offline process, and requires a different skill set.

Before Bing Crosby teamed up with Jack Mullin back in 1947 and seized on the potential for splicing tape offline to construct complete recorded performances, recording musicians had to execute a complete work flawlessly to the end while it was being recorded direct to phonograph disc—an online process. If they made a mistake, they had to go back to the beginning, scrap the disc, and start all over again.

Likewise, dialing a phone on a traditional land line is an online process. If you realize you’ve made a mistake, you have to abort—hang up—and begin again. Dialing a cell phone, on the other hand, is an offline process. You compose the number and, if you make a mistake, you go back a step and delete the wrong input—edit it out—and input the right number. When the entire telephone number has been constructed to your liking, you go online—literally, hit the green online button—and the call is executed by the service provider.

The ability to edit is what distinguishes offline from online processes.

Sound mixing for film used to be mostly an online activity. It was common practice in the early decades of film sound for an entire 10-minute reel to be mixed in a single pass, following one or more rehearsals. With the development of pick-up and record electronics for film dubbers making punching in possible, the two- or three-person re-recording team enjoyed the ability at last to go back and fix a flawed portion of a mix—usually refining their console settings listening to the sound backwards while the dubbers rewound in real time—without causing undue delay and excessive cost to the production.

Mix automation changed all that, from the introduction of console automation systems in the 1970s to today’s digital audio workstations featuring the ability to graph not just volume and mute, but just about every conceivable control parameter. Automation has allowed the offline construction of mixes to become standard operating procedure, with the mix being subsequently executed online in a single record pass or internal bounce-to-disk.

Now this has all changed again with the introduction of offline bounce in ProTools 11. This enables freezing a mix—that is, rendering the final mix up to 150 times faster than real time, according to Avid—and has made the notion of “online” something of a quaint curiosity.

Now a mix need never be onlined at all, since we are able to render into a single final file something that doesn’t ever need to be played through, prior to the playback for quality control checking and approval, after the fact.

The notion of online vs. offline, once so central to the production process and necessitating the development of the all-important EDL, is in the process of being relegated to the status of a quaint curiosity, a byway in the development of modern studio practices and procedures. It will soon be forgotten, along with such other bygone realities as the daily tape recorder alignment ritual, analog noise reduction devices, and uniformed gas station attendants.

It brings to mind the day that I finally sold my once invincible Synclavier and 16-track Direct-to-Disk recorder—to a couple of vintage synth collectors, no less. The only things I hung onto were two blank rack panels and an AC power bar. Some things, at least, are irreplaceable. 

Saturday, March 23, 2013

To Mix or Not to Mix? That is the Question

In the Monty Python film, The Meaning of Life, there is an unforgettable scene in an upscale French restaurant featuring this exchange between John Cleese’s fawning waiter and Terry Jones’ more-than-morbidly obese patron, Mr. Creosote:

“Today we have for appetizers moules mariniers, pâté de fois gras, beluga caviar, eggs benedict, tarte de poivre—that’s leek tart—frogs’ legs amandine, or oeufs de cailles—little quails’ eggs on a bed of pureéd mushrooms. It’s very delicate, very subtle.”

“I’ll have the lot,” replies Mr. Creosote.

“A wise choice, monsieur. And now, how would you like it served—all mixed up together in a bucket?”

“Yeah . . . with the eggs on top.”

While the humour in the scene is partly visual, the Pythons’ unique stamp of taking things to the brink of the ridiculous, and then vaulting over it, contrasts the list of the individual, highly refined dishes on the menu—representing the pinnacle of classic French cuisine—and the way these “very delicate, very subtle” elements are offered to the patron in a gross, vulgar, and repulsive manner, “all mixed up together in a bucket.”

Of course, no-one would willingly order a meal this way, much less be served in this fashion by a trained professional. Yet, that is much the way sound is often presented to theatre patrons: all mixed up together, with the eggs—or rather, the voices—on top. Occasionally delicate, not often subtle.

What is at issue here is the very notion of mixing, of combining disparate elements into a single channel (center cluster), two channels (L, R), a combination of these (L, C, R) or perhaps even on very rare occasions, a surround mix of four or five channels.

While mixing a large number of individual audio signals together into a few channels may be a very real requirement for the limited channel count of broadcast radio and television, as well as channel-restricted media such as consumer audio playback systems, this is certainly not the case for theatre and other staged entertainment. Until recently, however, theatrical and similar live events have largely been mixed in much the same way as broadcasts and recorded music.

This may be attributed in part to the large overlap in the designs of traditional recording, broadcast and live consoles; schools teaching audio (i.e., “recording schools”) continue to focus on the art and techniques of the mixdown; even one of the audio industry’s leading magazines proudly heralds the practice in its name, Mix. Originating in broadcast and recording sessions involving multiple microphones, and refined in multitrack recording studios producing mono or stereo masters, mixing has become entrenched in the industry and in the minds of many who dream of working in it, to the point where it’s almost as if no other way of working with sound is even remotely conceivable.

A great many shows are presented as if the audience were listening to a gargantuan stereo system, with massive line arrays hung to the left and right of the stage. Now this might not be inappropriate for a touring band well known from its recordings or for a big, dynamic rock musical where the design calls for a larger-than-life aspect.

Even so, many of the blockbuster musicals from the past quarter century benefited greatly from the creativity of such esteemed sound designers as Olivier Award winner Mick Potter, who, in the quest for more natural sound, have opted for separate vocal and orchestra mixes, striving simultaneously for clarity in the voices and power in the orchestra. Moreover, two voice mixes are sometimes derived, with one going to a duplicate set of loudspeakers in an A-B configuration pioneered in 1988 by Martin Levan for Aspects of Love, to eliminate electrical summing of mic signals and the ensuing phase problems that arise when performers are in close proximity to each other’s microphones.

For other production styles, however, an approach based on mixing may not be the most appropriate technique for conveying the nuances of theatre—including musical theatre, where sound systems have become ubiquitous—if the purpose of sound reinforcement is to allow every performer’s voice to be heard as it would unamplified in an optimum seat.