Showing posts with label recording. Show all posts
Showing posts with label recording. Show all posts

Thursday, September 12, 2013

Just Because it's Sound Doesn't Mean it has to be Mixed

Mixing is like driving—everybody does it, it gets you from here to there, and it seems like it’s been part of the culture forever.

For recording or broadcast requirements with a limited channel count, a stereo or mono mix will usually fit the bill, but for live events, perhaps we can do better.

As a case in point, consider a talker at a lectern in a large meeting room. Conventional practice would dictate routing the talker’s microphone to two loudspeakers at the front of the room via the left and right masters, and then feeding the signal with appropriate delays to additional loudspeakers throughout the audience area. A mono mix with the lectern midway between the loudspeakers will allow people sitting on or near the center line of the room to localize the talker more or less correctly by creating a phantom center image, but for everyone else, the talker will be localized incorrectly toward the front-of-house loudspeaker nearest them.

In contrast to a left-right loudspeaker system, natural sound in space does not take two paths to each of our ears. Discounting early reflections, which are not perceived as discrete sound sources, direct sound naturally takes only a single path to each ear. A bird singing in a tree, a speaking voice, a car driving past—all these sounds emanate from single sources. It is the localization of these single sources amid innumerable other individually localized sounds, each taking a single path to each of our two ears, that makes up the three-dimensional sound field in which we live. All the sounds we hear naturally, a complex series of pressure waves, are essentially “mixed” in the air acoustically with their individual localization cues intact.

Our binaural hearing mechanism employs inter-aural differences in the time-of-arrival and intensity of different sounds to localize them in three-dimensional space—left-right, front-back, up-down. This is something we’ve been doing automatically since birth, and it leaves no confusion about who is speaking or singing; the eyes easily follow the ears. By presenting us with direct sound from two points in space via two paths to each ear, however, conventional L-R sound reinforcement techniques subvert these differential inter-aural localization cues.

On this basis, we could take an alternative approach in our meeting room and feed the talker’s mic signal to a single nearby loudspeaker, perhaps one built into the front of the lectern, thus permitting pinpoint localization of the source. A number of loudspeakers with fairly narrow horizontal dispersion, hung over the audience area and in line with the direct sound so that each covers a fairly small portion of the audience, will subtly reinforce the direct sound as long as each loudspeaker is individually delayed so that its output is indistinguishable from early reflections in the target seats.

Such a system can achieve up to 8 dB of gain throughout the audience without the delay loudspeakers being perceived as discrete sources of sound, thanks to the well known Haas- or precedence-effect. A talker or singer with strong vocal projection may not even need a single “anchor” loudspeaker at the front at all.

As an added benefit to achieving intelligibility at a more natural level, the audience will tend to be unaware that there is a sound system in operation, an important step in reaching the elusive system design goal of transparency—people simply hear the talker clearly and intelligibly at a more or less normal level. This approach, which has been dubbed “source-oriented reinforcement,” precludes the sound system from acting as a barrier separating the performer from the audience, because it merely replicates what happens naturally, and does not disembody the voice through the removal of localization cues.

Traditional amplitude-based panning, which, as noted above, works only for those seated in the sweet spot along the centre axis of the venue, is replaced in this approach by time-based localization, which has been shown to work for better than 90 per cent of the audience, no matter where they are seated. Free from constraints related to phasing and comb-filtering that are imposed by a requirement for mono-compatibility or potential down-mixing—and that are largely irrelevant to live sound reinforcement—operators are empowered to manipulate delays to achieve pin-point localization of each performer for virtually every seat in the house.

Source-oriented reinforcement has been used successfully by a growing number of theatre sound designers, event producers and even DJs over the past 15 years or so, and this is where a large matrix comes into its own. Happily, many of today’s live sound boards are suitably equipped, with delay and EQ on the matrix outputs.

The situation becomes more complex when there is more than one talker, a wandering preacher, or a stage full of actors, but fortunately, such cases can be readily addressed as long as correct delays are established from each source zone to each and every loudspeaker on a one-to-one basis.

This requires more than a console level matrix with just output delays, or even assigning variable input delays to individual mics, since it necessitates a true delay-matrix allowing multiple independent time-alignments between each individual source zone and the distributed speaker system.

One such delay matrix that I have used successfully is the TiMax2 Soundhub, which offers control of both level and delay at each crosspoint in matrixes ranging from 16 x 16 up to 64 x 64 to define unique image definitions anywhere on the stage or field of play.

The Soundhub is easily added to a house system via analog, AES digital, and any of the various audio networks currently available, with the matrix typically being fed by input-channel direct outputs, or by a combination of console sends and/or output groups, as is the practice of the Royal Shakespeare Company, among others.

A familiar looking software interface allows for easy programming as well as real-time level control and 8-band parametric EQ on the outputs. A PanSpace graphical object-based pan programming screen allows the operator to drag input icons around a set of image definitions superimposed onto a jpg of the stage, a novel and intuitive way of localizing performers or manually panning sound effects.


 
The TiMax PanSpace graphical object-based pan programming screen


For complex productions involving up to 24 performers, designers can add the TiMax Tracker, a radar-based performer-tracking system that interpolates softly between image definitions as performers move around the stage, thus affording a degree of automation that is otherwise unattainable.

Where very high SPLs are not required, reinforcement of live events may best be achieved not by mixing voices and other sounds together, but by distributing them throughout the house with the location cues that maintain their separateness, which is, after all, a fundamental contributor to intelligibility, as anyone familiar with the “cocktail party” effect will attest.

As veteran West End sound designer Gareth Fry says, “I’m quite sure that in the coming years, source-oriented reinforcement will be the most common way to do vocal reinforcement in drama.”

While mixing a large number of individual audio signals together into a few channels may be a very real requirement for radio, television, cinema, and other channel-restricted media such as consumer audio playback systems, this is certainly not the case for corporate events, houses of worship, theatre and similar staged entertainment.

It may sound like heresy, but just because it’s sound doesn’t mean it has to be mixed. With the proliferation of matrix consoles, adequate DSP, and sound design devices such as the TiMax2 Soundhub and TiMax Tracker available to the sound system designer, mixing is no longer the only way to work with live sound—let alone the best way for every occasion.

Thursday, June 20, 2013

Putting "Excellence" Into Perspective

One of the hard lessons learned while working on an episodic TV series is that there is neither the time nor the budget to make everything "perfect." It's been said that a show is never finished, just abandoned when you run out of time or money or both.

In a recent blog entry on the ProTools Expert site, editor Russ Hughes wrote, "Perfection is said to be as much a curse as it is a blessing, especially for creative types. We record, edit, mix audio or shoot, cut and grade video and often we just can’t leave it alone, or indeed be satisfied with the end results.

"Our clients often never know the lengths we go to when working on their projects; they certainly won’t pay for half the work we did in the name of perfection. It’s a difficult balance being a creative professional with a budget on the one hand and a personal desire to do the best we can on the other. It’s the little things that take a project from good to great, as I’ve already alluded to most of us seldom feel we have done that, despite our best efforts," he said. (http://bit.ly/11IioLh)

One of the areas that causes us the most grief is all those annoying, unwanted noises that plague our recordings, especially production tracks for film and TV shows.

As supervising sound editor on 66 episodes of Relic Hunter a few years back, I calibrated our edit rooms to 82 dBA SPL and worked on eliminating unwanted noise perceptible at that level, and not at a higher level, since 82 was the level set in the mix theatre at Deluxe. Of course you would hear every little glitch and tick if you were to boost the monitor pot, but the average audience won't, and we're not going for absolute perfection, just doing the appropriate job for our client, the production company, within the available time and budgetary constraints without killing ourselves.

One takeaway from this is that for picture work, calibrate your monitor level appropriate for the program type and then DON'T TOUCH IT AGAIN. That's how it's done on the mix stage. In fact, I've seen film consoles with the monitor pot removed. Seasoned mixers know when dialog is at the right level just using their ears and never looking at a meter. If you mix consistently to, say, feature film level of 85 dB SPL each and every day for as little as 4 weeks without ever changing the level, you'll soon train yourself to accurately gauge level, and you'll love the freedom this brings to the work, along with a concomitant lack of stress over little things that will never be heard in the intended listening environment.

This is also the antidote to level creep in music mixes, where the monitor level goes up as the hours stretch on, and ultimately changes the track's spectral content due to the way we perceive the amount of bass and treble at different listening levels (due to the equal-loudness contours). Try to avoid this temptation at all costs and leave the monitor pot alone. Failing that, calibrate it to something like 85 or 90 dBA SPL and remove the knob. Try it. You may like it.

Wednesday, May 29, 2013

80 Years On: Getting it Right for Speech Reinforcement

April 27 marked the 80th anniversary of a historic milestone in the history of audio. On this date in 1933, the Philadelphia Orchestra under deputy conductor Alexander Smallens was picked up by three microphones at the Academy of Music in Philadelphia—left, center, and right of the orchestra stage—and the audio transmitted over wire lines to Constitution Hall in Washington, where it was replayed over three loudspeakers placed in similar positions to an audience of invited guests. Music director Leopold Stokowski manipulated the audio controls at the receiving end in Washington.

This historic event was reported and analyzed by audio pioneers Harvey Fletcher, J.C. Steinberg and W.B. Snow, E.C. Wente and A.L. Thuras, and others, in a collection of six papers published in January 1934 as the Symposium on Auditory Perspective by the IEEE, in Electrical Engineering. Paul Klipsch referred to the Symposium as "one of the most important papers in the field of audio."


Leopold Stowkowski and Harvey Fletcher
April 27, 1933: Leopold Stokowski at the controls with Harvey Fletcher observing
 
Prior to 1933, Fletcher had been working on what has since been termed the wall of sound. “Theoretically, there should be an infinite number of such ideal sets of microphones and sound projectors [i.e., loudspeakers] and each one should be infinitesimally small,” he wrote.

Fletcher's curtains of microphones and loudspeakers
Fletcher’s dual curtains of microphones and loudspeakers
 
Fletcher continued, “Practically, however, when the audience is at a considerable distance from the orchestra, as usually is the case, only a few of these sets are needed to give good auditory perspective; that is, to give depth and a sense of extensiveness to the source of the music.”

In this regard, Floyd Toole’s conclusions—following a career spent researching loudspeakers and listening rooms—are especially noteworthy. In his 2008 magnum opus, Sound Reproduction: Loudspeakers and Rooms, Toole noted that the “feeling of space”—apparent source width plus listener envelopment—which turns up in the research as the largest single factor in listener perceptions of “naturalness” and “pleasantness,” two general measures of quality, is increased by the use of surround loudspeakers in typical listening rooms and home theatres.

Given that these smaller spaces cannot be compared in either size or purpose to concert halls where sound is originally produced, Toole noted that in the 1933 experiment, “there was no need to capture ambient sounds, as the playback hall had its own reverberation."

Localization Errors

Recognizing that systems of as few as two and three channels were “far less ideal arrangements,” Steinberg and Snow observed that, nevertheless, “the 3-channel system was found to have an important advantage over the 2-channel system in that the shift of the virtual position for side observing positions was smaller."

In other words, for listeners away from the sweet spot along the hall’s center axis, localization errors due to shifts in the phantom images between loudspeakers were smaller in the case of a Left-Center-Right system compared with a Left-Right system.
Significantly, Fletcher did not include localization along with “depth and a sense of extensiveness” among the characteristics of "good auditory perspective.”

Regarding localization, Steinberg and Snow realized that “point-for-point correlation between pick-up stage and virtual stage positions is not obtained for 2-and 3-channel systems.” Further, they concluded that the listener “is not particularly critical of the exact apparent positions of the sounds so long as he receives a spatial impression. Consequently 2-channel reproduction of orchestral music gives good satisfaction, and the difference between it and 3-channel reproduction for music probably is less than for speech reproduction or the reproduction of sounds from moving sources.”

The 1933 experiment was intended to investigate “new possibilities for the reproduction and transmission of music,” in Fletcher’s words. Many, if not most, of the developments in multichannel sound have been motivated and financed by the film industry in the wake of Hollywood's massive financial investment in the "talkies" that single-handedly sounded the death knell of Vaudeville, and led to the conversion of a great many theatres into cinemas.

Given that the growth of the audio industry stemmed from research and development into the reproduction and transmission of sound for the burgeoning telephone, film, radio, television, and recorded music industries, it is curious that the term “theatre” continued (and still continues to this day) to be applied to the buildings and facilities of both cinemas and theatres. This reflects the confusion not only in their architecture, on which the noted theatre consultant Richard Pilbrow commented in his wonderful 2011 memoir A Theatre Project, but also in the development of their respective audio systems.

Theatre is Not Cinema: The Differing Requirements of Speech Reinforcement

Sound reinforcement was an early offshoot, eagerly adopted by demagogues and traveling salesmen alike to bend crowds to their way of thinking; yet, as Don Davis noted in 2013 in Sound System Engineering, “Even today, the most difficult systems to design, build, and operate are those used in the reinforcement of live speech. Systems that are notoriously poor at speech reinforcement often pass reinforcing music with flying colors. Mega churches find that the music reproduction and reinforcement systems are often best separated into two systems.”

The difference lies partly in the relatively low channel count of audio reproduction systems that makes localization of talkers next to impossible. Since delayed loudspeakers were widely introduced into the live sound industry in the 1970’s, they have been used almost exclusively to reinforce the main house sound system, not the performers themselves. This undoubtedly arose from the sheer magnitude of the sound pressure levels involved in the stadium rock concerts and outdoor festivals of the era.

However, in the case of, say, an opera singer, the depth, sense of extensiveness, and spatial impression that lent appeal to the reproduced sound of the symphony orchestra back in 1933, likely won’t prove satisfying in the absence of the ability to localize the sound image of the singer’s voice accurately. Perhaps this is one reason why “amplification” has become such a dirty word among opera aficionados.

In the 1980s, however, the English theatre sound designer Rick Clarke and others began to explore techniques of making sound appear to emanate from the lips of performers rather than from loudspeaker boxes. They were among a handful of pioneers who used the psychoacoustics of delay and the Haas effect “to pull the sound image into the heart of the action,” as sound designer David Collison recounted in his 2008 volume, The Sound of Theatre.

Out Board Electronics in the UK has since taken up the cause of speech sound reinforcement, with a unique delay-based input-output matrix in its TiMax2 Soundhub that enables each performer’s radio mic to be fed to dozens of loudspeakers—if necessary—arrayed throughout the house, with unique levels and delays to each loudspeaker such that more than 90 per cent of the audience is able to localize the voice back to the performer via Haas effect-based perceptual precedence, no matter where they are seated. Out Board refers to this approach as source-oriented reinforcement (SOR).

The delay matrix approach to SOR originated in the former DDR (East Germany), where in the 1970s, Gerhard Steinke, Peter Fels and Wolfgang Ahnert introduced the concept of Delta-Stereophony in an attempt to increase loudness in large auditoriums without compromising directional cues emanating from the stage. In the 1980s, Delta-Stereophony was licensed to AKG and embodied in the DSP 610 processor. While it offered only six inputs and 10 outputs, it came at the price of a small house.

Out Board started working on the concept in the early 1990s and released TiMax (now known as TiMax Classic) around the middle of the decade, progressively developing and enlarging the system up to the 64 x 64 input-output matrix, with 4,096 cross points, that characterizes the current generation, TiMax2.

The TiMax Tracker, an ingenious radar-based location system, locates performers to within six inches in any direction, so that the system can interpolate softly between pre-established location image definitions in the Soundhub for up to 24 performers simultaneously. The audience is thereby enabled to localize performers’ voices accurately as they move around the stage, or up and down on risers, thus addressing the deficiency of conventional systems regarding the localization of both speech and moving sound sources.

Source-Oriented Reinforcement

Out Board director Dave Haydon put it this way: “First thing to know about source-oriented reinforcement is that it’s not panning. Audio localization created using SOR makes the amplified sound actually appear to come from where the performers are on stage. With panning, the sound usually appears to come from the speakers, but biased to relate roughly to a performer’s position on stage. Most of us are also aware that level panning only really works for people sitting near the center line of the audience. In general, anybody sitting much off this center line will mostly perceive the sound to come from whichever stereo speaker channel they’re nearest to.

“This happens because our ear-brain combo localizes to the sound we hear first, not necessarily the loudest. We are all programmed to do this as part of our primitive survival mechanisms, and we all do it within similar parameters. We will localize even to a 1 ms early arrival, all the way up to about 25 ms, then our brain stops integrating the two arrivals and separates them out into an echo. Between 1 ms and about 10 ms arrival time differences, there will be varying coloration caused by phasing artifacts.

“This localization effect, called precedence or Haas Effect after the scientist who discovered it, works within a 6-8 dB level window. This means the first arrival can be up to 6-8 dB quieter than the second arrival and we’ll still localize to it. This is handy as it means we can actively apply this localization effect and at the same time achieve useful amplification.

“If we don’t control these different arrivals they will control us. All the various natural delay offsets between the loudspeakers, performers and the different seat positions cause widely different panoramic perceptions across the audience. You only to have to move 13 inches to create a differential delay of 1 ms, causing significant image shift. Pan pots just controlling level can't fix this for more than a few audience members near the center. You need to manage delays, and ideally control them differentially between every mic and every speaker, which requires a delay-matrix and a little cunning, coupled with a fairly simple understanding of the relevant physics and biology,” Haydon said.

Into the Mainstream

More and more theatres are adopting this approach, including New York’s City Center and the UK’s Royal Shakespeare Company. A number of Raymond Gubbay productions of opera-in-the-round at the notoriously difficult Royal Albert Hall—including Aida, Tosca, The King and I, La Bohème and Madam Butterfly—as well as Carmen at the O2 Arena, have benefited from source oriented reinforcement, as have recent productions of Les Miserables, Jesus Christ Superstar, Into the Woods, Beggar’s Opera, Marie Antoinette, Andromache, Tanz de Vampire, Lord of the Flies, Fela!, and many others at venues around the world.

Veteran West End sound designer Gareth Fry employed the technique earlier this year at the Barbican Theatre for The Master and Margarita, to make it possible for all audience members to continuously localize to the actors’ voices as they moved around the Barbican’s very wide stage. He noted that, in the three-hour show with a number of parallel story threads, this helped greatly with intelligibility to ensure the audience’s total immersion in the show’s complex plot lines.

Based on the experience, Fry said, “I’m quite sure that in the coming years, SOR will be the most common way to do vocal reinforcement in drama.”

As we mark the 80th anniversary of that historic first live stereo transmission, it’s worth noting that, in spite of the proliferation of surround formats for sound reproduction that has to date culminated in the cinematic marvel of 64-channel Dolby Atmos, we are only now getting onto the right track with regard to speech reinforcement.

It’s about time.

(photo source: http://www.stokowski.org) 

Saturday, April 13, 2013

The Passing of Online

The essential distinction between offline and online is that an offline process is one of construction; an online process, one of execution. In media production, online usually follows offline, as in the case of video editing, where a product that has been laboriously constructed in an offline edit suite—perhaps over the course of days or weeks—is executed by machinery following an edit decision list (EDL) in minutes or hours in an online suite.

Since the hourly rate of a well appointed online suite is typically several orders of magnitude higher than that of a small offline studio—often equipped with not much more than a desktop computer running editing software—the distinction between online and offline has long been etched into the steely heart of many a production manager.

Applying this distinction to the field of music, you might say that playing an instrument is generally an online process, and requires the talent to perform. Constructing a musical performance using MIDI step input, for example, is an offline process, and requires a different skill set.

Before Bing Crosby teamed up with Jack Mullin back in 1947 and seized on the potential for splicing tape offline to construct complete recorded performances, recording musicians had to execute a complete work flawlessly to the end while it was being recorded direct to phonograph disc—an online process. If they made a mistake, they had to go back to the beginning, scrap the disc, and start all over again.

Likewise, dialing a phone on a traditional land line is an online process. If you realize you’ve made a mistake, you have to abort—hang up—and begin again. Dialing a cell phone, on the other hand, is an offline process. You compose the number and, if you make a mistake, you go back a step and delete the wrong input—edit it out—and input the right number. When the entire telephone number has been constructed to your liking, you go online—literally, hit the green online button—and the call is executed by the service provider.

The ability to edit is what distinguishes offline from online processes.

Sound mixing for film used to be mostly an online activity. It was common practice in the early decades of film sound for an entire 10-minute reel to be mixed in a single pass, following one or more rehearsals. With the development of pick-up and record electronics for film dubbers making punching in possible, the two- or three-person re-recording team enjoyed the ability at last to go back and fix a flawed portion of a mix—usually refining their console settings listening to the sound backwards while the dubbers rewound in real time—without causing undue delay and excessive cost to the production.

Mix automation changed all that, from the introduction of console automation systems in the 1970s to today’s digital audio workstations featuring the ability to graph not just volume and mute, but just about every conceivable control parameter. Automation has allowed the offline construction of mixes to become standard operating procedure, with the mix being subsequently executed online in a single record pass or internal bounce-to-disk.

Now this has all changed again with the introduction of offline bounce in ProTools 11. This enables freezing a mix—that is, rendering the final mix up to 150 times faster than real time, according to Avid—and has made the notion of “online” something of a quaint curiosity.

Now a mix need never be onlined at all, since we are able to render into a single final file something that doesn’t ever need to be played through, prior to the playback for quality control checking and approval, after the fact.

The notion of online vs. offline, once so central to the production process and necessitating the development of the all-important EDL, is in the process of being relegated to the status of a quaint curiosity, a byway in the development of modern studio practices and procedures. It will soon be forgotten, along with such other bygone realities as the daily tape recorder alignment ritual, analog noise reduction devices, and uniformed gas station attendants.

It brings to mind the day that I finally sold my once invincible Synclavier and 16-track Direct-to-Disk recorder—to a couple of vintage synth collectors, no less. The only things I hung onto were two blank rack panels and an AC power bar. Some things, at least, are irreplaceable. 

Wednesday, February 29, 2012

Leap year and drop-frame time code are conceptually the same

For those in the media production industries, February 29th is a good day to revisit drop-frame SMPTE time code, because both leap year and drop-frame time code came into being for the sole purpose of reconciling two different time bases on which we do things with mundane regularity.

Take the calendar first: our calendar simply charts the sequence of the individual days that comprise a single year. The day is based, of course, on a single rotation of the earth on its axis, whereas the year is based on a single revolution of the earth around the sun. Rotation and revolution are the two different time bases on which our calendar is constructed.

Since it takes about 365.25 days for the earth to revolve around the sun, we collect four of those quarter days and add them together into a single day—February 29—that appears on the calendar once every four years.

We do this because there's no such thing as a quarter-day: you couldn't start a New Year at 6:00 a.m. After all, a day is a day and cannot be partitioned like that. It's an integer.

It's important to see that the concept of the yearly calendar comprises 366 days—February 29 is not imaginary. But rather than adding it every four years, what we are really doing is dropping it from the calendar in every year that is not a multiple of four. If the year is not divisible by 4, then we drop February 29 from our count of days in that year.

It's exactly the same with drop-frame time code, where frames are analogous to days, and hours to years. A video frame is a whole thing, an integer, and we count 30 of them in one second. But the rate at which they proceed is a bit less than 30 per second, more like 29.97 frames per second.

This is the same sort of fractional discrepancy that exists in the annual rate of 365.25 days per year.

We deal with it the same way, by dropping 2 frames from the count at the very beginning of every minute that is not a multiple of 10. In that first second, there are only 28 frames.

So frames 00 and 01 simply do not exist at the beginning of every minute of time code that doesn't have a zero at the end of it (10, 20, 30, 40, 50, and 00 minutes being the exceptions), just as February 29 does not exist in any year that can't be divided by 4. It's as simple as that.

Why go to the bother of doing this? For the calendar, it's long been considered important that the seasons start at roughly the same time every year: if we didn't have February 29 as a corrective, then the beginning of Spring, for example, would progress steadily back through February, January, December, and so on as the years rolled by.

For producers, it's important that the time displayed by your time code reader agrees with the real-time clock on the control room wall. Without drop-frame time code, a one-hour program as measured by your time code would actually run 3 seconds and 18 frames too long, and that would wreak havoc with broadcast schedules.

Note that what we are NOT doing is cutting out frames from our program and leaving them on the cutting room floor, as some of my former students at the Toronto Film School used to believe. Those "dropped" frames are simply never there in the first place, just as February 29 will not "be there" in 2013, 2014, and 2015. The calendar works as "drop-day" code.

The takeaway from this blog entry is that if you can intuitively grasp the concept of leap year, then you've already got the essence of drop-frame time code. Conceptually, they are one and the same.

Tuesday, February 28, 2012

Seller Beware—When You're Being Shopped for a Price

I got an email last week asking how much I would charge to mix 5 songs for a band's EP. I wrote back asking whether I'd be recording the original tracks, or just mixing tracks that someone else has already recorded for the band. Both, came the reply, and how much would it cost?

I wrote back to ask about the band, number of players, what instruments, etc., so that I could price out the job using the appropriate recording facility, and asking for a couple of possible dates when the band wanted to start recording. This is an important question, because in sales, there's a strong relationship between price, availability and delivery. You can sometimes get a better deal on studio time that would otherwise remain unbooked.

At this point, the answers started to get vague. What was clear, however, was that I was being shopped for a price. In other words, the prospect (not yet a client) had no intention of coming to me for the job, but was only trying to get a handle on the price of a job that most likely he was bidding on himself.

This happens to everyone from time to time, and is one of the reasons why it's not a great idea just to shoot out a price in response to an inquiry. Every job is different in some way, and a big part of the sales process is asking questions to qualify the buyer.

Asking questions not only keeps valuable business intelligence—your pricing policies— out of the hands of your competition, it also saves you from wasting time with tire kickers who would otherwise take up a lot of your time, but never end up buying anything.

The 80-20 rule seems to apply here: 80% of your business comes from 20% of your prospects. This is further refined so that in turn, 80% of your income comes from 20% of them. In other words, 4% (20% of 20%) of your potential clients are responsible for about two-thirds (80% of 80%) of your business.

Asking questions is your best line of defense here, and a genuine prospect will appreciate that you're drilling down in order to provide the best possible service.

Wednesday, January 18, 2012

MIDI Orchestration

I've just finished a first read-through of the 4th edition of Paul Gilreath's The Guide to MIDI Orchestration, published by Focal Press. Coming in at 600 pages, it's a pretty thorough introduction to the subject of orchestrating a musical composition using MIDI-based equipment and instrument sample libraries.

Several topics of interest to non-MIDI orchestrators and project studio folks alike are covered here, including instrument ranges and playing techniques, notation, voice leading, distribution of melodies and accompaniments to different instruments in the various sections, combining instruments to create different sounds, and achieving specific moods with orchestrations.

For the MIDI producer, there's a wealth of information on equipment, software, choosing sample libraries, and sequencing strings, woodwinds, brass, percussion, piano, harp and voices.

The author illustrates portions of the text with screenshots from his favourite digital audio workstations: Cubase/Nuendo, Logic, Digital Performer and Sonar. Unfortunately, Gilreath dismisses ProTools, saying it's "still working to catch up" in this area, which is too bad, given its market penetration together with the strides that ProTools has made in the whole area of MIDI sequencing and sampling since version 8 was released at the end of 2008.

In any event, readers working with ProTools can easily adapt the material to their way of working, which, in most respects, is not too different from the others.

Supplementary material is available online at www.midi-orchestration.com, including reviews of several good instrument sample libraries, supplementary tutorials, and audio examples, but you have to sign in to access it.

Much of this material—including several full-length, uncut chapters—is still freely available for download as a single zip file from Focal Press at www.focalpress.com/midiorchestrationfiles.aspx.

The book is a terrific reference, and it's refreshing that the author combines description and prescription in almost equal amounts, which is a rare feat. Anyone looking for a grounding in MIDI orchestration would do well to own this book, and will check in with it on a regular basis, if not frequently. It's a beautiful volume, well designed and easy to read, with a crisp and clear layout, and it should be on every music producer's reference shelf.

My criticisms are few. There are numerous errors relating to missing or misplaced illustrations or examples, including a missing reference section with bibliographical apparatus that the author himself refers to twice yet, strangely, is nowhere to be found!

There are also too many proofreading errors, not many of them spelling mistakes, which leads me to believe that spell-check may have served as a convenient substitute for a thorough proofing. This is a tad disappointing in a 600-page book priced at US$82.50 ($86.50 Canadian).

Fortunately, these shortcomings are overshadowed by the author's monumental achievement in turning out what was surely the crowning achievement of his career as a composer for film and television. May his new life as a dentist in Atlanta be as fulfilling!

The Guide to MIDI Orchestration, 4th Edition, by Paul Gilreath, published by Focal Press, 2010. ISBN: 978-0-240-81413-1

Monday, July 11, 2011

If your glass is more beautiful than the wine, change the wine

So says noted wine writer Tony Aspler.

Owners of smaller home and project studios who are tempted to hire top-notch professional recording engineers to help ramp up their business risk seeing clients follow these engineers to better studios.

I've seen this happen time and time again. Home and project studio owners need to understand that this is almost always inevitable when their reach starts to exceed their grasp and they want to compete with the big boys.

It may be prudent for smaller studio owners to consider a significant upgrade of their rooms and equipment before enlisting the services of established outside recording engineers.

It won't be the engineers' fault if clients seek to follow them to greener pastures.

Saturday, July 9, 2011

When a studio won't release the tracks you've paid for

Songwriters and musicians who use the services of small, home-based studios—and, for that matter, established commercial studios—would be well advised to establish the terms, conditions and policies of the studio before undertaking any recording.

I’m writing this because a beginning songwriter has come to me for advice. He is having a hard time getting a home studio owner to release his tracks, even though he has paid his bill in full—over $13,000! The songwriter wants to take the basic tracks of three songs that he recorded at this particular home studio to another, larger commercial studio for editing, mixing and mastering, but the owner of the home studio is refusing to release the material.

By way of explanation, the home studio owner wrote to me saying that he is “not going to let him take the files out of this studio in the condition they are. How he is going to come up with a better mix than any one of our engineers is beyond me to imagine. The files simply aren't ready to be exported or given away. Many guitar and other instrumental parts remain unedited and comped so even if he wanted them they are in no condition to give away.”

(The songwriter has made it clear that editing and comping are among the tasks he wants to complete at the new studio.)

The home studio owner then goes on to say, “We are not willing to give the files out. This is not normal practice. If he really wanted the files he'd have to buy us out but at this time I am not willing to even consider this.”

I’m not sure I understand why the owner thinks that releasing tracks that have been paid for is not “normal practice.” It’s also not clear to me what he means about the songwriter having to “buy us out,” given that he paid his bill in full over six months ago!

The studio owner concludes, “If this project goes out of the studio I have no guarantee if it will be mixed to a certain standard and I've brought in some heavy players that [the songwriter] got at cost.”

Why the studio owner considers it his business that the songs will be “mixed to a certain standard” is beyond me. That is not his responsibility. And the part about providing players “at cost” seems to indicate that the studio owner is in the habit of marking up session players’ fees and then taking a piece for himself. Maybe that’s how the cost of three unfinished demos climbed to beyond $13,000!

This might never have become an issue if the songwriter had clearly established the ground rules at the outset. At this point, it looks like a case for small claims court.

Songwriters be warned: Make sure you know at the outset what policies, practices or procedures a studio considers to be normative before you record a single note there. Second, make it your business to pay session players directly. Don’t accept an all-in deal with the studio, where you pay everything to the studio. In fact, it should be your job—or your producer’s job—to hire the session musicians in the first place.

I know of one instance where a guitar player was unable to attend a session because his wife went into labour that morning with their first child. I was there when the studio owner actually called the rental department of a local music store (Long & McQuade in Toronto) to find an on-the-spot replacement. When the replacement guitarist arrived, it became painfully apparent that he couldn’t read the chart. In fact, he couldn’t even tune his guitar, and he was sent away with return cab fare paid by—you guessed it—the songwriter!

At that point, the songwriter should have called it quits, but he was too cowed by the studio owner to voice his displeasure.

My final recommendation is that songwriters should bring a USB drive—preferably 8 GB or 16 GB—to their sessions, so that they can take a backup of their recordings away with them—provided, of course, that their account with the studio is paid up to date. After all, possession of the recorded material is the only security a studio has against non-payment for services rendered.

I have not named the offending studio in this blog entry, but readers who wish to continue the discussion can email me at buzz@abcbuzz.com.

Friday, November 19, 2010

When a studio's clients want to leave—and take you with them

As a freelance recording engineer, what do you do when artists approach you directly and ask you to produce their recordings in your own right, after you've made their acquantance in someone else's studio and on someone else's dime?

This is a situation I've found myself in a few times over the years. On two occasions, clients told me that the studio we met in didn't sound good and wanted to go elsewhere—with me—to record in the future. In another case, a performer said that the producer was too hyper, "not laid back enough," and could we record somewhere else where the atmosphere was more relaxed. Similarly, in another instance, an artist complained of being stressed out in the presence of the studio owner, and consequently couldn't turn in a first-rate performance.

When this started happening, I thought it best to speak frankly with the producer or studio owner, one of whom told me that none of his other freelancers would ever dream of "poaching" his clients, whose business he had worked diligently to acquire over months and years.

I pointed out that I wasn't in the habit of poaching clients—the industry is far too small and such behaviour is ruinous to one's reputation. I suggested that those who were dissatisfied were eventually going to go elsewhere anyway, and rather than lose their business entirely, why not work out a finder's fee arrangement for him—call it a commission or kick-back if you will—so that his role in securing the business was recognized tangibly.

He rejected this suggestion out of hand, insisting they were HIS clients, and that I should endeavour to convince them to stay despite their expressed concerns.

Well, no one is automatically entitled to a client's business for life, and as the saying goes, whatever it took to get you here isn't enough to keep you here—you're only as good as your last gig. Clients are free to go where they will.

Knowing that some clients may approach competent staff members directly in an attempt to secure their services at more favourable rates, some employers insist that their employees sign a non-compete agreement.

But this doesn't wash with many freelancers, something studio owners should bear in mind when building a business based on out-sourcing the work. It's a two-edged sword—freelancers are not employees, and do not enjoy the same benefits and security that employees do, and turnabout is fair play.

Some savvy studio owners offer freelance engineers a piece of the action—participation in the business akin to stock options—in order to make it more attractive for freelancers to discourage the studio's clients from going elsewhere. I have suggested this on several occasions, with mixed results.

In the end, it's difficult, if not impossible, to maintain a good working relationship with a studio when its clients become dissatisfied and want you to take them elsewhere to record. But for a studio owner, recognizing that you're in a business relationship with freelance engineers, and not a social one, is a necessary first step in arriving at a business solution to what might otherwise become a thorny personal problem.

For a studio owner, it should serve as a wake-up call that all is not as it should be when clients express their dissatisfaction to a sympathetic ear behind the board. The solution may be as simple as staying away from the session, however tempting it may be for a studio owner to participate in the proceedings. From a "strictly-business" perspective, this is the most straightforward solution.

In the case of a home or project studio, this is not as easily accomplished, and sharing a business with a home may open up the studio operation to a level of personal micromanagement that may be detrimental to its success, exposing clients to everything from a simple request to remove their shoes, to disagreeable cooking odors emanating from the kitchen.

When home studio rates are well below prevailing commercial rates, such irritants may well be tolerable, but if home studio owners set their rates equivalent to—or higher than—commercial operations, they may need to adjust their expectations and behaviour accordingly.

Tuesday, October 26, 2010

How much headroom do you need when recording? Part 1

Some years back, I recorded a Stravinsky symphonic work to analog tape running at 15 ips half-track, dbx type 1. When I transferred the recording to digital for archiving, the big orchestral bass drum gave me problems. Had I been recording to digital on location, it would have been a mess: I would have needed to set 0 VU at -25 dBFS to keep the digital meter out of the red. That's how much energy the bass drum was putting out. 

Standards organizations specify how much headroom should be available above operating level (0 VU) in the digital domain. In my experience, the European EBU standard of 0 VU = -18 dBFS does not afford nearly enough headroom, and neither does the North American SMPTE standard of -20 dBFS. Granted, these were "reasonable" compromises in the days of 16-bit technology, when 93dB dynamic range was about all you could expect to get in the real world (as opposed to the theoretical 96 dB, calculated at 6 dB per bit).

Analog headroom of 24 dB—which many manufacturers of professional grade equipment achieve with maximum output levels of +28 dBu (ref 0 VU = +4 dBu)—should be considered the minimum standard during production. Even then, the Stravinsky would have been into overload by about 1 dB, so you might occasionally require even more headroom. 


In our current 24-bit world, I would say that it's not unreasonable to demand 28 dB headroom when recording wide dynamic range program material, such as symphonic works. It still gives you a working signal-to-noise ratio of better than 100 dB, and 28 dB of headroom includes a small comfort margin so you can enjoy the program without stressing over the levels during recording, knowing that you will most likely never go into the red.

Incidentally, I see a lot of "analog channels" being marketed as quality front end processing for recording into ProTools and other DAWs. Many of these boxes include a compressor after the mic preamp, probably because few recordists stop to consider how much analog headroom is really needed in a given situation. Instead of backing the level off to allow for enough headroom without compression, they tend to run the ProTools meters high up into the yellow, recording with compression on individual tracks at 24-bit resolution. It's as if a little bit of green flickering at the low end of the meter must be avoided at all costs. This is foolishness.


24-bit technology allows you to record at a moderate level with 28 dB of headroom and still accumulate no perceptible noise in the recording. Save the compression for mixing and mastering, when it becomes a creative tool rather than a protective device. And then, whatever else you do, don't normalize! Oversampling digital-to-analog converters, which are pretty much the norm these days, routinely create signal peaks greater than 0 dBFS between samples that measure 0 dBFS on disc. But that's another subject for another time.

Tuesday, September 14, 2010

Plan ahead for split-track recording when making documentary videos

A wedding videographer came to me last week with an audio problem in a wedding video he was editing. Just before their first dance as husband and wife, the happy couple had taken to the dance floor with handheld mics to sing a torch ballad over a karaoke track. A huge crowd-pleaser, this would have been a "highlight" of the wedding video. Unfortunately, on tape the backing track was so loud that the lyrics were completely unintelligible—in fact, you couldn't hear the man's voice at all. Could I save the track? And his reputation?

In a nutshell, no, not this time. The voices were so faint in the two-track mono mix and the electric piano so loud  that no amount of EQ could bring them forward. The client even sent me the original wav file of the karoake track, hoping that if I mixed it out of phase (reversed polarity) with the botched recording, it would cancel out some of the music track, leaving the voices more intelligible.

For the cancellation trick to work, the music in the end recording would have to be almost identical to the original karoake file. It wasn't: the karaoke was in stereo not mono, and it ran slightly faster than in the final recording.

Moreover, there was tons of ambience in the duet recording, because the karaoke playback was picked up not only directly from the audio mixer at line level, but also from the loudspeakers in the hall by the singers' microphones. So the final recording is muddied and colored by the reproduced sound from the loudspeakers and all sorts of reflection from the walls and floor, which, of course, is not present in the original track. 

This also partially explains why the music was so much louder than the vocals: the music was being picked up twice, once as a direct feed, and secondly from the singers' microphones. 

What should the operator have done? Since the goal was apparently not to make a stereo recording—no attempt was made to preserve the original left-right stereo of the instrumental track in making a two-track mono recording—the operator should have fed both sides of the original instrumental track to track 1 of the camera, and the microphones to track 2.

This simple split-track technique would have allowed for the voices to be properly balanced with the music in post production. It's done all the time in recording dialog for films and TV on 2 tracks, especially when there is a mic on a boom or fishpole, and the actors are wearing lavs. Boom goes to track 1, lavs to track 2. 

When shouting and screaming are anticipated in a scene (perhaps not at a wedding), the same audio is printed on both tracks but at a reduced level—say, 10 dB lower than normal—on track 2. So if the scream is distorted on track 1, the sound editor can lay up the lower, undistorted scream from track 2, and match the levels during mixdown.

It's simple and it works. It just requires a little planning ahead, but when you only have one shot at it, split-track recording can save the day. And your reputation.