The Supersaw Is Not a Waveform, It’s a Stack
A supersaw is not an oscillator type. It is a sum of several sawtooth oscillators, each shifted slightly in frequency from the others, mixed into one voice. That distinction matters because it means you can build one anywhere — in any synth with unison, or by stacking oscillators by hand — and you can change every property of it.
The Roland JP-8000’s implementation, the one everything else references, used seven sawtooth voices. Two details from published analyses of that oscillator explain most of its character. First, the detune control does not spread the voices linearly: the mapping from knob position to cent offset is a curve, so small knob movements at the bottom of the range do almost nothing and the interesting region sits in a narrow band. Second, the seven voices are not mixed at equal level — the center voice sits above the six detuned ones, and the side voices come up in level as detune increases. That asymmetric mix is why the JP-8000 sound keeps a clear pitch center instead of turning into a wash.
If your mental model is “supersaw = preset I load,” replace it with “supersaw = a stack I control.” Everything below is a parameter of that stack.
Detune: What Widens the Sound Is Actually Phase Drift
Two sawtooths at slightly different frequencies do not stay in a fixed relationship. Their phase difference advances continuously, and every partial in the spectrum drifts against its neighbor at a rate proportional to its harmonic number. The ear reads that drift as beating — a slow amplitude and timbral movement that we call chorusing when it is gentle and call out-of-tune when it is not.
The rate is arithmetic. A detune of c cents at fundamental f produces a frequency difference of roughly f × c / 1731 Hz. At A4 (440 Hz), 10 cents gives about 2.5 Hz of beating at the fundamental and around 25 Hz at the tenth harmonic. That is why supersaws shimmer in the upper mids before they sound wide at the bottom.

Practical ranges, measured as total spread across the stack: under 5 cents you get thickening without obvious movement; 8–20 cents is the classic trance chorus; 25–40 cents starts reading as deliberate detuning, useful for a hoover or a rave stab; past 50 cents the pitch center dissolves and the lead stops carrying melody.
The trap is that cents are proportional, not absolute. The same 20-cent spread that shimmers at C5 produces sub-1 Hz beating at C2, and the harmonics of the detuned voices sit close enough to collide into low-frequency mud rather than movement. Two fixes: reduce detune on lower notes with a key-tracking modulation if your synth offers it, or simply never play the wide supersaw layer below C3 and give the low register to a separate mono layer.
The Trade-off Between Voice Count and Density
Adding voices does not make the sound louder or denser in a linear way. It fills the spectral gaps between the partials of neighboring voices. Three voices leave audible space; the beating is discrete and you can almost count it. Seven voices fill most of that space, which is exactly why the JP-8000 landed there. Going to 16 adds little new spectral information — the gaps are already filled — while multiplying CPU cost, increasing the chance that partials cancel each other, and making the mono sum less predictable.
Choose by role. A single-note lead that has to cut through a busy drop wants 5–7 voices: enough width to sound modern, few enough that the pitch center stays sharp and the mono fold-down stays stable. A pad playing three- or four-note chords is already summing 3–4 stacks, so 3–5 voices per note is usually enough; pushing 16 voices per note across a chord produces a smear where individual chord tones stop being identifiable. If your synth only offers even voice counts, prefer the option that keeps a center-weighted voice, because an even spread with no true center loses the pitch anchor.
How Stereo Width Is Built: Pan Spread and Mono Compatibility
Width comes from distributing the voices across the stereo field, not from detune alone. Detune alone, panned center, gives you a thick mono sound. The standard arrangement keeps one voice in the center as the reference and spreads the remaining voices symmetrically left and right, so the left and right channels carry different phase relationships of the same spectrum.
That difference is the width, and it is also the risk. When left and right sum to mono, partials that are near-opposite in phase cancel. Which partials cancel depends on where the drifting phases happen to be, so the mono sum of a wide supersaw is not a fixed timbre — it moves. Typical symptoms are a thinner, hollow sound and an unstable low-mid region.
Check it with a correlation meter: a wide supersaw layer will sit somewhere between 0 and +0.5, and sustained excursions below 0 mean significant cancellation on fold-down. The mono button on your monitor controller tells you the same thing faster. This matters because large club systems commonly sum low frequencies to mono below the crossover, and anyone standing off-axis from one side of the PA hears something close to one channel. The fix is not less width overall — it is keeping the low and low-mid content narrow and letting width live above roughly 300–500 Hz.
Phase Start: Hearing the Same Sound on Every Note
With free-running oscillators, the stack’s phases keep advancing whether a note is held or not. Every note-on therefore catches the voices in a different relationship. The result is that the attack transient and the amount of low-end energy change from note to note — sometimes the partials reinforce, sometimes they partially cancel, and a repeated sixteenth-note pattern comes out uneven.
Phase retrigger resets all voices to a fixed start phase at each note-on. For staccato trance leads and gated plucks, turn it on: the transient becomes identical every time, the low end is consistent, and the rhythm reads as tight rather than wobbly. For sustained pads and rolling chord beds, leave it off — free-running phase is what makes long notes feel alive, and there is no transient to protect. If retrigger produces an audible click, the cause is the amp envelope attack being too short, not the phase reset; a 2–5 ms attack removes it.
Building a Supersaw Lead From Scratch: The Signal Chain

The order is not arbitrary. The supersaw stack and the mono sub saw are separate layers because they need different width: the sub carries pitch and weight and must stay center, while the stack carries width. High-passing comes before the filter envelope so the envelope shapes only the content you intend to keep; high-passing at 100–150 Hz on the wide layer removes the detuned low partials that would otherwise cancel in mono. The low-pass and its envelope then create the movement that makes the lead breathe under the mix.
Saturation goes after unison, not before. Saturating each voice separately just adds harmonics that then detune against each other; saturating the summed stack generates intermodulation between the voices, which is what glues them into one instrument rather than seven. Then EQ: the 200–400 Hz region accumulates energy from every voice and is the single most common source of a lead that sounds huge soloed and muddy in context — a 2–4 dB dip there usually does it. Stereo imaging comes last so it acts on the final spectrum, and so any width you add is not later re-narrowed by a filter.
Sitting the Lead in the Mix: Sidechain, EQ and Reverb Decisions
A wide supersaw occupies a lot of space by design, and in trance it competes with the two elements that cannot move: kick and bass. Solve the overlap with EQ first. High-pass the lead until it sounds slightly thin in solo — in context it will not be — and confirm the bass fundamental region is clear. Only then reach for compression.
Use sidechain to make room, not to pump. The pump is a stylistic choice in some subgenres, but a fast release (40–80 ms) with modest reduction (2–4 dB) on the lead is enough to keep the kick transient readable without turning the lead into a tremolo. If you need audible pumping as an effect, apply it to a duplicate layer or to the reverb return, not to the dry lead.
Reverb: feed it mono and return it stereo. A mono send removes the unpredictable phase relationships of the wide source before they reach the reverb algorithm, so the tail is dense and controllable instead of blurry. High-pass the send around 300–400 Hz and low-pass it around 8 kHz so the tail does not add to the low-mid buildup or fight the hats.
Delay adds width cheaply — a ping-pong or a short stereo offset produces wide movement without the density of reverb. It turns to mud when the delay times fill the gaps in the melody: dotted eighths on a sixteenth-note lead leave no space, and the repeats stack into a drone. Match the delay subdivision to the rhythmic space the melody actually leaves, and high-pass the return the same way you do the reverb.
Your First Step From Preset to Your Own Sound
Do this twice, and write down what each step changes.
First pass: open a supersaw preset you like. Set detune to zero and voice count to 3. You should now hear a plain saw with slight thickening — that is your baseline. Bring the voices back one step at a time, then raise detune in small increments, then open pan spread from zero to full. At each change, note what moved: pitch stability, perceived width, low-end weight, mono sum.
Second pass: start from an empty patch with a single saw oscillator and rebuild the same settings by hand, using the chain in the diagram above. The preset’s designer made dozens of decisions you never saw — filter slope, saturation type, hidden EQ, the detune curve itself. Your sound’s signature lives in the difference between the two passes. Find that difference, and you stop borrowing leads.