How SonaLab Reads Your Voice
Open the app and you will see a wall of charts. Where do those numbers come from? — From your voice.
This page is the prerequisite primer for all SonaLab FAQ documents. You need no acoustics knowledge — just 10–15 minutes of hands-on play. We will start from "what is sound" and work our way to every line on the spectrum, and in the end you will discover: every number in the app is a part of your voice.
The Journey from Sound to Numbers
Your voice travels this pipeline to become the charts in the app. Each chapter on this page lights up one link of the chain. Try scrolling.
Chapters 1–4
Chapter 5
Chapters 6–11
Thickness · Closure · Balance
Chapter 12 ties it together
- Sound is recommended (top-right button): every experiment lets you actually "hear" what you are dragging. Without it you still see the animations, but half the experience is missing.
- All values match the parameter meanings inside the app. Chart names in the text link to the chart gallery.
What Is Sound: Vibration & Air Pressure Waves
Sound = a vibrating object pushing air into alternating sparse and dense waves.
When you speak or sing, your vocal folds vibrate, opening and closing rapidly in your throat. Each closure gives the air a push, like dropping pebbles into a pond — ripples spread outward and finally hit your eardrum, and that is when you "hear" sound.
Experiment 1 · Vocal Folds Opening & Closing Create Waves
Drag the frequency slider and watch three things change in sync: how fast the vocal folds open and close, how the air molecules bunch and spread, and the waveform on the oscilloscope. They are three views of the same vibration.
Vibration → air waves → eardrum vibration → sound. Frequency is how many times this vibration opens and closes per second. Every analysis in the app starts from this stream of waves captured by the microphone.
Frequency & Pitch: Speed Decides Height
Faster vibration = higher frequency = higher pitch. Frequency is measured in Hz (cycles per second).
If your vocal folds open and close 200 times per second, that is 200 Hz. The faster they open and close, the "sharper" (higher) the sound; the slower, the "duller" (lower). Human ears hear roughly 20 Hz – 20,000 Hz. When you sing, your pitch is the vibration frequency of your vocal folds — in vocal terminology, F0 (fundamental frequency).
Experiment 2 · Drag Faster or Slower, Hear the Height
Drag the slider to change frequency: the waveform becomes denser/sparser (visual) and the pitch rises/falls (audio). Note: frequency is perceived logarithmically — going from 100→200 Hz and 400→800 Hz sounds like the same leap (both double, an interval of one octave).
Frequency determines pitch. The app's "pitch" chart measures your F0 in real time. That is why female voices are usually higher than male voices — shorter, thinner vocal folds vibrate faster.
Amplitude & Loudness: Height Decides Size
How high the waveform "surges" is amplitude — the larger the amplitude, the louder the sound.
Say "ah—" softly and then belt it: the difference is amplitude — the air is pushed with more force. Loudness is measured in dB (decibels), a logarithmic scale — multiplying amplitude by 10 only adds 20 dB. This matches how our ears perceive loudness (it is not linear).
Experiment 3 · Amplitude & dB
Drag the amplitude slider: the waveform height changes (visual) and the volume changes (audio). Try "halving the amplitude" and listen — does it only sound a little quieter? That is why dB is useful.
Amplitude decides loudness, and dB is the auditory language of "doubling/halving". Many app parameters (H1−H2, SPR) are expressed in dB because that is exactly a measure of "relative strength".
Complex Waves: Why the Human Voice Is Not a Single String
Real sound = fundamental + a stack of integer-multiple overtones. How strong each overtone is decides the timbre.
If sound were a single sine wave, every instrument and every voice would sound identical — a monotonous "beep". Real sound is many frequencies vibrating at once: besides the fundamental F0 there are its integer multiples 2×F0, 3×F0… which we call harmonics H1 (=F0), H2 (=2×F0), H3 (=3×F0). How strong each harmonic is decides the timbre.
Experiment 4 · Take Sound Apart and Stack It Back
Drag the 8 sliders below to adjust how strong each harmonic is. Each row above is one harmonic's waveform; the bottom row is the real sound waveform after they are summed. Try "turning everything off except H1" (it becomes a pure tone), then gradually add H2, H3 back — the "timbre" you hear changes.
Harmonics are "integer-multiple friends of the same fundamental". They decide whether a voice sounds thick, bright, or dark. In the next chapter you will see: a spectrum chart is simply these "friends" lined up by frequency.
FFT & the Spectrum: The Acoustic Prism
The waveform is what sound "looks like"; the spectrum is what sound "is made of". FFT is the algorithm that turns a waveform into a spectrum.
White light passes through a prism and becomes a rainbow — each color is a different component of the white light. Similarly, FFT "takes apart" a complex waveform and tells us "how strong each harmonic is". The result, drawn out, is the spectrum: x-axis = frequency, y-axis = energy (dB). The waveform and the spectrum are two views of the same sound.
Experiment 5 · Open the Prism
Above is the waveform (like white light: all components stacked together); below is its spectrum (like the rainbow the prism splits out). Click "Open prism" to watch the transformation animation: white light → decomposition → rainbow spectral lines, each line being one frequency component — the higher the line, the stronger the energy. The sliders below are exactly the same as Chapter 4 — there you stacked harmonics to see the waveform; here you take the other view and split the waveform into its lines.
Waveform = how you look at time; spectrum = what exists in frequency. The next 6 parameters are all "read" from the spectrum.
F0: The First Line on the Spectrum
F0 (fundamental) = pitch. On the spectrum it is the leftmost, usually tallest line.
Remember? Harmonics are integer multiples of F0: H1 = 1×F0, H2 = 2×F0, H3 = 3×F0… So when F0 changes, all the spectral lines shift together (like an accordion). The app uses the YIN algorithm to find this "most basic vibration period" in the waveform — in other words, it finds F0.
Experiment 6 · Drag F0, the Whole Spectrum Shifts
Drag the F0 slider (the red highlighted line marks where F0 is). Watch all spectral lines move together like an accordion while the pitch changes. This is what "changing a note" looks like on the spectrum.
F0 = the leftmost line on the spectrum = your pitch. The app's "pitch" chart and register detection both start from F0.
H1 & H2: Two Peaks on the Spectral Lines
H1 is the strength of the first harmonic, H2 the second. The height difference between them hides the secret of vocal fold thickness.
The height of a spectral line = how strong that harmonic is. Let's single out H1 (at F0) and H2 (at 2×F0): strong H1 → round, thick waveform (like male voices and chest voice); strong H2 → sharp, bright waveform (like female voices and head voice). On the spectrum it simply comes down to which of the two lines is taller.
Experiment 7 · Drag the H1/H2 Heights
Drag the heights of the two spectral lines (or use the sliders below) and watch the waveform change shape. Try raising H1 while lowering H2, then the reverse — the "thick/thin" feel of the sound changes accordingly.
H1 and H2 are two neighboring lines on the spectrum. They are not just any two lines — their "height difference" is a dedicated acoustic metric, coming up in the next chapter.
H1−H2: The Thickness Metric
H1−H2 (dB) = the amplitude difference between the first and second harmonics, directly reflecting vocal fold thickness.
This is a core metric in vocal analysis. Why? Thick vocal folds (TA contraction) → low harmonics are strong → H1 above H2 → H1−H2 is negative; thin vocal folds (CT stretch) → high harmonics are strong → H2 above H1 → H1−H2 is positive. The ideal range in the app is −3 dB to +8 dB (the green zone) — too thick or too thin both lose points.
Experiment 8 · Find the Balance Between −3 and +8 dB
Drag the slider to change H1−H2 (the height difference between the two lines); the green vertical band is the ideal range. Switch between three vocal fold states with one click and watch the line shapes: Chest (H1 high, H2 low, negative) → Balanced mix (inside the range) → Head (H2 high, H1 low, positive).
H1−H2 = the "thermometer" of vocal fold thickness. This is where the app's thickness balance score comes from, matching the TA/CT pendulum model. Related: TA/CT balance chart
F1 & Formants: The Vocal Tract Is a Filter
The vocal folds are the source (producing all frequencies); the vocal tract is the "filter" — amplifying certain frequencies into formants.
As the sound from the vocal folds passes through the mouth and pharynx it gets "colored": formants are the frequency regions amplified by the vocal tract (F1, F2, F3…), appearing as "humps" on the spectral envelope. The key rule: lower larynx → longer vocal tract → lower F1. A low F1 means a "deep", grounded sound — this is the core of the app's resonance depth score (30% weight).
Experiment 9 · Move the Larynx, F1 Follows
The left side shows a vocal tract cross-section: the larynx (red) position decides the tract length. Drag "larynx position" and watch the tract lengthen/shorten while the F1 hump on the spectral envelope (yellow dashed line) moves left and right.
Formants = frequencies amplified by the vocal tract "loudspeaker". Lower F1 = lower larynx = deeper sound. Related: resonance depth chart
SPR: Singer's Formant Ratio (Projection)
SPR = energy in the high band (2–4 kHz) ÷ energy in the low band (0–2 kHz). It measures whether your voice "carries".
In a trained singer's voice there is a concentrated "singer's formant" around 2.5–3.5 kHz — an acoustic "lens" created by opening the pharynx. SPR quantifies it: SPR = 10×log₁₀(high-band energy / low-band energy). The closer to 0, the brighter and more "metallic" the voice. An orchestra's energy happens to dip right at 2.5–3.5 kHz — so the singer's formant lets a voice cut through the band.
Experiment 10 · Open the Pharynx, Grow a High-Frequency Peak
Drag the "pharynx opening" slider: the pharynx (orange region) widens, an energy bump appears in the 2–4 kHz region of the spectrum (orange zone), and the SPR value updates live. Note: the pharynx is the vertical chamber behind the oral cavity (behind the tongue root, above the larynx). Opening it = the pharyngeal wall moves back + the soft palate lifts, while the oral cavity stays the same size — it is not "a bigger pharynx means a smaller mouth".
SPR = the "projection reading" of your voice. It is the core metric of the app's compression chart (40% weight). Related: singer's formant chart
NAQ: Vocal Fold Closure Quality (A Parameter Invisible on the Spectrum)
NAQ measures how "tightly and cleanly" the vocal folds close — it hides behind the spectrum and can only be seen through "inverse filtering".
The previous parameters can all be "read" off the spectrum. But NAQ (Normalized Amplitude Quotient) cannot — it describes how the vocal folds close. The sound is colored by the vocal tract before reaching the mic, so the app models the tract with LPC, then performs inverse filtering to "peel off" the tract and recover the glottal airflow pulse, then measures: total airflow (AAC) ÷ closure speed (dpeak).
Experiment 11 · Three Fold States, Three Pulses
Switch between three closure modes and watch the fold animation plus the inverse-filtered glottal airflow pulse: Pressed (hard closure, airflow trapped, NAQ<0.1) · Flow (crisp closure, smooth airflow, 0.1–0.25 ideal zone) · Breathy (doesn't close fully, leaking, NAQ>0.3).
NAQ is a "closure quality" reading, not a leakage percentage. 0.5 does not mean 50% of your air leaks — it is a dimensionless ratio. This is exactly what the app's vocal fold closure chart shows. Related: vocal fold closure chart
The Master Workbench: Put All 6 Parameters Back on One Spectrum
F0, H1−H2, F1, SPR, NAQ all come from the same spectrum of the same sound — now watch them interact.
Each chapter so far looked at one parameter. But in real sound they exist together and affect each other. This workbench puts all 6 parameters back on one spectrum: pick a phonation mode, or freely drag every slider, and watch how the spectrum and the 4 dimension scores (depth / compression / thickness / throatiness) interact.
Experiment 12 · Phonation Mode Panorama
- 📖 TA/CT balance chart — depth + compression + thickness − throatiness = balance
- 📖 Resonance depth chart — F1, low-frequency energy, H1
- 📖 Singer's formant chart — SPR and projection
- 📖 Vocal fold closure chart — NAQ and airflow balance
The Journey of Sound, Back to the Start
Remember the pipeline in the prologue? Now you have played with every link yourself:
What sound looks like
What sound is made of
The first line
Height difference of two lines
The vocal tract filter
High-frequency energy peak
The secret of inverse filtering
Next, go read the TA/CT balance chart and you will find every formula now has a picture in your mind. If any parameter still feels fuzzy, return to its chapter and play again — listening with your ears and watching with your eyes beats memorizing definitions.
Sound Fundamentals · a teaching page for the SonaLab Learning Center · values match the app parameters