How Voxa analyses a voice
What Voxa measures in a recording, how the eight delivery scores are built from those measurements, and what the numbers cannot tell you. Every definition here is the one the analyser itself uses.
What happens to your recording
The eight scores
Each score runs from 0 to 100 and is a weighted average of simple curves over the measurements below. The drivers are listed largest weight first.
Confidence
How self-assured the delivery sounds: an even pace, few fillers, controlled pauses and a voice that does not waver.
Driven by:
- fillers (weight 1.2) — pushes it up when: few filler words
- speech rate (weight 1) — pushes it up when: well-paced delivery
- pause ratio (weight 0.8) — pushes it up when: controlled pausing
- pitch stability (weight 0.8) — pushes it up when: stable, un-shaky pitch
- loudness variation (weight 0.6) — pushes it up when: steady loudness
- jitter (weight 0.6) — pushes it up when: clear, un-strained voice
Caveat: Measures steadiness, not conviction. A calm speaker who is deeply unsure will still score high.
Charisma
How compelling you are to listen to — pitch range, intonation movement, dynamic loudness and a resonant voice.
Driven by:
- pitch range (weight 1.2) — pushes it up when: expressive pitch range
- pitch variation (weight 1) — pushes it up when: dynamic intonation
- loudness variation (weight 0.9) — pushes it up when: dynamic loudness
- voice clarity (HNR) (weight 0.6) — pushes it up when: clear, resonant voice
- speech rate (weight 0.5) — pushes it up when: engaging pace
Caveat: Rewards expressiveness. A deliberately understated, authoritative style scores lower here without being worse.
Authority
How much weight the delivery carries: a grounded pitch, an unhurried pace, pauses that let points land, and clear projection.
Driven by:
- pitch (mean) (weight 0.9) — pushes it up when: lower, grounded pitch
- speech rate (weight 0.9) — pushes it up when: deliberate pace
- pause ratio (weight 0.8) — pushes it up when: strategic pausing
- voice clarity (HNR) (weight 0.7) — pushes it up when: clear projection
- jitter (weight 0.7) — pushes it up when: steady voice
- fillers (weight 0.7) — pushes it up when: assertive, filler-free
Caveat: Pitch height contributes, which builds in a bias towards lower voices. Weigh it against the pace and pause evidence rather than the total.
Clarity
How easy you are to follow: a clean signal, a moderate rate, distinct articulation and clear phrase breaks.
Driven by:
- voice clarity (HNR) (weight 1.2) — pushes it up when: crisp, clear voice
- fillers (weight 1) — pushes it up when: clean, filler-free speech
- speech rate (weight 1) — pushes it up when: easy-to-follow pace
- pause ratio (weight 0.7) — pushes it up when: clear phrase breaks
- articulation rate (weight 0.6) — pushes it up when: distinct articulation
- shimmer (weight 0.4) — pushes it up when: steady tone
Caveat: Partly measures your microphone and room, through the HNR term.
Warmth
How approachable you sound — melodic intonation, a smooth unstrained voice and an unhurried pace.
Driven by:
- pitch variation (weight 1) — pushes it up when: melodic, engaged intonation
- shimmer (weight 0.8) — pushes it up when: smooth voice
- speech rate (weight 0.7) — pushes it up when: unhurried, personable pace
- voice clarity (HNR) (weight 0.6) — pushes it up when: open, resonant tone
- pitch (mean) (weight 0.5) — pushes it up when: inviting pitch
- jitter (weight 0.5) — pushes it up when: relaxed voice
Caveat: Cultural and language norms for 'warm' speech vary widely; the bands are set for English.
Enthusiasm
How much energy comes through: punchy loudness dynamics, wide animated pitch and a lively pace.
Driven by:
- loudness variation (weight 1.1) — pushes it up when: dynamic, punchy delivery
- pitch range (weight 1) — pushes it up when: wide, animated pitch
- speech rate (weight 0.8) — pushes it up when: lively pace
- pitch (mean) (weight 0.7) — pushes it up when: raised, energised pitch
- voice clarity (HNR) (weight 0.4) — pushes it up when: bright voice
Caveat: Close to charisma by design, but weighted towards raw energy rather than expressive control.
Professionalism
How polished the delivery is: fluent, low-filler speech, a composed pace, structured pauses and a clear voice.
Driven by:
- fillers (weight 1.1) — pushes it up when: fluent, low-filler speech
- voice clarity (HNR) (weight 0.9) — pushes it up when: clear articulation/recording
- speech rate (weight 0.8) — pushes it up when: composed pace
- pause ratio (weight 0.7) — pushes it up when: well-structured pauses
- shimmer (weight 0.5) — pushes it up when: controlled voice
- pitch stability (weight 0.5) — pushes it up when: composed intonation
Caveat: A proxy for polish, not for competence or seniority.
Nervousness higher = worse
Acoustic signs of tension: vocal micro-tremor, amplitude instability, hesitation markers, shaky pitch and pressured speed.
Driven by:
- jitter (weight 1.1) — pushes it up when: vocal micro-tremor
- fillers (weight 1) — pushes it up when: many fillers/hesitations
- shimmer (weight 0.9) — pushes it up when: amplitude instability
- pitch stability (weight 0.9) — pushes it up when: shaky pitch
- speech rate (weight 0.7) — pushes it up when: pressured fast speech
- mean pause (weight 0.5) — pushes it up when: long stalling pauses
Caveat: Higher is worse here — the only dimension where that is true. Some voices are naturally less steady, so read it against your own baseline over several recordings.
The measurements
Pitch (mean) (Hz)
The average fundamental frequency of your voice — how high or low you sound.
How: Normalised autocorrelation per 40 ms frame; frames with no clear periodicity are dropped, and the median of the rest is taken.
Typical: Adult voices commonly sit around 85–180 Hz or 165–255 Hz depending on the speaker. There is no 'good' value.
Caveat: Anatomy, not skill. It is used only where pitch height genuinely correlates with perception (authority, energy), and never as a quality score on its own.
Pitch range (st)
The spread between your lowest and highest pitch, in semitones.
How: The 5th-to-95th-percentile span of the tracked F0, converted to semitones so it is comparable between high and low voices.
Typical: Engaged speech typically spans 8–14 semitones. Under ~4 reads as monotone.
Caveat: Percentiles, not extremes, so one squeak does not inflate it.
Pitch variation (st)
How much your pitch moves around moment to moment — your intonation.
How: Standard deviation of the semitone-converted F0 track.
Typical: Around 3–9 semitones reads as lively but controlled. Very high can read as unsteady rather than expressive.
Pitch stability
Pitch variability relative to your own average — a shakiness proxy.
How: Coefficient of variation of F0 in Hz (standard deviation ÷ mean).
Typical: Roughly 0.06–0.16 is a stable, natural voice. Near zero is flat; high is shaky.
Caveat: Normalising by your own mean makes this comparable across speakers, unlike raw Hz spread.
Loudness (dB)
Average signal level.
How: RMS energy per short window, in dB.
Typical: Depends entirely on mic and distance — informational only.
Caveat: Not scored: it measures your recording setup, not your delivery.
Loudness variation (dB)
How much you vary your volume — vocal dynamics and emphasis.
How: Standard deviation of the per-window loudness in dB, over speech only.
Typical: About 2–6 dB is dynamic, engaged delivery. Under ~1.5 dB is flat.
Caveat: Mic compression and auto-gain (common on phones and headsets) flatten this artificially.
Speech rate (wpm)
Words per minute across the whole recording, pauses included.
How: Word count from the transcript ÷ total duration.
Typical: Conversational presenting is roughly 125–170 wpm. Above ~200 reads as rushed.
Caveat: Needs a transcript. Language-dependent: the bands are set for English.
Articulation rate (w/s)
Words per second of actual speaking, with pauses removed.
How: Word count ÷ speaking time (total duration minus detected pauses).
Typical: About 2.2–3.6 words per second is comfortable to follow.
Caveat: Separating this from speech rate distinguishes 'talks fast' from 'pauses little' — two different habits with the same wpm.
Fillers (/min)
Filled pauses and hedges per minute: um, uh, like, you know, I mean…
How: Regex match over the transcript against an editable filler list, scaled to a per-minute rate.
Typical: Under ~3/min is barely noticeable. Above ~10/min is distracting.
Caveat: Indicative, not exact: 'like' and 'actually' have legitimate uses and are counted anyway. The word list is in features/transcript.py.
Pause ratio
The fraction of the recording that is silence.
How: Total detected pause time ÷ total duration.
Typical: About 0.12–0.35 is well-structured speech. Very low leaves listeners no room; very high reads as hesitant.
Caveat: Room noise raises the silence floor and can hide short pauses.
Mean pause (s)
Average length of the pauses you take.
How: Mean duration of silent stretches longer than the detection minimum.
Typical: Roughly 0.3–0.8 s is normal phrasing. Over ~1.5 s starts to read as stalling.
Pauses
How many distinct pauses were detected.
How: Count of silent stretches above the minimum length.
Typical: Scales with clip length; read it together with pause ratio.
Jitter
Cycle-to-cycle wobble in pitch period — a micro-tremor in the voice.
How: Mean absolute difference between consecutive glottal period lengths, divided by the mean period.
Typical: On this implementation's scale, roughly 0.02–0.04 is a steady voice.
Caveat: This clean-room measurement runs about 2× higher in absolute terms than Praat's, so do not compare the number to published Praat thresholds. It is internally consistent, and the rubric bands are calibrated to this scale.
Shimmer
Cycle-to-cycle wobble in amplitude — unsteadiness in vocal power.
How: Mean absolute difference between consecutive period peak amplitudes, divided by the mean amplitude.
Typical: Roughly 0.05–0.12 on this scale reads as controlled.
Caveat: Same calibration caveat as jitter. Compression and clipping distort it.
Voice clarity (HNR) (dB)
Harmonics-to-noise ratio: how much of your voice is clean tone versus breath and noise.
How: Ratio of periodic to aperiodic energy in voiced frames, in dB, from the autocorrelation peak.
Typical: Higher is clearer. Above ~15 dB is a crisp signal; under ~8 dB is breathy, noisy, or a poor recording.
Caveat: This measures the recording as much as the voice — a bad room or a distant mic will lower it no matter how well you speak.
Voiced fraction
Share of frames where a pitch could be tracked at all.
How: Voiced frames ÷ total frames.
Typical: A quality check rather than a score: low values mean noise, whispering or very little speech.
Words
Number of words transcribed.
How: Token count over the transcript.
Typical: Under about 40 words leaves the rate and filler figures noisy.
Duration (s)
Length of the recording.
How: Sample count ÷ sample rate.
Typical: Aim for 60–180 s. Below ~8 s nothing here is reliable; below ~14 s there is no timeline.
Reading the results
How a score is built
Every dimension is a weighted average of small curves. Each curve takes one measured feature and returns 0–1: ramps say 'more is better' (or 'less is'), trapezoids say 'this band is best and both extremes cost you'. The weighted mean of the available curves, times 100, is the score. Nothing is trained, fitted or hidden — the curves and weights are a hundred lines of readable rules, and changing them changes the scores.
Data confidence
The share of a dimension's intended evidence that was actually measurable. A dimension expecting six features that only got four of them (say, because there was no transcript) reports a lower confidence, and the score itself is computed from just those four. A 70 at 55% confidence is a much weaker claim than a 70 at 100% — treat the confidence as the error bar.
Evidence lines
Under each score are the measured features that moved it most, largest weight first, each with its value and whether it helped or hurt. If a score surprises you, the evidence is where to look: it names the number responsible, and that number is something you can go and change.
The over-time view
The same scoring re-run on overlapping windows across the recording, so a single average does not hide the shape of the talk. Each sparkline is drawn so that up is always better — including nervousness, which is inverted. The trend arrow compares the first third against the last third, so it answers 'did I settle in or fall apart?'. Needs about 14 seconds of audio.
The emotional read
Two crude axes. Arousal (energy) is a weighted blend of loudness dynamics, tempo, pitch range and pitch height, and is reasonably well grounded in acoustics. Valence (positive vs negative) is a much weaker proxy built from expressiveness, clarity and pace — acoustics alone genuinely cannot distinguish excited from agitated. Read the label as a prompt, not a finding.
Getting a good recording
Sixty to a hundred and eighty seconds of continuous speaking on one subject. A quiet room, a consistent distance from the mic, and no headset noise-suppression or auto-gain if you can turn it off — those flatten exactly the loudness dynamics being measured. Speak as you would to an audience rather than reading, and keep conditions identical between recordings so the comparison means something.
What this cannot tell you
These are indicative reads of vocal delivery, not facts about you. Accent, native language, gender, age, microphone, room acoustics and background noise all move the numbers, and the rubric bands are set for English speech recorded in a quiet room. Nothing here is validated against human ratings. The honest use is longitudinal: record yourself under the same conditions over time and watch your own trend, rather than comparing your absolute score with anyone else's.
Where your audio goes
Everything runs on your device. The browser decodes the recording, a WebAssembly module measures and scores it, and the optional transcription runs Whisper inside the page. The speech model is downloaded once; the recording is never uploaded, stored or shared.