Voxa

How Voxa analyses a voice

What Voxa measures in a recording, how the eight delivery scores are built from those measurements, and what the numbers cannot tell you. Every definition here is the one the analyser itself uses.

What happens to your recording

1 · Decode
Your file becomes a mono waveform.
Whatever you hand over (wav, mp3, m4a, webm, a video file…) is decoded to a single channel of 32-bit floating-point samples at the file's own sample rate. Stereo is averaged to mono; nothing is normalised or filtered, because loudness dynamics are one of the things being measured.
2 · Prosody
Pitch, loudness and pauses over time.
Fundamental frequency (F0) is tracked frame by frame with normalised autocorrelation, giving mean pitch plus how much it moves. Loudness is a short-window RMS in dB. Silence below an adaptive threshold, held for long enough, is a pause — which yields pause count, mean pause length and the fraction of the recording that is silence.
3 · Voice quality
Cycle-to-cycle steadiness of the voice.
Within voiced stretches, individual glottal periods are located and compared with their neighbours. Variation in period length is jitter, variation in peak amplitude is shimmer, and the ratio of periodic to aperiodic energy is the harmonics-to-noise ratio. These are the classic acoustic correlates of vocal tension, fatigue and breathiness.
4 · Transcript
Words, timings, speaking rate and fillers.
Optional. Speech-to-text produces the text plus a start/end time for every word. From that come word count, speech rate over the whole clip, articulation rate over speaking time only, and the filled-pause count. Without it, roughly a third of the evidence is missing and the affected scores are marked with a lower data-confidence.
5 · Scoring
Features become eight dimension scores.
Each dimension is a weighted average of membership functions — small curves that map one measured feature onto 0–1. Some are ramps (more is better), most are sweet spots (a band is best, and both extremes cost you). Only features that were actually measured take part; the share of the intended weight that was available becomes the data-confidence.
6 · Over time
The same scoring, re-run on a sliding window.
For clips longer than about 14 seconds, the recording is cut into overlapping windows and each is scored independently. Comparing the first third with the last third gives each dimension a trend, and the highest/lowest scoring windows point you at your strongest and weakest moments.

The eight scores

Each score runs from 0 to 100 and is a weighted average of simple curves over the measurements below. The drivers are listed largest weight first.

Confidence

How self-assured the delivery sounds: an even pace, few fillers, controlled pauses and a voice that does not waver.

Driven by:

  • fillers (weight 1.2) — pushes it up when: few filler words
  • speech rate (weight 1) — pushes it up when: well-paced delivery
  • pause ratio (weight 0.8) — pushes it up when: controlled pausing
  • pitch stability (weight 0.8) — pushes it up when: stable, un-shaky pitch
  • loudness variation (weight 0.6) — pushes it up when: steady loudness
  • jitter (weight 0.6) — pushes it up when: clear, un-strained voice

Caveat: Measures steadiness, not conviction. A calm speaker who is deeply unsure will still score high.

Charisma

How compelling you are to listen to — pitch range, intonation movement, dynamic loudness and a resonant voice.

Driven by:

  • pitch range (weight 1.2) — pushes it up when: expressive pitch range
  • pitch variation (weight 1) — pushes it up when: dynamic intonation
  • loudness variation (weight 0.9) — pushes it up when: dynamic loudness
  • voice clarity (HNR) (weight 0.6) — pushes it up when: clear, resonant voice
  • speech rate (weight 0.5) — pushes it up when: engaging pace

Caveat: Rewards expressiveness. A deliberately understated, authoritative style scores lower here without being worse.

Authority

How much weight the delivery carries: a grounded pitch, an unhurried pace, pauses that let points land, and clear projection.

Driven by:

  • pitch (mean) (weight 0.9) — pushes it up when: lower, grounded pitch
  • speech rate (weight 0.9) — pushes it up when: deliberate pace
  • pause ratio (weight 0.8) — pushes it up when: strategic pausing
  • voice clarity (HNR) (weight 0.7) — pushes it up when: clear projection
  • jitter (weight 0.7) — pushes it up when: steady voice
  • fillers (weight 0.7) — pushes it up when: assertive, filler-free

Caveat: Pitch height contributes, which builds in a bias towards lower voices. Weigh it against the pace and pause evidence rather than the total.

Clarity

How easy you are to follow: a clean signal, a moderate rate, distinct articulation and clear phrase breaks.

Driven by:

  • voice clarity (HNR) (weight 1.2) — pushes it up when: crisp, clear voice
  • fillers (weight 1) — pushes it up when: clean, filler-free speech
  • speech rate (weight 1) — pushes it up when: easy-to-follow pace
  • pause ratio (weight 0.7) — pushes it up when: clear phrase breaks
  • articulation rate (weight 0.6) — pushes it up when: distinct articulation
  • shimmer (weight 0.4) — pushes it up when: steady tone

Caveat: Partly measures your microphone and room, through the HNR term.

Warmth

How approachable you sound — melodic intonation, a smooth unstrained voice and an unhurried pace.

Driven by:

  • pitch variation (weight 1) — pushes it up when: melodic, engaged intonation
  • shimmer (weight 0.8) — pushes it up when: smooth voice
  • speech rate (weight 0.7) — pushes it up when: unhurried, personable pace
  • voice clarity (HNR) (weight 0.6) — pushes it up when: open, resonant tone
  • pitch (mean) (weight 0.5) — pushes it up when: inviting pitch
  • jitter (weight 0.5) — pushes it up when: relaxed voice

Caveat: Cultural and language norms for 'warm' speech vary widely; the bands are set for English.

Enthusiasm

How much energy comes through: punchy loudness dynamics, wide animated pitch and a lively pace.

Driven by:

  • loudness variation (weight 1.1) — pushes it up when: dynamic, punchy delivery
  • pitch range (weight 1) — pushes it up when: wide, animated pitch
  • speech rate (weight 0.8) — pushes it up when: lively pace
  • pitch (mean) (weight 0.7) — pushes it up when: raised, energised pitch
  • voice clarity (HNR) (weight 0.4) — pushes it up when: bright voice

Caveat: Close to charisma by design, but weighted towards raw energy rather than expressive control.

Professionalism

How polished the delivery is: fluent, low-filler speech, a composed pace, structured pauses and a clear voice.

Driven by:

  • fillers (weight 1.1) — pushes it up when: fluent, low-filler speech
  • voice clarity (HNR) (weight 0.9) — pushes it up when: clear articulation/recording
  • speech rate (weight 0.8) — pushes it up when: composed pace
  • pause ratio (weight 0.7) — pushes it up when: well-structured pauses
  • shimmer (weight 0.5) — pushes it up when: controlled voice
  • pitch stability (weight 0.5) — pushes it up when: composed intonation

Caveat: A proxy for polish, not for competence or seniority.

Nervousness higher = worse

Acoustic signs of tension: vocal micro-tremor, amplitude instability, hesitation markers, shaky pitch and pressured speed.

Driven by:

  • jitter (weight 1.1) — pushes it up when: vocal micro-tremor
  • fillers (weight 1) — pushes it up when: many fillers/hesitations
  • shimmer (weight 0.9) — pushes it up when: amplitude instability
  • pitch stability (weight 0.9) — pushes it up when: shaky pitch
  • speech rate (weight 0.7) — pushes it up when: pressured fast speech
  • mean pause (weight 0.5) — pushes it up when: long stalling pauses

Caveat: Higher is worse here — the only dimension where that is true. Some voices are naturally less steady, so read it against your own baseline over several recordings.

The measurements

Pitch (mean) (Hz)

The average fundamental frequency of your voice — how high or low you sound.

How: Normalised autocorrelation per 40 ms frame; frames with no clear periodicity are dropped, and the median of the rest is taken.

Typical: Adult voices commonly sit around 85–180 Hz or 165–255 Hz depending on the speaker. There is no 'good' value.

Caveat: Anatomy, not skill. It is used only where pitch height genuinely correlates with perception (authority, energy), and never as a quality score on its own.

Pitch range (st)

The spread between your lowest and highest pitch, in semitones.

How: The 5th-to-95th-percentile span of the tracked F0, converted to semitones so it is comparable between high and low voices.

Typical: Engaged speech typically spans 8–14 semitones. Under ~4 reads as monotone.

Caveat: Percentiles, not extremes, so one squeak does not inflate it.

Pitch variation (st)

How much your pitch moves around moment to moment — your intonation.

How: Standard deviation of the semitone-converted F0 track.

Typical: Around 3–9 semitones reads as lively but controlled. Very high can read as unsteady rather than expressive.

Pitch stability

Pitch variability relative to your own average — a shakiness proxy.

How: Coefficient of variation of F0 in Hz (standard deviation ÷ mean).

Typical: Roughly 0.06–0.16 is a stable, natural voice. Near zero is flat; high is shaky.

Caveat: Normalising by your own mean makes this comparable across speakers, unlike raw Hz spread.

Loudness (dB)

Average signal level.

How: RMS energy per short window, in dB.

Typical: Depends entirely on mic and distance — informational only.

Caveat: Not scored: it measures your recording setup, not your delivery.

Loudness variation (dB)

How much you vary your volume — vocal dynamics and emphasis.

How: Standard deviation of the per-window loudness in dB, over speech only.

Typical: About 2–6 dB is dynamic, engaged delivery. Under ~1.5 dB is flat.

Caveat: Mic compression and auto-gain (common on phones and headsets) flatten this artificially.

Speech rate (wpm)

Words per minute across the whole recording, pauses included.

How: Word count from the transcript ÷ total duration.

Typical: Conversational presenting is roughly 125–170 wpm. Above ~200 reads as rushed.

Caveat: Needs a transcript. Language-dependent: the bands are set for English.

Articulation rate (w/s)

Words per second of actual speaking, with pauses removed.

How: Word count ÷ speaking time (total duration minus detected pauses).

Typical: About 2.2–3.6 words per second is comfortable to follow.

Caveat: Separating this from speech rate distinguishes 'talks fast' from 'pauses little' — two different habits with the same wpm.

Fillers (/min)

Filled pauses and hedges per minute: um, uh, like, you know, I mean…

How: Regex match over the transcript against an editable filler list, scaled to a per-minute rate.

Typical: Under ~3/min is barely noticeable. Above ~10/min is distracting.

Caveat: Indicative, not exact: 'like' and 'actually' have legitimate uses and are counted anyway. The word list is in features/transcript.py.

Pause ratio

The fraction of the recording that is silence.

How: Total detected pause time ÷ total duration.

Typical: About 0.12–0.35 is well-structured speech. Very low leaves listeners no room; very high reads as hesitant.

Caveat: Room noise raises the silence floor and can hide short pauses.

Mean pause (s)

Average length of the pauses you take.

How: Mean duration of silent stretches longer than the detection minimum.

Typical: Roughly 0.3–0.8 s is normal phrasing. Over ~1.5 s starts to read as stalling.

Pauses

How many distinct pauses were detected.

How: Count of silent stretches above the minimum length.

Typical: Scales with clip length; read it together with pause ratio.

Jitter

Cycle-to-cycle wobble in pitch period — a micro-tremor in the voice.

How: Mean absolute difference between consecutive glottal period lengths, divided by the mean period.

Typical: On this implementation's scale, roughly 0.02–0.04 is a steady voice.

Caveat: This clean-room measurement runs about 2× higher in absolute terms than Praat's, so do not compare the number to published Praat thresholds. It is internally consistent, and the rubric bands are calibrated to this scale.

Shimmer

Cycle-to-cycle wobble in amplitude — unsteadiness in vocal power.

How: Mean absolute difference between consecutive period peak amplitudes, divided by the mean amplitude.

Typical: Roughly 0.05–0.12 on this scale reads as controlled.

Caveat: Same calibration caveat as jitter. Compression and clipping distort it.

Voice clarity (HNR) (dB)

Harmonics-to-noise ratio: how much of your voice is clean tone versus breath and noise.

How: Ratio of periodic to aperiodic energy in voiced frames, in dB, from the autocorrelation peak.

Typical: Higher is clearer. Above ~15 dB is a crisp signal; under ~8 dB is breathy, noisy, or a poor recording.

Caveat: This measures the recording as much as the voice — a bad room or a distant mic will lower it no matter how well you speak.

Voiced fraction

Share of frames where a pitch could be tracked at all.

How: Voiced frames ÷ total frames.

Typical: A quality check rather than a score: low values mean noise, whispering or very little speech.

Words

Number of words transcribed.

How: Token count over the transcript.

Typical: Under about 40 words leaves the rate and filler figures noisy.

Duration (s)

Length of the recording.

How: Sample count ÷ sample rate.

Typical: Aim for 60–180 s. Below ~8 s nothing here is reliable; below ~14 s there is no timeline.

Reading the results

How a score is built

Every dimension is a weighted average of small curves. Each curve takes one measured feature and returns 0–1: ramps say 'more is better' (or 'less is'), trapezoids say 'this band is best and both extremes cost you'. The weighted mean of the available curves, times 100, is the score. Nothing is trained, fitted or hidden — the curves and weights are a hundred lines of readable rules, and changing them changes the scores.

Data confidence

The share of a dimension's intended evidence that was actually measurable. A dimension expecting six features that only got four of them (say, because there was no transcript) reports a lower confidence, and the score itself is computed from just those four. A 70 at 55% confidence is a much weaker claim than a 70 at 100% — treat the confidence as the error bar.

Evidence lines

Under each score are the measured features that moved it most, largest weight first, each with its value and whether it helped or hurt. If a score surprises you, the evidence is where to look: it names the number responsible, and that number is something you can go and change.

The over-time view

The same scoring re-run on overlapping windows across the recording, so a single average does not hide the shape of the talk. Each sparkline is drawn so that up is always better — including nervousness, which is inverted. The trend arrow compares the first third against the last third, so it answers 'did I settle in or fall apart?'. Needs about 14 seconds of audio.

The emotional read

Two crude axes. Arousal (energy) is a weighted blend of loudness dynamics, tempo, pitch range and pitch height, and is reasonably well grounded in acoustics. Valence (positive vs negative) is a much weaker proxy built from expressiveness, clarity and pace — acoustics alone genuinely cannot distinguish excited from agitated. Read the label as a prompt, not a finding.

Getting a good recording

Sixty to a hundred and eighty seconds of continuous speaking on one subject. A quiet room, a consistent distance from the mic, and no headset noise-suppression or auto-gain if you can turn it off — those flatten exactly the loudness dynamics being measured. Speak as you would to an audience rather than reading, and keep conditions identical between recordings so the comparison means something.

What this cannot tell you

These are indicative reads of vocal delivery, not facts about you. Accent, native language, gender, age, microphone, room acoustics and background noise all move the numbers, and the rubric bands are set for English speech recorded in a quiet room. Nothing here is validated against human ratings. The honest use is longitudinal: record yourself under the same conditions over time and watch your own trend, rather than comparing your absolute score with anyone else's.

Where your audio goes

Everything runs on your device. The browser decodes the recording, a WebAssembly module measures and scores it, and the optional transcription runs Whisper inside the page. The speech model is downloaded once; the recording is never uploaded, stored or shared.

Analyse your voice