Most language apps give your pronunciation a score. Few of them say what they measured. This page says.
When you make a vowel, your tongue and jaw shape a tube, and the tube has resonances. The first two are called F1 and F2. F1 goes with how open your mouth is. F2 goes with how far forward your tongue sits. Together they place a vowel, and they are the reason bit and bet are different words.
The app reads F1 and F2 off your microphone and puts a dot on a chart. The chart is your vowel space, not a general one — a three-second calibration at the start finds your own four corners, because a six-foot man and a nine-year-old do not share a mouth.
The engine is tested against synthesised vowels whose formants are known before the test runs. On the five Patois vowels it lands within a few percent of the target every time, and it gives the same answer at 48 kHz and at 16 kHz, because phones and laptops disagree about sample rate and a reading that changes with the hardware is not a reading.
Silence gets you no dot. Room noise gets you no dot. A clip too short to hold a vowel gets you no dot. This sounds like a small thing and it is the most important behaviour in the whole feature: a dot that keeps drifting while you sit there saying nothing teaches you to stop looking at it.
A high-pitched voice makes F1 harder to measure. That is physics, not a fault in the app, and the honest response is to say so. High readings are flagged and drawn soft, so you know the number is holding less weight than usual.
Pitch is tracked with YIN, an algorithm that finds the repeating period in a waveform. On silence it reports nothing, same rule as the vowels.
Rhythm is measured as the variation between one syllable and the next. Patois and English distribute weight differently, and that difference is a bigger part of sounding wrong than any single consonant. The app describes what it heard — very even, even, mixed, bouncy — and never grades it. There is no correct rhythm to score you against.
It does not score authenticity, and it says so on the page where you use it. Nothing in a formant tells you whether a Jamaican would take you for one. A machine that measured that would be measuring something it has no way of knowing.
The target ring on the chart is drawn from how a sound is articulated, not from a recording of any particular speaker, and it is labelled that way.
The audio is decoded, measured, and dropped. No clip is stored, no reading history is kept, and nothing about how you sound is written down. The calibration keeps four corner points and nothing else. There is a test that fails the build if using the pronunciation check makes a single network request.