Freak-cha

Blog

Tuning the tongue detector after HackHarvard

We recorded 36 short webcam trials to find out why Freak-cha missed deliberate movements and sometimes answered when nobody meant to. Then we tuned the rules that turn the model's predictions into gestures.

A prediction still needs a decision

Freak-cha asks you to answer with your tongue. Up and down mean yes. Left and right mean no. For each camera frame, the model predicts a tongue position and gives it a confidence score. The code around that model decides whether a run of predictions counts as a hold or repeated movement, and when to accept an answer.

That code is the harness we tuned. We changed its confidence cutoffs, timing, and handling of interrupted movements. The model's weights stayed the same throughout these trials.

One tester, four rounds

We built a recording page with a countdown, a neutral lead-in, six seconds of a prompted action, and a short recovery. The tester replayed and confirmed each clip before moving on. Prompts covered holds and repeated movements, plus ordinary behavior like talking that should leave the detector quiet.

Each confirmed trial saved the video, the exact images sent to the model, its position predictions and confidence scores, and the gestures the app emitted. The recordings stayed local and captured no audio. We could replay new rules against the same predictions without rewriting what happened during recording. An answer that arrived after the action ended counted as late, not successful.

Guided recorder with six trial steps, an empty camera preview, the centered tongue prompt, and the countdown and review instructions.
The recorder before round four. The prompt sits beside the camera, with the countdown and review steps spelled out. The camera is off here, and the local recording folder is hidden. Select either screenshot to view it at full size.

What the recordings changed

The first rules waited for 15 frames. At the roughly four frames per second we observed, that meant nearly four seconds of waiting. We switched to a 2.4-second window measured by elapsed time. Holds now need at least 1.4 seconds of evidence, five confident samples, and 90% agreement. Repeated movement gets checked before holds, so a pause at one end does not immediately become the answer.

We also made a sustained gesture answer once. Half a second of confident neutral behavior releases it for another answer. This stopped a long hold from repeatedly triggering the app.

Round three exposed two narrower problems. The model called an up pose correctly in 22 of 24 action frames, but many scores sat around 0.64, below the 0.65 hold cutoff. One wrong prediction during an up/down transition could also reset the whole movement sequence.

Version three lowered the hold cutoff to 0.60 while keeping the duration and agreement requirements. It also tolerates one wrong movement label for less than 350 milliseconds. Two consecutive confident incompatible labels, or a longer interruption, still reset the sequence. A broader experiment with average confidence made talking trigger an answer. We rejected it.

What passed, and what that tells us

On replay, version three recovered the missed up hold and slow vertical movement in round three. We then recorded six fresh trials. All four intended gestures registered, and the centered pose and talking produced no answer. The gestures took 1.7 to 2.3 seconds after the action cue, with no duplicate answers.

Saved results after four rounds, showing recorded outcomes alongside version two and three replays, plus the first three trial rows.
The results page after all 36 trials. Round summaries separate the answers recorded at the time from later rule replays. The first rows still show the original centered-pose false trigger.
Replay of all 36 saved clips
Recognition rulesCorrect first outcomes
Version two29 / 36
Version three31 / 36
A correct first outcome means the first answer matched the prompt, or the detector stayed quiet when no answer was expected. Both versions used the same saved model predictions.

Version two also passed all six fresh clips on replay. Those clips had better raw model predictions, so we cannot credit the rule changes for every success. Across all 36 clips, version three added two correct first outcomes without adding false or duplicate answers. That is evidence from one tester, not an accuracy claim for everyone who tries the app.

The rules are in the app now

The camera now calls the model inside the Next.js server. Both the current camera API and the older endpoint use one detection service, which applies version three to the model's outputs. We checked the live path with an actual model prediction around 0.64 and confirmed that a sustained up pose answered once.

Earlier recordings still contain centered poses labeled down and right poses labeled left. Timing rules cannot reliably fix those image mistakes. The next useful test is more people, cameras, and lighting conditions, followed by model training if those errors persist. For now, we have recordings that show which part needs work.

A personal reflection

The speed at which we could spin up a little app for such a specific improvement is what stays with me. We wanted to understand a few missed tongue movements, and we could build a tool to record them, inspect the predictions, and test a change.

To me, this is the beginning of a new age. AI gives me the ability to do work I couldn't do before, including making tools for these small improvements as I need them. All of this was possible with GPT 6.1 Sol.