Research
Experiments in language, memory, and learning.
Wisp Vision Model
for Speech Recognition
Audio + Vision → Fusion → Transcript
Question
Can visual information improve spoken-English recognition for language learners?
Approach
Combine audio signal processing with visual lip-reading cues from the Wisp Vision Model. Fuse both streams into a single transcript, then compare against the target subtitle line.
Observations
Learners can speak the line they hear instead of typing it, lowering the barrier for users who struggle with spelling.
The model returns a confidence score and highlights words that differ from the original caption, giving immediate pronunciation feedback.
Visual lip-reading cues may disambiguate similar phonemes such as /p/ and /b/ or /l/ and /r/.
The system could support shadowing practice — repeating the line and scoring rhythm, intonation, and word accuracy.
The long-term goal is a multi-modal tool that adapts to typing, listening, and speaking in one flow.
Still Research · 01