Quy P.
Lab Notebook

Research

Experiments in language, memory, and learning.

AI · Vision · Audio·September 2025

Wisp Vision Model

for Speech Recognition

Status: Experimental
AUDIOVISION
│           │└──────┬───────┘       ▼FUSION       ▼TRANSCRIPT

Audio + Vision → Fusion → Transcript

Question

Can visual information improve spoken-English recognition for language learners?

Approach

Combine audio signal processing with visual lip-reading cues from the Wisp Vision Model. Fuse both streams into a single transcript, then compare against the target subtitle line.

Observations

01

Learners can speak the line they hear instead of typing it, lowering the barrier for users who struggle with spelling.

02

The model returns a confidence score and highlights words that differ from the original caption, giving immediate pronunciation feedback.

03

Visual lip-reading cues may disambiguate similar phonemes such as /p/ and /b/ or /l/ and /r/.

04

The system could support shadowing practice — repeating the line and scoring rhythm, intonation, and word accuracy.

05

The long-term goal is a multi-modal tool that adapts to typing, listening, and speaking in one flow.

Still Research · 01