What does voice to video AI usually mean?
Most voice-to-video tools ask you to upload narration, wait for analysis, and receive an edited collection of stock footage, captions, avatars, or generated clips. That is useful for repurposing a podcast or voice note, but the voice is source material for a future export.
AIR Live uses voice differently: speech is a live control surface. You speak while a generated scene is already playing, and the next part of that same scene responds.
How live voice direction works
- The microphone button requests permission and starts listening.
- Browser speech recognition transcribes your opening scene description.
- Evolink cleans the completed sentence into a visual direction. If cleanup is unavailable, the original words remain usable.
- The first direction starts a paid fal H3 Max Director session once the account has enough credits.
- Later spoken directions update the scene over the same connection.
- Generated video and audio arrive over WebRTC. Stopping the session saves the recording to your account through cloud storage.
For example, “hold on, keep the same woman, move the camera behind her, and make the city lose power” becomes a concise direction that preserves the named subject, camera request, and visual event while removing conversational filler.
Why not use browser transcription alone?
Browser speech recognition produces words, but a spoken sentence can include hesitation or a correction. Evolink helps turn that sentence into a usable direction while retaining the current scene context. It is not the audio transport. Gemini Live is an optional alternative for streaming speech and structured directions; it is not required for the browser recognition path. Typed input remains available when speech recognition is unsupported.
Privacy boundary
AIR Live does not intentionally store raw microphone audio. The browser streams audio to the configured speech provider for realtime processing. Final scene directions may be retained with generation records. You can stop listening at any time or direct entirely by text.
Try speaking to a generated world
Open AIR Live, press the microphone button, and describe your opening scene. Live mode requires 72 credits for 60 seconds; a five-second typed Preview costs 4 credits.