Voice¶
The film's code
import manimgx as m
class VoiceHero(m.Scene):
def construct(self) -> None:
square = m.MathTex("x^2", font_size=96).shift(2 * m.LEFT)
derivative = m.MathTex("2x", font_size=96).shift(2 * m.RIGHT)
arrow = m.Arrow(square.get_right(), derivative.get_left(), buff=0.4)
script = m.Text(
"The derivative of [x squared] is [two x].", font_size=32
).to_edge(m.DOWN)
self.add(square, script)
self.play(m.Indicate(square))
self.play(m.GrowArrow(arrow), m.Write(derivative))
self.wait()
say speaks a line with the scene's voice, and plays animations while it speaks. Put the words of each animation in brackets, and give the animations in the same order: each starts with the first of its words. The scene waits until the line ends, so lines don't overlap; change the words, and the animations follow the new words.
The scene's voice speaks through a text-to-speech service. manimgx keeps what it says in a
folder named voice/, next to the scene's file, so the scene renders again without the
service, and only a changed line goes to it again. What the voice says becomes the video's
captions.
say ¶
Speak a script, playing animations as it is said; the scene waits for both.
Mark in brackets the words an animation belongs to: the animations given play,
in order, during the bracketed words, each starting as its words start and
lasting as long as they take to say, or its own run time if that is longer (an
empty [] starts one where it stands). Without brackets, the animations play
during the whole speech. The brackets are not spoken.
scriptWhat to say, with a bracketed span for each animation (or none).
*animationsThe animations, one per span.
voiceWho says it; by default the scene's
voice.
Returns The speech, with its words and when each is said.
Source
src/manimgx/scene.py
speech ¶
What the scene's voice says for text, without playing it.
Spoken once, then kept in voice/ beside the scene's file (named by the text), so
the scene renders again without speaking again; replace a file there with your
own recording of the same words, and it is used instead. Play the speech, add it
(add_sound: it speaks while the scene goes on), or
read its words.
textWhat to say.
voiceWho says it; by default the scene's
voice.
Returns The speech.
Source
src/manimgx/scene.py
Speech ¶
Speech: a sound that knows what it says and when it says each word.
A voice returns one. It plays as any sound does; Scene.say
also times animations by its words.
sourceThe audio: a file's path, a file's bytes, or samples.
textWhat it says.
wordsEach word and when it is said, if the voice knows; otherwise they are estimated from the audio.
rateThe samples' rate; only for samples.
Source
src/manimgx/audio/sound.py
words ¶
Its words, in order, with when each is said (in seconds from its start): the voice's, or estimated from the audio; a trim or a change of speed moves them with the audio.
Source
src/manimgx/audio/sound.py
Voices¶
The voices manimgx brings are fal.ai's: m.voices.Fal, with the model you choose and its
settings. Your editor shows each model's settings.
Fal ¶
A scene's voice, in its class:
voice = m.voices.Fal(
"fal-ai/minimax/speech-2.8-hd",
voice_setting={"voice_id": "Calm_Woman", "speed": 1.1},
)
A text-to-speech model on fal.ai, as a voice: Fal(model, **settings), where the
settings are the model's own (ElevenV3 for the default).
modelThe model's ID on fal.
keyfal's API key; by default
FAL_KEY.
The model's settings.
Source
src/manimgx/audio/fal.py
Models¶
ElevenV3 ¶
The settings of ElevenLabs' Eleven v3 (fal-ai/elevenlabs/tts/eleven-v3), which times
its words. Its text may carry audio tags: [whispers], [laughs], [excited]…
voiceA voice's name or ID in ElevenLabs' library: Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily, Bill… (default "Rachel").
stabilityHow steady the voice is, from 0 (expressive) to 1 (steady) (default 0.5).
language_codeThe language, as an ISO 639-1 code ("en", "de"…) (default: the text's).
apply_text_normalizationWhether numbers and the like are spelled out before they are said (default "auto": as the model sees fit).
voice ¶
A voice's name or ID in ElevenLabs' library: Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily, Bill… (default "Rachel").
stability ¶
How steady the voice is, from 0 (expressive) to 1 (steady) (default 0.5).
language_code ¶
The language, as an ISO 639-1 code ("en", "de"…) (default: the text's).
apply_text_normalization ¶
Whether numbers and the like are spelled out before they are said (default "auto": as the model sees fit).
MiniMax ¶
The settings of MiniMax Speech 2.8 HD (fal-ai/minimax/speech-2.8-hd). Its text may
carry pauses, <#0.5#> (in seconds), and interjections: (laughs), (sighs), (coughs),
(clears throat), (gasps), (sniffs), (groans), (yawns).
voice_settingThe voice and how it speaks (
MiniMaxVoice).audio_settingThe audio it makes (
MiniMaxAudio).language_boostA language or dialect it listens for (default: none).
normalization_settingHow it evens out the loudness (
MiniMaxLoudness).voice_modifyHow it changes the voice itself (
MiniMaxTimbre).pronunciation_dictHow it pronounces particular words (
MiniMaxPronunciation).
voice_setting ¶
The voice and how it speaks (MiniMaxVoice).
audio_setting ¶
The audio it makes (MiniMaxAudio).
language_boost ¶
A language or dialect it listens for (default: none).
normalization_setting ¶
How it evens out the loudness
(MiniMaxLoudness).
voice_modify ¶
How it changes the voice itself (MiniMaxTimbre).
pronunciation_dict ¶
How it pronounces particular words
(MiniMaxPronunciation).
MiniMaxVoice ¶
A MiniMax voice and how it speaks.
voice_idThe voice: Wise_Woman, Friendly_Person, Inspirational_girl, Deep_Voice_Man, Calm_Woman, Casual_Guy, Lively_Girl, Patient_Man, Young_Knight, Determined_Man, Lovely_Girl, Decent_Boy, Imposing_Manner, Elegant_Man, Abbess, Sweet_Girl_2, Exuberant_Girl, or a cloned voice's ID (default "Wise_Woman").
speedHow fast, from 0.5 to 2 (default 1).
volHow loud, from 0.01 to 10 (default 1).
pitchHow high, in semitones from -12 to 12 (default 0).
emotionThe emotion it speaks with (default: as the text reads).
english_normalizationWhether English numbers are read more carefully, a little slower (default False).
voice_id ¶
The voice: Wise_Woman, Friendly_Person, Inspirational_girl, Deep_Voice_Man, Calm_Woman, Casual_Guy, Lively_Girl, Patient_Man, Young_Knight, Determined_Man, Lovely_Girl, Decent_Boy, Imposing_Manner, Elegant_Man, Abbess, Sweet_Girl_2, Exuberant_Girl, or a cloned voice's ID (default "Wise_Woman").
speed ¶
How fast, from 0.5 to 2 (default 1).
vol ¶
How loud, from 0.01 to 10 (default 1).
pitch ¶
How high, in semitones from -12 to 12 (default 0).
emotion ¶
The emotion it speaks with (default: as the text reads).
english_normalization ¶
Whether English numbers are read more carefully, a little slower (default False).
MiniMaxAudio ¶
The audio MiniMax makes.
sample_rateSamples a second (default 32000).
bitrateBits a second, for MP3 (default 128000).
formatThe file's format (default "mp3").
channelChannels: 1 (mono) or 2 (stereo) (default 1).
MiniMaxLoudness ¶
How MiniMax evens out the audio's loudness.
enabledWhether it does (default True).
target_loudnessThe loudness it aims at, in LUFS, from -70 to -10 (default -18).
target_rangeThe loudness range, in LU, from 0 to 20 (default 8).
target_peakThe highest peak, in dBTP, from -3 to 0 (default -0.5).
MiniMaxTimbre ¶
How MiniMax changes the voice itself.
pitchHigher or lower, from -100 to 100 (default 0).
intensityMore or less energetic, from -100 to 100 (default 0).
timbreIts tone's color, from -100 to 100 (default 0).
MiniMaxPronunciation ¶
How MiniMax pronounces particular words.
tone_listEach word and its pronunciation:
"text/(pronunciation)"(Chinese tones 1 to 5:"燕少飞/(yan4)(shao3)(fei1)").
tone_list ¶
Each word and its pronunciation: "text/(pronunciation)" (Chinese tones 1 to 5:
"燕少飞/(yan4)(shao3)(fei1)").
Gemini ¶
The settings of Google's Gemini 3.8 Flash TTS (google/gemini-3.8-flash-tts). Its text
may carry vocal events: <laugh>, <sigh>.
voiceThe voice (default "Kore").
style_instructionsHow to say it, in words: "Warm and unhurried, like a patient teacher." (default: none).
Inworld ¶
The settings of Inworld TTS-1.5 Max (fal-ai/inworld-tts), which speaks up to 2,000
characters a line.
voiceThe voice, with its language (default "Craig (en)").
sample_rate_hertzSamples a second (default 48000).
Qwen ¶
The settings of Alibaba's Qwen3-TTS 1.7B (fal-ai/qwen-3-tts/text-to-speech/1.7b).
voiceThe voice (default: the model's).
languageThe language (default "Auto": the text's).
promptHow to say it, in words (not what to say) (default: none).
speaker_voice_embedding_file_urlA cloned voice: the URL of the speaker embedding
fal-ai/qwen-3-tts/clone-voicemade; it overridesvoiceandprompt(default: none).reference_textWhat the cloned voice's recording says, which helps it sound like it (default: none).
temperatureHow varied the delivery is, above 0 and up to 1 (default 0.9).
top_kSampling: from the k likeliest sounds (default 50).
top_pSampling: from the likeliest sounds that make up p of the chance, 0 to 1 (default 1).
repetition_penaltyHow strongly it avoids repeating itself (default 1.05).
subtalker_dosampleWhether its second stage samples (default True).
subtalker_temperatureIts second stage's temperature, 0 to 1 (default 0.9).
subtalker_top_kIts second stage's top k (default 50).
subtalker_top_pIts second stage's top p, 0 to 1 (default 1).
max_new_tokensThe most sound it makes, in its codec's tokens, 1 to 8192 (default 8192 here: fal's own, 200, can cut a long line short).
voice ¶
The voice (default: the model's).
language ¶
The language (default "Auto": the text's).
prompt ¶
How to say it, in words (not what to say) (default: none).
speaker_voice_embedding_file_url ¶
A cloned voice: the URL of the speaker embedding fal-ai/qwen-3-tts/clone-voice
made; it overrides voice and prompt (default: none).
reference_text ¶
What the cloned voice's recording says, which helps it sound like it (default: none).
temperature ¶
How varied the delivery is, above 0 and up to 1 (default 0.9).
top_k ¶
Sampling: from the k likeliest sounds (default 50).
top_p ¶
Sampling: from the likeliest sounds that make up p of the chance, 0 to 1 (default 1).
repetition_penalty ¶
How strongly it avoids repeating itself (default 1.05).
subtalker_dosample ¶
Whether its second stage samples (default True).
subtalker_temperature ¶
Its second stage's temperature, 0 to 1 (default 0.9).
subtalker_top_k ¶
Its second stage's top k (default 50).
subtalker_top_p ¶
Its second stage's top p, 0 to 1 (default 1).
max_new_tokens ¶
The most sound it makes, in its codec's tokens, 1 to 8192 (default 8192 here: fal's own, 200, can cut a long line short).
Captions¶
add_subcaption ¶
Caption the film: content shown from now plus offset, for duration seconds.
The film's captions hold it, with those of its
speech; manimgx render writes them beside the video, as subtitles.
contentThe caption's text.
durationHow long it shows, in seconds.
offsetHow long after now it begins, in seconds.
Source
src/manimgx/scene.py
A voice of your own¶
A voice is any function from a text to a Speech: a sound that knows its words, and when each is said. Any text-to-speech service can be one.
Voice ¶
Anything that speaks: a function from text to Speech. Bring any
text-to-speech: wrap its call in a function that returns Speech(audio, text=text).
Word ¶
timed ¶
The words of text, timed by a voice's timed pieces of it, in order (its characters,
or its tokens): each word from the start of the piece its first letter is in to the end
of the piece its last letter is in. Empty if the pieces don't spell the text's letters
(a voice that said "two" for "2"): its words are then estimated.
Source
src/manimgx/audio/sound.py
estimate ¶
When each word of text is said in samples, estimated: the audio's voiced span
shared among the words by their letters, with a share more for the pause after a comma or
a full stop. Against voices' own times, it is off by about 0.1 s on average.
Source
src/manimgx/audio/sound.py
cached ¶
What voice says for text: spoken once, then kept in folder as its audio and a
JSON of its words, named by the text and a digest of the voice and the text.
The audio may be replaced (a recording of the same words, in any format the engine reads, under the same name): its words are then timed again, from it.
Source
src/manimgx/audio/sound.py