Ltx 2.3 problem with dialogs

Good afternoon, we need the help of experienced specialists. I’m already exhausted. When generating videos, the neural network constantly distorts the accents, although I clearly put in large letters in which places it is necessary to speak with pleasure. But it doesn’t help. Does anyone know how to defeat this or what to prescribe in a negative draft so that the system starts to follow the instructions clearly?

I’m not an experienced specialist, but assuming this is about LTX-2.3’s built-in audio generation, I wonder if the prompt may be specifying the speech in a slightly mismatched way:


Short answer

If by “accent” you mean word stress / emphasis / prosody — for example, “make this word sound stronger” — then capital letters alone are probably too weak.

Instead of treating it like typography:

"I will NEVER go back."

I would try writing it as speech performance direction:

The woman says slowly and clearly, "I will never go back." She pauses briefly before the word "never", then says "never" more firmly and slightly louder than the rest of the sentence. The audio is crisp close-mic speech with quiet room tone.

If by “accent” you mean a regional accent — for example British, French, Russian, American, etc. — that is a harder problem. You can specify it in the prompt, but if exact accent or pronunciation is important, prompt-only generation may not be reliable enough. Reference audio, audio-to-video, LipDub, or replacing/cleaning the final audio may be more practical.

Why I would not start with the negative prompt

For this specific problem, I would not start by asking “what should I put in the negative prompt?”

A negative prompt is usually better for removing broad unwanted concepts. But here the issue is probably not just “avoid bad accent.” It is more like:

  • which word should be stressed
  • where the pause should be
  • how fast the line should be spoken
  • whether the voice should be calm, angry, hesitant, loud, soft, tense, etc.
  • whether “accent” means regional pronunciation or just emphasis

So I would first improve the positive prompt: describe the speech you want, not only the speech you do not want.

“Accent” may mean two different things

This is important because the fix depends on the meaning.

If “accent” means… Example Better first approach
Regional accent British accent, French accent, Russian accent Specify language, regional accent, and voice style, but expect limited reliability. Use reference audio if exact fidelity matters.
Word stress / emphasis / prosody Make “never” stronger; pause before a word; say one phrase more firmly Do not rely only on capital letters. Describe the delivery: pause, pace, volume, firmness, emotional beat, and the exact word to emphasize.
Pronunciation of a specific word A name, acronym, unusual word, technical term Rewrite the phrase more clearly, but exact pronunciation may require reference audio, external TTS, or post-production.
Wrong speaker / wrong dialogue assignment The wrong character speaks; two people speak at once Use clearer speaker references, shorter dialogue chunks, and check whether the Space/ComfyUI workflow adds or rewrites prompt text.

LTX-style prompts seem closer to directing a performance

I would not treat this as a universal rule, because prompt behavior ultimately depends on the model weights, training data, text encoder, conditioning design, guidance settings, sampler, seed, and sometimes the UI or workflow around the model.

But different models do have practical prompting habits.

For LTX-2.3, the official prompt guide strongly points toward a cinematic / performance-direction style rather than a short keyword or typography-based style. The guide says LTX-2.3 responds better to long, specific prompts, and it specifically mentions facial expressions, timing, pauses, emotional beats, voice qualities, and breaking dialogue into short phrases with acting directions between them.

Useful references:

The LTX documentation also says to clearly describe audio, including ambient sound, music, speech, or singing, and to put spoken dialogue in quotation marks. It also says to specify language and accent if needed.

So for LTX, I would think less like this:

keyword, keyword, keyword, BIG EMPHASIS, no bad accent

and more like this:

A short cinematic close-up. The woman speaks English in a calm but tense voice. She says, "I will never go back." She pauses briefly before "never", then says that word more firmly and slightly louder. Her jaw tightens as she finishes the sentence. The audio is crisp close-mic speech with quiet room tone.

Why capital letters are probably not enough

Capital letters are a text convention. They may sometimes help a model infer emphasis, but they do not fully specify speech delivery.

Even in ordinary TTS systems, prosody is often controlled with explicit mechanisms such as pauses, pitch, rate, volume, pronunciation, and emphasis. For example, SSML exists partly because speech synthesis often needs structured controls for pronunciation, volume, pitch, rate, and related properties. Google’s Text-to-Speech SSML documentation similarly describes controls such as prosody, pitch, speaking rate, volume, breaks, and pronunciation-related markup.

LTX-2.3 is not just an ordinary TTS engine, though. It is an audio-video generation model. The LTX-2 paper describes LTX-2 as a joint audio-visual model with video and audio streams connected by cross-modal attention. That means the speech can be affected by the scene, character, timing, facial motion, acoustic environment, and video prompt — not only by the quoted text.

So I would avoid relying on typography alone.

Better prompt patterns to try

1. For word emphasis

Instead of:

The man says: "This is VERY important."

Try:

The man speaks slowly and seriously. He says, "This is very important." He places the strongest emphasis on the word "very", saying it slightly louder and more firmly than the rest of the sentence.

Or:

The man says, "This is..." He pauses briefly, then says "very important" with a firmer tone and stronger emphasis on "very". The audio is clear speech with quiet room tone.

2. For a pause before a key word

Instead of:

The woman says: "I will NEVER go back."

Try:

The woman says softly, "I will..." She pauses, looks directly at him, and her jaw tightens. She then says "never" more firmly and slightly louder: "never go back." The audio is crisp, close-mic speech with quiet room tone.

3. For regional accent

Instead of:

The woman speaks with an accent.

Try:

The woman speaks English with a light French accent, in a calm conversational voice. She says: "I do not think we should go back." The audio is crisp, with quiet room tone.

But I would not expect perfect control here. Regional accent is not just one simple attribute. It involves pronunciation, rhythm, vowel/consonant patterns, speaking style, and sometimes speaker identity. If the exact accent matters, prompt-only control may not be the best tool.

4. For clearer dialogue timing

Instead of putting a long line in one block:

She says, "So anyway, we went to watch the movie, and spoiler, someone dies at the end", then she laughs.

Try splitting it into shorter beats:

The woman speaks in a casual YouTuber-style close-up. She says, "So anyway, we went to watch the movie." She pauses and hides a small laugh. Then she continues, "And spoiler: someone dies at the end." The audio is clear speech with natural room tone.

This style is closer to what the LTX-2.3 prompt guide suggests: short dialogue chunks, acting directions, and clear physical cues.

A useful mental model

I would use this mental model:

Capital letters are typography. LTX-style prompting is closer to directing an actor.

So instead of asking only:

How do I mark the accented word?

ask:

How should the character perform the line?

Then describe:

  • who is speaking
  • the exact quoted dialogue
  • the language
  • regional accent, if relevant
  • the word to emphasize
  • pause placement
  • pace
  • volume
  • firmness / softness
  • emotional beat
  • facial or body cue
  • acoustic environment

For example:

A tight close-up of a tired man sitting in a quiet kitchen at night. He speaks English in a low, controlled voice. He says, "I told you..." He pauses, looks down, then continues more firmly: "I am not going back." He places the strongest emphasis on "not", saying it slightly louder and slower than the surrounding words. The audio is crisp close-mic speech with quiet room tone and no dramatic music.

What I would try first

I would test in this order:

  1. Clarify the meaning of “accent”

    • Regional accent?
    • Word stress?
    • Pronunciation?
    • Speaker voice?
  2. Use positive speech direction

    • Do not start with negative prompt.
    • Describe the desired delivery directly.
  3. Quote the dialogue

    • Keep spoken words inside quotation marks.
  4. Split long dialogue into short beats

    • Especially if the line contains pauses, laughter, hesitation, or emphasis.
  5. Describe the emphasis in natural speech terms

    • “pauses before the word”
    • “says it slightly louder”
    • “says it more firmly”
    • “slows down on that word”
    • “voice becomes sharper/softer/tenser”
  6. Describe the audio environment

    • “crisp close-mic speech”
    • “quiet room tone”
    • “subtle natural ambience”
    • “no dramatic music” only if needed, but focus mainly on what you want to hear
  7. If exact accent matters, use a stronger control surface

    • reference audio
    • audio-to-video
    • LipDub
    • external TTS
    • post-production audio replacement or cleanup

If you are using a Space, ComfyUI workflow, or wrapper

One more thing: if this is not the raw model call, the prompt may be passing through another layer.

For example, a Space, workflow, prompt enhancer, default text node, or UI wrapper may add or rewrite text before LTX sees it. So if the model seems to ignore a very short instruction, it may be worth checking:

  • whether a prompt enhancer is enabled
  • whether default prompt text is being appended
  • whether the workflow separates visual prompt and audio prompt
  • whether the workflow has audio guidance settings
  • whether the model is actually receiving the exact dialogue you typed

This matters because LTX prompting often appears to work best when the final prompt is a complete scene description, not a tiny isolated speech instruction.

When prompt-only may not be enough

If you need the voice to have a specific regional accent, exact pronunciation, or consistent speaker identity, I would not rely only on text prompting.

There are LTX-related routes that are closer to that kind of control, such as:

Those are probably more relevant when the problem is not just “make this word stronger,” but “make this voice/accent/pronunciation stay consistent.”

Summary

My guess is:

  • LTX-2.3 does have built-in audio generation, so this is not automatically an external TTS issue.
  • But “accent” is ambiguous.
  • If you mean word emphasis, capitalization alone is probably too weak.
  • Try writing the desired speech as performance direction: quoted dialogue, short phrases, pauses, pace, volume, firmness, and the exact word to emphasize.
  • If you mean regional accent or exact pronunciation, prompt-only control may be unreliable, and reference audio / audio-to-video / LipDub / post-production may be more realistic.
  • I would improve the positive prompt first before trying to solve this with the negative prompt.

Thank you very much. for such a detailed answer, now I’ll dig even deeper) By accent, I mean the emphasis on a specific letter in a word. The language is Russian. ltx surprises me, sometimes he reads in perfect Russian without any hints with big letters in sentences, and sometimes he stumbles on simple words. Example: I needed him to say the word “imagine” in Russian and indicate which letter in the word should be stressed. No matter what I did, he was still saying the wrong thing. And it’s quite clear to me why the system with a capital letter works perfectly in suno, but there are such difficulties?

Ah, so that is the kind of “accent” you mean. That clarification helps. In that case, whether this is controllable depends quite a lot on LTX itself, but I think the practical answer is roughly below. Also, the fact that Suno, LTX, and other models have different “dialects” for this kind of control is actually a pretty fundamental issue:


Short answer

If the goal is correct Russian word stress, I would treat this less as a normal “negative prompt” problem and more as a pronunciation-control problem.

For LTX, I would split the answer into two levels:

Route My rough expectation
Text prompt only Worth testing, but probably unreliable for exact Russian stress
Text prompt with stress mark + syllable hint + explicit pronunciation instruction Better, but still not guaranteed
Generate/record correct Russian audio first, then use LTX audio-to-video / image-audio-to-video / custom-audio I2V Much more promising
Final mux/post-production with verified audio Most reliable for pronunciation correctness

So the practical summary is:

Do not treat capital letters as a universal pronunciation API.
For LTX text-only prompting, try stress marks, syllable hints, and explicit delivery instructions.
But if Russian stress must be correct, make the correct audio first and use LTX as an audio-driven video generator.

That audio-first route is not just theoretical. LTX has official audio-to-video API support, and the LTX/ComfyUI community already has workflows around custom audio, image-audio-to-video, Qwen/Fish-style TTS, and voice-driven talking video.

Why Suno and LTX can behave differently

This is the core issue.

Different generative models learn different informal “control languages.”

Suno is strongly connected to music, lyrics, song structure, vocal emphasis, line breaks, and lyric formatting. In that world, things like:

ALL CAPS
line breaks
repeated syllables
elongated words
[spoken]
[whispered]
[chorus]

can become useful signals because the model has likely seen many examples where typography and lyric formatting correlate with vocal delivery.

LTX is different. LTX-2.3 is described as a diffusion-based audio-video foundation model that generates synchronized video and audio in a single model. It is not simply a lyrics-to-song model or a normal standalone TTS engine. Its speech is entangled with:

  • the person on screen
  • mouth movement
  • facial expression
  • camera motion
  • scene timing
  • environment
  • emotion
  • ambient sound
  • video/audio synchronization

So the same trick can work differently:

Model family Likely stronger signal
Song / lyric model Typography, line breaks, lyric structure, repeated syllables
Ordinary TTS text normalization, lexicon, SSML, phonemes, voice settings
LTX-style audio-video model scene description, quoted dialogue, audio prompt, timing, character action, reference image/audio
Lip-sync model input audio waveform, face crop/identity, mouth motion constraints

That is why this is not just a “bad prompt” issue. The model may simply not have learned the same convention that Suno learned.

Russian stress is especially difficult

Russian word stress is not a simple typographic effect.

In normal Russian writing, stress marks are usually omitted. The model often sees a plain word and must infer where the stress should be. That can depend on lexical knowledge, word form, context, and training data coverage.

For example, learning materials may write stress like this:

предста́вь
вообрази́

But ordinary text usually does not include those marks. So if the model sees:

представь

it has to know the pronunciation from memory. If it does not know it reliably, capitalization may not fix the problem.

This is very different from simply saying:

Say this word louder.

Russian lexical stress is closer to:

Pronounce this specific word with the correct stressed vowel.

That is why a text-only video model may be inconsistent: sometimes it knows the word, sometimes it guesses, sometimes it follows visual/timing constraints more strongly than the intended stress hint.

Text-only LTX prompt: what is still worth trying

I would still test text-only prompting, but I would treat it as experimental.

Instead of only doing:

предстАвь

I would try redundant hints:

A close-up of one man in a quiet room, speaking directly to the camera. He speaks Russian slowly and clearly. He says only one word: "предста́вь". The stress is on the vowel "а", pronounced пред-СТАВЬ. The stressed syllable is slightly longer and louder than the others. The audio is crisp close-mic Russian speech with quiet room tone and no music.

This gives the model several different signals:

Signal type Example
Normal Russian spelling представь
Stress mark предста́вь
Syllable/stress hint пред-СТАВЬ
Natural-language explanation “The stress is on the vowel а
Acoustic explanation “slightly longer and louder”
Delivery instruction “slowly and clearly”
Scene simplification one speaker, quiet room, close-up

This may still fail, but it is a more LTX-like prompt than just capitalizing one letter.

Why negative prompt probably will not solve it

I would not start with:

negative prompt: wrong stress, bad Russian pronunciation, incorrect accent

That does not tell the model what the correct pronunciation is.

For this kind of problem, a positive target is more useful:

The stress is on the vowel "а": пред-СТАВЬ.
The stressed syllable is slightly longer and louder.
He says the word slowly and clearly in Russian.

A negative prompt can suppress broad unwanted things. But Russian word stress is not a broad unwanted artifact. It is a specific pronunciation target.

The more promising route: generate the Russian audio first

For exact Russian stress, I would probably move the pronunciation problem out of the LTX text prompt.

A better production-style route is:

Russian-capable TTS / voice clone / human recording
        ↓
verify stress and pronunciation
        ↓
LTX audio-to-video / image-audio-to-video / custom-audio I2V
        ↓
optional mux/post-production with verified audio

This is not just a workaround. It matches LTX’s strengths better.

LTX’s audio-to-video API is explicitly designed to generate video driven by an audio track. The documentation says you can supply dialogue, music, or ambient sound, and the model produces visuals synchronized to the audio. LTX’s Audio-to-Video capability page also describes audio as the primary conditioning signal, where voice, music, and sound drive motion, pacing, and scene structure.

That is exactly the kind of control surface you want when pronunciation matters.

Instead of asking LTX:

Please infer the correct Russian stress from text.

you give it:

Here is the already-correct Russian audio. Generate the visual performance around it.

That is a much stronger signal.

Practical precedent: people are already doing this

There are already LTX community workflows that look very close to this route.

For example:

  • In a Kijai/LTX2.3_comfy discussion, RuneXX shared an I2V & T2V with Custom Audio workflow described as “Use your own audio files with lip sync, and synced motion.”
  • In the same ecosystem, users discuss workflows where they first generate or clone voice with Qwen TTS, then use LTX I2V with custom audio.
  • Comfy has an LTX-2.3 Image Audio to Video workflow where a portrait image and audio file are used to create a lip-synced talking video.
  • There are also LTX/Comfy workflows combining Qwen TTS, Fish Audio, or other TTS/voice-cloning tools with LTX image/audio video generation.

So the route is not only theoretical:

TTS / voice clone / verified recording
        ↓
LTX custom audio / IA2V / A2V
        ↓
talking video

is already a practical pattern in the LTX 2.3 ComfyUI ecosystem.

Important caveat: custom audio is not always automatic lip-sync

I would still be careful.

Providing custom audio does not always guarantee perfect lip-sync. Some users report cases where the audio is present in the output but does not properly drive the mouth, or behaves more like voice-over narration.

A useful practical trick from the LTX ComfyUI ecosystem is:

Provide both the audio file and a transcript/description of the spoken line in the prompt.

For example, if your audio says:

Предста́вь, что это правда.

then the prompt should not only say:

A man speaks Russian.

It should say something like:

A close-up of a man speaking directly to the camera in Russian. He says: "Предста́вь, что это правда." His mouth movements are synchronized to the provided audio. The scene is quiet, with clear close-mic speech and no music.

This helps the model understand that the audio is meant to be character dialogue, not just background narration or ambience.

Recommended audio-first workflow

If I were trying to get correct Russian stress in LTX, I would test this pipeline.

Step 1 — Make the Russian audio outside LTX

Use one of:

  • a Russian-capable TTS
  • a voice-cloning TTS
  • a human recording
  • a manually edited recording
  • a TTS system with SSML/phoneme/lexicon controls, if available

The important point is: verify the pronunciation before giving it to LTX.

Step 2 — Keep the audio simple

For the first test:

  • one speaker
  • short phrase
  • no background music
  • no echo
  • no heavy reverb
  • clean volume
  • clear Russian speech

Do not start with a long dramatic scene.

Step 3 — Use LTX audio-to-video or image-audio-to-video

Use the verified audio as the main conditioning signal.

If using an image:

  • visible face
  • visible mouth
  • not too stylized
  • not too side-profile
  • stable lighting
  • one speaker only

Step 4 — Put the transcript in the prompt

Example:

A close-up portrait of one man speaking Russian directly to the camera. He says: "Предста́вь, что это правда." His mouth movements are synchronized to the provided audio. The delivery is calm and clear. The audio is close-mic Russian speech with quiet room tone.

Step 5 — Check whether LTX preserves or changes the audio

Depending on the workflow, the generated output audio may not be exactly the same as your verified source audio.

If the pronunciation is correct in the source audio but degraded in the output, then simply mux the verified audio back into the final video.

Conceptually:

generated_video.mp4 + verified_russian_audio.wav -> final_video.mp4

Step 6 — If lip-sync is weak, simplify before changing models

Try:

  • shorter clip
  • clearer face
  • more frontal portrait
  • less camera motion
  • no second speaker
  • no music
  • transcript in prompt
  • different seed
  • different workflow version

Only after that would I move to a dedicated lip-sync model.

When to use dedicated lip-sync tools

If LTX gives a good video but poor mouth movement, a dedicated lip-sync step may be better.

Tools in this category include:

These tools solve a narrower problem:

given video face + given audio -> lip-synced face video

That is narrower than LTX’s job:

scene + character + motion + audio + camera + style -> full audiovisual generation

So if the visual scene is already good and only the mouth timing is wrong, a dedicated lip-sync tool may be the better final step.

Audio-to-audio / voice conversion

Audio-to-audio or voice conversion may also be useful, but I would separate it from pronunciation correction.

Voice conversion is useful when the issue is:

  • voice identity
  • timbre
  • speaker style
  • accent color
  • emotional tone
  • making one generated voice sound more like another voice

But for Russian lexical stress, I would not rely on voice conversion as the main fix.

If the source audio has the wrong stress, a voice converter may preserve the wrong stress. It may change the voice color while keeping the same pronunciation error.

So for this problem, I would prioritize:

correct Russian audio first

then use voice conversion only if needed:

correct Russian audio
        ↓
optional voice conversion / voice cloning
        ↓
LTX audio-to-video / lip-sync

Why this is better than forcing text prompting

The reason is simple:

Task Best control signal
Correct Russian stress verified Russian audio
Speaker voice identity reference audio / voice clone
Face and video generation LTX image/video prompt
Mouth timing audio-driven generation or lip-sync
Final audio correctness mux verified audio back in

Text prompt alone asks LTX to solve too many tasks at once:

Read Russian correctly
infer word stress
generate voice
generate face
generate mouth motion
generate scene
align audio and video
follow camera/action prompt

Audio-first separates the tasks:

TTS/recording handles pronunciation.
LTX handles audiovisual performance.
Post-production handles final audio correctness.

That is usually more controllable.

Minimal LTX text-only fallback

If you still want to test text-only LTX first, I would use a minimal diagnostic prompt:

A close-up of one man in a quiet room, speaking directly to the camera. He speaks Russian slowly and clearly. He says only one word: "предста́вь". The stress is on the vowel "а", pronounced пред-СТАВЬ. The stressed syllable is slightly longer and louder than the others. The audio is crisp close-mic Russian speech with quiet room tone and no music.

If that works, gradually add complexity.

If that fails, I would not spend too much time on capitalization tricks. I would switch to the audio-first route.

Suggested production route

For your case, I would rank the routes like this:

Rank Route Why
1 Russian TTS / recording → verify stress → LTX audio-to-video or IA2V Best match for pronunciation-sensitive generation
2 Russian TTS / recording → LTX video → mux verified audio back Best if LTX changes/degrades audio
3 Russian TTS / recording → LTX video → dedicated lip-sync cleanup Best if mouth movement is weak
4 Text-only prompt with stress mark + syllable hints Worth trying, but not robust
5 Negative prompt Probably least useful for this exact problem

My current guess

My guess is:

  • Suno may respond to capitalization because it has learned lyric/music formatting conventions.
  • LTX text-only prompting may not treat capitalization inside Russian words as a reliable pronunciation-control marker.
  • Russian stress is hard because it is usually not written in ordinary spelling.
  • LTX is especially complicated because speech is generated together with video, face motion, timing, and scene context.
  • The best practical route is to stop making LTX infer the pronunciation from text.
  • Create the correct Russian audio first, then use LTX’s audio-driven workflow.

So I would say:

If exact Russian stress matters, do not make typography carry the whole burden.
Use text-only LTX prompting as a quick experiment, but use verified audio as the real control signal.