MiniMax H3 Microexpression Prompts: Direct Subtle, Natural Emotion

E
Emma Chen·9 min read·Sep 10, 2026
Share on X
MiniMax H3 Microexpression Prompts: Direct Subtle, Natural Emotion

AI Overview

How do I write a MiniMax H3 microexpression prompt?

Name the emotion, the visible facial evidence, the trigger, and the moment it settles. Use one restrained change—such as eyes softening and a jaw releasing—rather than a stack of dramatic emotion words.

Why does an H3 character look blank or robotic?

The prompt may describe an internal feeling without an observable action, or ask for too many competing beats. Give the face a cause, one readable response, natural blinking and breathing, then specify a neutral resting state.

Can MiniMax H3 copy an expression from a reference?

Reference mode can use images or video to communicate identity and performance signals. Assign each file one explicit role, separate face identity from acting reference, and verify that expression transfer does not replace the character.

Should dialogue and facial reactions be prompted together?

Yes, when the reaction is tied to a specific line or pause. Place the expression before, during or after the spoken beat, keep mouth motion compatible with the words, and avoid simultaneous conflicting camera direction.

What a Microexpression Prompt Actually Controls

A microexpression prompt should direct a small change that the viewer can see over time. “She feels nervous” is an acting note, but it does not tell the model which signal carries the feeling. “Her inner eyebrows rise slightly, she presses her lips for a moment, then exhales and lets her jaw relax” gives a sequence of observable actions. The second version is easier to stage, review and revise.

MiniMax H3 accepts text, image, video and audio context, and its official materials describe four-to-fifteen-second outputs at 24 fps. That is enough for a facial beat to emerge, peak and settle, but not for a complicated emotional arc plus several camera moves. Treat the prompt as a compact shot brief: character, trigger, visible response, camera, sound and ending.

A woman on a rain-lit platform holding relief and worry at once

The useful target is not “sad face.” It is a readable combination: moist eyes, a jaw beginning to release, controlled lips and a gaze that remains alert.

Microexpression is a creative direction term here, not clinical emotion detection. Use small visible cues to communicate a scripted feeling while preserving identity, anatomy and lip synchronization.

Build the Expression as a Timed Beat

Start with a neutral baseline. Without one, the model may begin at maximum intensity and have nowhere to move. Next, state the trigger: a train leaves, a name is announced, a message ends, or another character pauses. Then choose two or three compatible facial signals. Finish by describing the new resting state. This creates a four-part beat that can be copied into almost any H3 prompt:

  1. Baseline: composed face, steady gaze, relaxed shoulders.
  2. Trigger: she hears the offscreen answer after a short pause.
  3. Response: her eyes soften, lower eyelids lift slightly, and one corner of her mouth almost rises.
  4. Settle: she exhales quietly and returns to professional calm, with concern still visible in her gaze.

Avoid giving every facial feature a separate command. A stack of brow, eye, nose, jaw, lip and smile instructions is often contradictory. Prioritize one transition and let breathing, posture and gaze support it. Natural blinks work best as occasional punctuation, not a mechanical schedule.

A job candidate trying to hide nervousness outside an interview

For restrained tension, inspect the relationship between brow, lip pressure, grip and posture. The face should not jump to theatrical fear.

Camera distance determines whether the audience can read the cue. Use a close-up or medium close-up for eyelid, mouth-corner and jaw changes. If you need a wide establishing shot, move to the face before the emotional peak or use a separate shot. The MiniMax H3 face-at-distance guide covers cases where the face is too small to carry the performance reliably.

Four Copy-Ready Microexpression Prompt Patterns

These templates are newly written production patterns, not claimed official H3 outputs. Replace the bracketed story details and keep the observable sequence. Each pattern asks for one dominant performance beat.

Concealed nerves: “Medium close-up of [character] waiting outside [place]. The face begins composed. After hearing [trigger], the inner eyebrows lift slightly and the lips press together for less than a second. A shallow breath raises the chest; the fingers loosen around [object]. The expression settles into deliberate calm. Natural blink timing, no grimace, no exaggerated fear. Camera remains still.”

Relief with concern: “Close-up of [character] reading [result]. The eyes brighten first, the shoulders release and a small smile almost appears. Then the gaze returns to [offscreen subject], keeping a trace of concern in the lower eyelids. One quiet exhale. End on controlled professional composure, not celebration.”

A chef registering a small disappointment while tasting soup

A practical prompt names the tasting trigger and a brief response—narrowed eyes, nose tension and controlled lips—before the chef returns to problem-solving.

Private disappointment: “Tight close-up as [character] tastes or inspects [object]. For a brief beat, the nose tightens and the eyes narrow while the mouth stays closed. The character catches the reaction, breathes through the nose and looks back at the task with focused resolve. No disgust face, no head jerk, no comedy.”

Listening reaction during dialogue: “Over-the-shoulder close-up on [listener] while [speaker] delivers the final sentence offscreen. The listener holds eye contact, blinks once after the key word, and one mouth corner lifts while the eyes remain sad. Wait half a beat before the reply. Keep breathing and head motion subtle; preserve lip sync for the response.”

Do not paste all four into one request. Choose the emotional job, adapt the trigger and remove unused language. If the clip includes spoken dialogue, write the line and reaction in playback order. The page about keeping voices consistent offers a useful review method for matching vocal and visual performance across shots, even when your generation model differs.

Choose Mode and Reference Roles

Text-to-video asks H3 to invent the person, setting and performance, so describe appearance only enough to keep the actor readable. First-frame image-to-video already owns the opening face and composition; spend prompt space on motion, expression timing, gaze and what must remain unchanged. First-and-last-frame mode is useful when the final emotion must land on a specific image, but the transition still needs a plausible trigger and restrained path.

Reference-to-video is useful when signals come from separate assets. Assign roles explicitly: Picture 1 owns facial identity, Picture 2 owns wardrobe, Video 1 supplies the restrained listening performance, and Audio 1 supplies voice timbre. State that acting movement and timing transfer without replacing the target person's identity.

A doctor reading a result and returning to professional composure

Mixed emotion becomes readable when the prompt defines order: eyes brighten, shoulders soften, an almost-smile appears, and composure returns.

Use the Ref2V versus FL2VA decision guide before attaching media. More references are not automatically better. Competing face angles, expression peaks or lighting conditions can weaken the signal. For a simple portrait, one strong identity image and one focused acting reference are often easier to diagnose than a folder of similar files.

Fix Blank, Robotic, or Overacted Faces

A blank face usually means the prompt describes mood but not change. Add a trigger and two visible cues. A robotic face often combines fixed staring, perfectly timed blinking and a locked head; restore natural breathing, a small gaze adjustment and asymmetry. An overacted face usually comes from repeated intensifiers such as “extremely terrified, shocked and devastated.” Replace them with one restrained transition and a negative boundary: no grimace, no wide eyes, no sudden head snap.

If the mouth moves strangely, separate silent reaction from speech. Let the listener react during the other person's line, pause, then deliver a short response. If identity changes as emotion peaks, reduce camera travel, strengthen the identity reference and simplify the expression. If the result ignores subtle wording, check the text/vision encoder and reference path; the Qwen encoder precision guide explains how to compare difficult semantic prompts without confusing quantization with another workflow change.

MiniMax H3 reference-to-video motion sample

This existing H3 reference-to-video result is an inspection sample, not a microexpression benchmark. Review the full motion for identity stability, gaze, blink rhythm and whether the face settles naturally.

When testing, keep the input, seed, duration, model, resolution and camera instruction fixed. Change one expression phrase at a time. A useful three-run sequence is baseline, reduced instruction and one clarified trigger. If the shorter prompt works, add only the detail required for the story rather than restoring the entire original stack.

Review the Full Performance in Production

Approve emotion in motion, not from a flattering paused frame. Watch at normal speed with sound, muted to judge timing, and audio-only to judge the vocal beat. Check identity, gaze, blink rhythm, mouth articulation, head motion, shoulder tension and whether the camera stays close enough during the key moment.

An elderly father listening to a voice message with mixed joy and sadness

Mixed feeling does not require a dramatic face: softened eyes, asymmetric mouth movement and calm breathing can carry the scene.

For a multi-shot scene, write an emotional state handoff after each approved take: “guarded before the call,” “relief begins after the name,” and “calm but still uncertain.” Store the chosen reference, prompt, seed and acceptance note with the shot. Seedance Agent can organize those references, plan the close-up at the right beat, collect reviewer decisions and rerun only the failed performance segment instead of regenerating the whole sequence.

Use the MiniMax H3 model page when you want the model and mode in one hosted workflow, or start from the image-to-video workspace when a strong portrait already defines the actor. The creative goal stays the same: a small, motivated facial change that supports the story without calling attention to the generation process.

Conclusion

Effective MiniMax H3 microexpression prompts translate an internal feeling into a short visible timeline: neutral baseline, clear trigger, two or three compatible facial and body cues, and a controlled resting state. Keep the camera close enough, assign each reference one role, place reactions around dialogue in playback order, and remove exaggerated or contradictory instructions. Review the complete clip with and without sound, record the accepted emotional state between shots, and revise one phrase at a time. When reference management, timing, approvals and partial reruns become the harder problem, build the performance workflow with Seedance Agent and keep the attention on the actor's finished moment.

Ready to try it yourself?

Put the steps from this guide into practice with Seedance and turn prompts or images into polished videos in minutes.

Free credits on signup. Plans from $20/month.