Skip to content
Get started
GENERATION API CALLS
Video Generation

MiniMax

This page is auto-generated from model configurations. Last updated: 2026-09-03.

This reference lists all available MiniMax video generation models and their parameters. Use these parameter names when calling the Generation API.


A cinematography model offering advanced camera control for cinematic storytelling and professional framing.

Model ID: model_minimax-video-01-director

Capabilities: txt2video, img2video

LLM Markdown: https://app.scenario.com/api/models/model_minimax-video-01-director/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringYes----Describe your video. Try Prompt Spark or check our Help Center for detailed prompt tips for this model.
firstFrameImagefileNo----Image used as the first frame of the video.
promptOptimizerbooleanNotrue---Use prompt optimizer

MiniMax H3 2K / 24 fps video generation with native stereo audio: text, first/last frame, or multimodal references (images, videos, optional audio).

Model ID: model_minimax-h3

Capabilities: txt2video, img2video, video2video

LLM Markdown: https://app.scenario.com/api/models/model_minimax-h3/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringYes----Describe the video you want.
firstFrameImagefileNo----An image to use as the video’s opening frame. The video’s shape follows this image. Can’t be combined with reference images, videos, or audio.
lastFrameImagefileNo----An image to use as the video’s closing frame. Only works when a first frame is also set. Can’t be combined with reference images, videos, or audio.
referenceImagesfile_arrayNo----Images to guide the subjects or style (up to 9, max 30 MB each). Can’t be combined with a first or last frame. First 5 are free; additional images are billed.
referenceVideosfile_arrayNo----Videos to guide the motion (up to 3, each 2–15s and 2–15s in total, max 50 MB each). Can’t be combined with a first or last frame. Billed per second of uploaded duration.
referenceAudiofile_arrayNo----Audio to guide the video (up to 3, each 2–15s and 2–15s in total, max 15 MB each). Requires at least one reference image or video.
durationnumberNo5515-How long the video lasts, in seconds (5-15). Longer videos cost more.
resolutionstringNo2K--768P, 2KThe output quality. Higher resolutions cost more.
aspectRatiostringNoadaptive--21:9, 16:9, 4:3, 1:1, 3:4, 9:16, adaptiveThe shape of the video. Auto picks a fitting shape. Ignored when a first frame is set, since the shape follows that image.

Fal’s post-trained MiniMax H3 Max image-to-video: animate a first frame (optional last frame), 5–15s at 480p or 768p.

Model ID: model_minimax-h3-max-i2v

Capabilities: img2video

LLM Markdown: https://app.scenario.com/api/models/model_minimax-h3-max-i2v/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringYes----Describe the video you want.
firstFrameImagefileYes----An image to use as the video’s opening frame. The video’s shape follows this image.
lastFrameImagefileNo----An image to use as the video’s closing frame. Only works when a first frame is also set.
durationnumberNo5515-How long the video lasts, in seconds (5-15). Longer videos cost more.
resolutionstringNo768P--480P, 768PThe output quality. Higher resolutions cost more.
promptExpansionModestringNobalanced--disabled, balanced, qualityHow much to rewrite the prompt before generation. Disabled skips it. Balanced takes about a second. Quality spends up to ~30s on a richer prompt.
seednumberNo----Use a seed for reproducible results. Leave blank for a random seed.

MiniMax H3 Max reference-to-video via Fal: stronger prompt adherence and aesthetics from images, videos, and optional audio, with native stereo sound at 480P or 768P.

Model ID: model_minimax-h3-max-reference-to-video

Capabilities: img2video, video2video

LLM Markdown: https://app.scenario.com/api/models/model_minimax-h3-max-reference-to-video/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringYes----Describe the video. Refer to uploads by order: Image 1, Image 2, Video 1, Audio 1, and so on.
referenceImagesfile_arrayNo----Images to guide subjects or style (up to 9, max 30 MB each). Refer to them as Image 1, Image 2, and so on. Combined images, videos, and audio must not exceed 12 files. First 4 (1024×1024) are included; additional images are billed.
referenceVideosfile_arrayNo----Videos to guide motion (up to 3, each 2–15s and 2–15s in total, max 50 MB each). Refer to them as Video 1, Video 2, and so on. Billed per second of uploaded duration at the output resolution.
referenceAudiofile_arrayNo----Audio to guide the video (up to 3, each 2–15s and 2–15s in total, max 15 MB each). Requires at least one reference image or video.
durationnumberNo5515-How long the video lasts, in seconds (5–15). Longer videos cost more.
resolutionstringNo768P--480P, 768PThe native generation resolution. Reference video billing depends on this value.
aspectRatiostringNoadaptive--21:9, 16:9, 4:3, 1:1, 3:4, 9:16, adaptiveThe shape of the video. Auto picks a fitting shape.
promptExpansionModestringNobalanced--balanced, qualityHow much to rewrite the prompt before generation. Balanced returns in about a second; quality spends up to ~30s on a richer prompt.
seednumberNo-02147483647-Optional seed for reproducible results. A random seed is used when omitted.

Fal’s post-trained MiniMax H3 Max text-to-video: stronger prompt adherence and aesthetics, 5–15s at 480p or 768p.

Model ID: model_minimax-h3-max-t2v

Capabilities: txt2video

LLM Markdown: https://app.scenario.com/api/models/model_minimax-h3-max-t2v/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringYes----Describe the video you want.
durationnumberNo5515-How long the video lasts, in seconds (5-15). Longer videos cost more.
resolutionstringNo768P--480P, 768PThe output quality. Higher resolutions cost more.
aspectRatiostringNo16:9--21:9, 16:9, 4:3, 1:1, 3:4, 9:16The shape of the video.
promptExpansionModestringNobalanced--disabled, balanced, qualityHow much to rewrite the prompt before generation. Disabled skips it. Balanced takes about a second. Quality spends up to ~30s on a richer prompt.
seednumberNo----Use a seed for reproducible results. Leave blank for a random seed.

Fal’s post-trained MiniMax H3 Max Turbo image-to-video: faster inference with stronger prompt adherence, 5–15s at 480p or 768p.

Model ID: model_minimax-h3-max-turbo-i2v

Capabilities: img2video

LLM Markdown: https://app.scenario.com/api/models/model_minimax-h3-max-turbo-i2v/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringYes----Describe the video you want.
firstFrameImagefileYes----An image to use as the video’s opening frame. The video’s shape follows this image.
lastFrameImagefileNo----An image to use as the video’s closing frame. Only works when a first frame is also set.
durationnumberNo5515-How long the video lasts, in seconds (5-15). Longer videos cost more.
resolutionstringNo768P--480P, 768PThe output quality. Higher resolutions cost more.
promptExpansionModestringNobalanced--disabled, balanced, qualityHow much to rewrite the prompt before generation. Disabled skips it. Balanced takes about a second. Quality spends up to ~30s on a richer prompt.
seednumberNo----Use a seed for reproducible results. Leave blank for a random seed.

Fal’s post-trained MiniMax H3 Max Turbo text-to-video: faster inference with stronger prompt adherence, 5–15s at 480p or 768p.

Model ID: model_minimax-h3-max-turbo-t2v

Capabilities: txt2video

LLM Markdown: https://app.scenario.com/api/models/model_minimax-h3-max-turbo-t2v/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringYes----Describe the video you want.
durationnumberNo5515-How long the video lasts, in seconds (5-15). Longer videos cost more.
resolutionstringNo768P--480P, 768PThe output quality. Higher resolutions cost more.
aspectRatiostringNo16:9--21:9, 16:9, 4:3, 1:1, 3:4, 9:16The shape of the video.
promptExpansionModestringNobalanced--disabled, balanced, qualityHow much to rewrite the prompt before generation. Disabled skips it. Balanced takes about a second. Quality spends up to ~30s on a richer prompt.
seednumberNo----Use a seed for reproducible results. Leave blank for a random seed.

Model ID: model_minimax-hailuo-02

Capabilities: txt2video, img2video

LLM Markdown: https://app.scenario.com/api/models/model_minimax-hailuo-02/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringYes----Describe your video
firstFrameImagefileNo----Image used as the first frame of the video. The output video will have the same aspect ratio as this image.
lastFrameImagefileNo----Used to generate a video that transitions from the first frame to this image. Requires a first frame image.
durationnumberNo6--6, 10Duration of the video in seconds. 10 seconds is only available for 768p resolution.
resolutionstringNo1080p--768p, 1080pPick between standard 768p, or pro 1080p resolution. The pro model is not just high resolution, it is also higher quality.
promptOptimizerbooleanNotrue---Use prompt optimizer

A high-fidelity video generation model optimized for realistic human motion, cinematic VFX, expressive characters, and strong prompt and style adherence across both text-to-video and image-to-video workflows

Model ID: model_minimax-hailuo-2-3

Capabilities: txt2video, img2video

LLM Markdown: https://app.scenario.com/api/models/model_minimax-hailuo-2-3/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringYes----Text prompt for generation
firstFrameImagefileNo----First frame image for video generation. The output video will have the same aspect ratio as this image.
durationnumberNo6--6, 10Duration of the video in seconds. 10 seconds is only available for 768p resolution.
resolutionstringNo768p--768p, 1080pPick between 768p or 1080p resolution. 1080p supports only 6-second duration.
promptOptimizerbooleanNotrue---Use prompt optimizer

A lower-latency image-to-video version of Hailuo 2.3 that preserves core motion quality, visual consistency, and stylization performance while enabling faster iteration cycles.

Model ID: model_minimax-hailuo-2-3-fast

Capabilities: txt2video, img2video

LLM Markdown: https://app.scenario.com/api/models/model_minimax-hailuo-2-3-fast/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringYes----Text prompt for generation
firstFrameImagefileNo----First frame image for video generation. The output video will have the same aspect ratio as this image.
durationnumberNo6--6, 10Duration of the video in seconds. 10 seconds is only available for 768p resolution.
resolutionstringNo768p--768p, 1080pPick between 768p or 1080p resolution. 1080p supports only 6-second duration.
promptOptimizerbooleanNotrue---Use prompt optimizer

Minimax Video-01 is a versatile model generating high-quality videos with 720p resolution and smooth motion.

Model ID: model_minimax-video-01

Capabilities: txt2video, img2video

LLM Markdown: https://app.scenario.com/api/models/model_minimax-video-01/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringYes----Describe your video
firstFrameImagefileNo----Image used as the first frame of the video.
subjectReferencefileNo----An optional character reference image to use as the subject in the generated video
promptOptimizerbooleanNotrue---Use prompt optimizer