MiniMax H3 is a next-generation, open-weights, general-purpose multimodal video model with synchronized stereo sound and opt-in mature-content support; 4-step H3 Turbo delivers about 3.5× faster iteration.
Standard quality or about 3.5× faster Turbo—picture and stereo sound together.
MiniMax H3 is a next-generation, open-weights, general-purpose multimodal video model—not a Seedance derivative—with short-form capabilities comparable to Seedance 2.0. On Sogni, H3 supports opt-in mature-content workflows while remaining subject to Sogni's Terms of Use and MiniMax's license and safety requirements. It creates the moving image and stereo soundtrack in one generation, so you can direct the action, camera, dialogue, ambience, sound effects, and music in the same prompt and get a complete audiovisual shot instead of a silent clip that still needs a soundtrack.
Choose how much control and speed you want. Start from text for a blank canvas, animate a starting image when the composition is already set, provide opening and closing frames when the shot must arrive at a specific visual, or condition a scene on labelled image, video, and audio references. H3 Turbo adds 4-step text-to-video, image-to-video, and first-and-last-frame workflows; Ref2VA remains a standard 20-step workflow.
Use H3 Turbo for drafts and rapid prompt iteration. Sogni's warm reference runs measured about 3.5× faster image-to-video and 3.7× faster text-to-video, conservatively summarized as about 3.5× faster; actual time varies by mode, input, clip length, resolution, worker hardware, and queue state. The upstream v0.1 release is a preview and may trade some fine visual detail and audio polish for speed, so standard H3 remains the quality-first choice.
On Sogni, H3 produces roughly 5–15 seconds of 768p-class video at 24 fps with 32 kHz stereo audio. Sogni Web offers six square, landscape, cinema, and portrait presets; direct SDK/API jobs accept a 32 px grid up to 1344 px per axis within the model's pixel budget. MiniMax reports stable dialogue support in 11 languages.
All seven current H3 modes are available through the Sogni API and Creative Agent Skill: four standard workflows (text-to-video, image-to-video, first-and-last-frame, and Ref2VA) plus three Turbo workflows (text-to-video, image-to-video, and first-and-last-frame). Sogni Web currently exposes standard text-to-video, image-to-video, and Ref2VA. Ref2VA supports up to 9 images, 3 videos, and 3 audio clips (12 files total, with at least one visual reference).
Use standard 20-step H3 when detail and audio polish come first, or a 4-step Turbo FL2VA workflow for about 3.5× faster iteration. Ref2VA multi-reference video currently has no Turbo variant.
| Workflow | Details | Access | |
|---|---|---|---|
| Text to video Standard | Workflow: Create video and audio from a written scene Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: No image needed Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | Create → |
| First frame Standard | Workflow: Animate a starting image with video and audio Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: One opening frame Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | Create → |
| First and last frames Standard | Workflow: Direct the motion between two compositions Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: One opening frame and one closing frame Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | API / agent workflow |
| Multi-reference Standard | Workflow: Condition a scene on labelled image, video, and audio references Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: 0–9 images · up to 3 videos · up to 3 audio clips · 12 files total Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | Create → |
| Turbo text to video Turbo | Workflow: Create video and audio from text with the Turbo path Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: No image needed Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | API / agent workflow |
| Turbo first frame Turbo | Workflow: Animate a starting image with the Turbo path Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: One opening frame Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | API / agent workflow |
| Turbo first and last frames Turbo | Workflow: Connect two anchor images with the Turbo path Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: One opening frame and one closing frame Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | API / agent workflow |
For Creative Agent, write a compact shot brief: what we see, what changes, how the camera moves, and what we hear. The agent expands it into H3's production prompt. Direct SDK users should follow MiniMax's structured prompt guide.
Read the Sogni Engineering field guide, Your Prompt Is Now a Director, for side-by-side H3 and H3 Turbo tests across text-to-video, image-to-video, first-and-last-frame, and reference-to-video workflows.
Choose standard 20-step H3 when fine visual detail and audio polish matter most. Choose H3 Turbo for drafts, prompt iteration, timing tests, and other work where turnaround matters most. Sogni's warm reference runs measured about 3.5× faster image-to-video and 3.7× faster text-to-video, conservatively summarized as about 3.5× faster. Actual time varies with workflow, input conditioning, clip length, resolution, worker hardware, and queue state.
LightX2V labels the current v0.1 Turbo weights a preview and notes that image detail still needs improvement. Turbo is a speed/quality trade-off, not a higher-quality replacement for standard H3.
MiniMax H3 generates picture and synchronized stereo sound together: dialogue, ambience, effects, and music can all be directed in the same brief. It is useful for short brand films, dialogue scenes, product and fashion clips, animated posters and titles, motion design, music visuals, and stylized social video.
Start from text when the scene is still an idea. Supply an opening frame when identity, product shape, palette, or composition is established. Add a closing frame when a transition must land at a specific visual. Standard Ref2VA conditions a scene on labelled uploads such as <Picture 1>, <Video 1>, and <Audio 1>; it accepts up to nine images, three videos, and three audio clips, with no more than 12 files total and at least one visual reference. Turbo currently has no Ref2VA variant.
MiniMax H3 is MiniMax's next-generation open-weights model, with a short-form multimodal scope comparable to Seedance 2.0. Sogni supports opt-in mature-content generation with H3; "uncensored" describes that wider adult creative scope, not an absence of rules. Sogni's Terms of Use and the applicable MiniMax and LightX2V licenses still apply.
Sogni Web supports standard H3 text-to-video, image-to-video, and Ref2VA reference-to-video. The Sogni API and Creative Agent Skill additionally support standard first-and-last-frame video and the exact Turbo t2v, i2v, and flf2v IDs listed above. Sogni generates 768p-class output and does not expose MiniMax's hosted editing or 2K regeneration stages.
Standard H3 costs 16 Spark ($0.08) per second of finished video; H3 Turbo costs 6 Spark ($0.03) per second — about 2.67× cheaper. Within each speed tier, duration is the only price factor: layout, workflow, and reference count do not change the rate. The 24 fps frame grid produces billable clips from roughly 5.17 seconds (82.67 Spark / $0.41 standard, 31 Spark / $0.155 Turbo) to 15.08 seconds (241.33 Spark / $1.21 standard, 90.5 Spark / $0.4525 Turbo).
Review the official MiniMax H3 model card, MiniMax H3 Community License, LightX2V H3 Turbo model card, and LightX2V Turbo implementation. Sogni has received written authorization from MiniMax to offer H3 through the platform.
Standard MiniMax H3 is billed at a flat 16 Spark ($0.08) per second of finished video; H3 Turbo is 6 Spark ($0.03) per second — about 2.67× cheaper. Within each speed tier, cost scales with clip duration only — layout and workflow do not change the per-second rate, and references are never charged extra. Use pay-as-you-go Spark, or eligible fair-use coverage on an active Unlimited plan.
Start with a natural creative brief. Creative Agent expands it into the production prompt MiniMax H3 needs and runs the matching workflow.
const response = await fetch('https://api.sogni.ai/v1/chat/completions', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${process.env.SOGNI_API_KEY}`,
},
body: JSON.stringify({
messages: [{ role: 'user', content: "Create an 8-second MiniMax H3 video from this brief: A ceramic artist opens a glowing kiln as the camera slowly pushes in; fire crackles, tools clink softly, and a restrained string score begins" }],
sogni_tools: 'creative-agent',
sogni_tool_execution: true,
}),
});
if (!response.ok) throw new Error(await response.text());
const { choices } = await response.json();
console.log(choices[0].message.content); curl https://api.sogni.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $SOGNI_API_KEY" \
-d '{
"messages": [{ "role": "user", "content": "Create an 8-second MiniMax H3 video from this brief: A ceramic artist opens a glowing kiln as the camera slowly pushes in; fire crackles, tools clink softly, and a restrained string score begins" }],
"sogni_tools": "creative-agent",
"sogni_tool_execution": true
}' Creative Agent accepts the short brief above and prepares MiniMax H3's structured production prompt. Calling a worker directly? Follow the MiniMax H3 integration guide and use one of these exact model IDs: minimax-h3-fl2va-fp8_t2v, minimax-h3-fl2va-fp8_i2v, minimax-h3-fl2va-fp8_flf2v, minimax-h3-ref2va-fp8_r2v, minimax-h3-fl2va-fp8_t2v_turbo, minimax-h3-fl2va-fp8_i2v_turbo, minimax-h3-fl2va-fp8_flf2v_turbo. Full reference at docs.sogni.ai.
Real generations from Sogni. Hit Use prompt to open the app with MiniMax H3 selected and the full prompt preloaded.
Use a flat monthly plan for credit-free fair-use generation, or buy Spark packs when pay-as-you-go fits better. Both run on the same creator-owned GPU network.
One flat price in the app. Generate under fair use without a per-image meter.
Image, video, music, and language models in one workspace and one API key.
Prefer pay-as-you-go? Call MiniMax H3 by id and pay with Spark packs.
Runs on a decentralized GPU network where workers share subscription revenue.
For creators searching for an 'uncensored Seedance 2.0' alternative, MiniMax H3 is a strong Sogni match because both cover short-form text, image, reference, dialogue, and native-sound workflows. H3 is a separate MiniMax model, not a Seedance derivative. Sogni supports opt-in mature-content workflows with H3, but it is not restriction-free: Sogni's Terms of Use and MiniMax's license and safety requirements still apply.
H3 is a strong fit for short brand films, dialogue scenes, product reveals, fashion clips, animated posters and title cards, motion design, interface animation, music visuals, and stylized social video where sound matters as much as motion.
MiniMax H3 Turbo applies LightX2V's 4-step distillation LoRA to H3 FL2VA text-to-video, image-to-video, and first-and-last-frame workflows. Sogni measured about 3.5× faster I2V and 3.7× faster T2V in warm reference runs, conservatively summarized as about 3.5× faster. Actual time varies, and the upstream v0.1 preview may give up some fine visual detail and audio polish. There is no Turbo Ref2VA workflow.
Not yet. Sogni currently supports up to 768p-class output because that is the highest-resolution stage MiniMax has released as open weights. MiniMax has committed to a second-phase internal upscaler that takes the workflow to 2K; Sogni plans to integrate it as soon as those weights are available. The upgrade will be included automatically for existing subscribers. Unless a release is specifically marked otherwise, newly supported open-weight models and related updates are automatically included in both existing and new subscriptions.
Current end-to-end averages, including queue wait, are about 13–15 minutes for a full-quality standard H3 video and 3–4 minutes for H3 Turbo. Capacity and demand vary day to day, so these are estimates rather than guarantees. Premium Spark and SOGNI pay-as-you-go jobs run ahead of subscription jobs, and Unlimited Pro jobs run ahead of Unlimited jobs. Within the same priority class, Sogni's dynamic scheduler favors accounts with the least H3 usage that UTC day. This helps prevent queue buildup for casual and regular fair-use subscribers while making throughput available to as many users as possible.
Yes. H3 creates synchronized 32 kHz stereo sound and 24 fps video together.
All seven current H3 modes are available through both the Sogni API and Creative Agent Skill: standard text-to-video, first-frame image-to-video, first-and-last-frame video, and Ref2VA reference-to-video, plus 4-step Turbo text-to-video, image-to-video, and first-and-last-frame workflows. Sogni Web currently exposes standard text-to-video, image-to-video, and Ref2VA.
Sogni generates roughly 5–15 second clips at 24 fps. Sogni Web offers six presets from square through portrait and cinema layouts; direct SDK/API jobs accept dimensions on a 32 px grid, up to 1344 px per axis and 1,032,192 pixels total.
Standard MiniMax H3 is billed at 16 Spark ($0.08) per second of finished video, and H3 Turbo at 6 Spark ($0.03) per second — about 2.67× cheaper — where 1 Spark = $0.005. Within each speed tier, duration is the only price factor: layout and workflow do not change the per-second rate, and references carry no surcharge. On standard H3, a 5.17 second clip is 82.67 Spark ($0.41), 10.13 seconds is 162 Spark ($0.81), and the longest 15.08 second clip is 241.33 Spark ($1.21). On H3 Turbo, the same 5.17 second clip is 31 Spark ($0.155) and the longest 15.08 second clip is 90.5 Spark ($0.4525). Pay with Spark, or use eligible fair-use coverage on an active Unlimited plan.
There is no hard daily video cap, but Sogni intentionally does not publish a fixed count because that would encourage usage aimed at a limit rather than normal fair use. Capacity is designed to exceed the needs of an average user. At the standard $0.08-per-second pay-as-you-go rate, a full-quality 15-second H3 render represents $1.20 in render capacity. Just 17 such renders would exceed the $20 monthly cost of Unlimited, and 42 would exceed the $50 monthly cost of Unlimited Pro—in a single day. Those examples illustrate plan value, not promised daily quotas or usage targets.
Yes. Premium Spark and SOGNI pay-as-you-go jobs receive priority over subscription jobs, Unlimited Pro receives priority over Unlimited, and standard H3 uses more dynamic throttling than H3 Turbo because of heavier demand. There is no hard daily limit. Instead, available concurrency is halved as an account crosses each higher daily usage threshold. The thresholds adjust with available capacity and demand and reset at the start of each UTC day. Fewer than 10% of Unlimited subscribers currently experience a fair-use over-usage slowdown. Active-job concurrency and queue-size limits also apply.
Yes. Sogni is designed for an individual creator to keep a substantial queue moving throughout the day, with high queue limits plus API and Creative Agent access. For automated production, use the Sogni Creative Agent Skill for Claude, Codex, or Hermes. Fair use still applies: unattended continuous 24/7 infrastructure or multi-user production workloads require pay-as-you-go Spark or an Enterprise arrangement. Read the Creative Agent Skill guide.
Yes. The standard Ref2VA workflow in Sogni Web, the API, and Creative Agent accepts 0–9 images, up to 3 videos, and up to 3 audio clips, with at least one visual reference (image or video) and no more than 12 files total. Turbo currently has no Ref2VA variant. Sogni does not expose MiniMax's hosted editing or 2K regeneration stages.
All seven current MiniMax H3 modes are available through the Sogni API and Creative Agent Skill: four standard workflows and three Turbo workflows. Sogni Web currently exposes standard text-to-video, image-to-video, and Ref2VA reference-to-video. Use Spark pay-as-you-go or eligible Unlimited fair-use coverage.
Yes. Sogni has received written authorization directly from MiniMax to offer MiniMax H3 through the Sogni platform. Reports that H3 cannot legally be used in certain regions are incorrect: MiniMax licenses H3 for deployment in the United States, European Union, United Kingdom, and South Korea through its formal authorization process. The published weights remain under the MiniMax H3 Community License Agreement, which governs self-hosted deployments rather than generations you run on Sogni.
For Creative Agent, write a compact shot brief: subject and setting, action over time, camera framing and movement, visual style, then dialogue, ambience, effects, and music. The agent expands it for H3. Direct SDK users should follow MiniMax's structured three-field prompt guide.
MiniMax reports stable dialogue support for Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish, with additional languages supported to varying degrees.
Sogni is built with privacy and creative freedom in mind. Your work remains your own, and inference runs through the Sogni Supernet, a decentralized network of creator GPUs, instead of requiring local hardware or a separate model host. Use is still governed by Sogni's Privacy Policy and Terms of Use.
Create in the app, or build with the API. Your call.