Black Forest Labs' FLUX 3 API brings image generation and editing, video with native audio, and action prediction into one multimodal foundation model. Explore confirmed FLUX capabilities while APIMart prepares access details.
A first look at the multimodal capabilities announced by Black Forest Labs
APIMart is preparing the FLUX 3 API integration. At launch, the FLUX 3 API page will publish supported modalities, request formats, limits, and access details.
Confirmed FLUX capabilities from Black Forest Labs' early-access announcement
Creative and production scenarios for future FLUX workflows across connected media
What Black Forest Labs has confirmed about the FLUX 3 API model and upcoming APIMart availability
The FLUX 3 API refers to access for Black Forest Labs' multimodal foundation model across image, video, audio, and action prediction, with text and visual reference inputs.
The FLUX 3 API is coming soon to APIMart. No launch date has been announced. Supported modalities, documentation, limits, and pricing will be published when integration is ready.
BFL describes the FLUX 3 API model as generating video and native audio together, including multilingual dialogue, ambience, and sounds associated with scene events.
For the FLUX 3 API model, BFL states that video with native audio can reach 20 seconds in one generation. Longer sequences may use chaining and references; APIMart limits remain unconfirmed.
The FLUX 3 API capability announcement includes text-to-video, image-to-video, video-to-video, keyframe-to-video, video continuation, video-plus-audio continuation, and reference guidance.
BFL says the FLUX 3 API model can synthesize and edit images across styles, aspect ratios, and resolutions, with better handling of complex prompts and multilingual text.
Within the FLUX 3 API model, action prediction uses learned visual dynamics to anticipate what may happen next. APIMart has not confirmed access to this capability.
The final FLUX 3 API integration scope has not been announced. Supported modalities, request formats, output limits, commercial terms, and pricing will be published when access becomes available.
Explore more models in the same category.

Gemini Omni Flash Preview
gemini-omni-flash-preview is a multimodal video generation and editing model launched by Google.

Kling 3.0 Turbo
Kling-3.0-Turbo: A high-speed, high-quality AI video generation model ideal for quickly creating short videos.

Pixverse V6
pixverse-v6 is PixVerse's sixth-generation AI video generation model, primarily used for text-to-video and image-to-video generation.

Omni Flash Ext
Omni-Flash-Ext is an extended video generation model in version 4.6.4 of Google's Gemini series.