Fuser Apps are here 🚀
Fuser Apps are here. Free generations for the next month 💫
Let's GobyBlack Forest Labs
Surgically edit, restyle, and combine visual scenes through plain text instructions without masking or tedious rework
From single-image tweaks to multi-reference composites, edit scenes naturally using conversational directions instead of complex node trees.
Built on a 12-billion parameter multimodal flow transformer, Kontext blends requested changes into native lighting, geometry, and spatial context.
Explore how surgical object swaps, era transfers, and precision typography retain consistent character and environmental cohesion.
See how art directors, product designers, and campaign visualizers use instruction-guided editing to streamline creative production.
The Max variant provides enhanced spatial layout reasoning, tighter prompt adherence, and higher precision for typography and multi-image blends at 110.3 credits per generation. The standard Pro variant processes at 55.2 credits per generation, offering faster iteration speeds for everyday edits.
Use the multi-image endpoint when combining elements from up to four distinct references, such as applying a specific outfit from one image to a subject in another. For localized modifications to an existing frame, the single-image endpoint provides faster, targeted results.
FLUX.1 Kontext excels at surgical local object edits, style transfers, context-aware typography replacement, character consistency across scenes, and historical photo restoration without requiring manual inpainting masks.
Avoid stacking multiple distinct changes into a single complex instruction, as sequential edits yield cleaner results. Additionally, avoid feeding more than three reference images at once, using vague pronouns like 'it' or 'her', or writing non-English prompts.
FLUX.1 Kontext was created by Black Forest Labs and runs on a 12-billion parameter multimodal flow transformer architecture that processes text instructions and image latent representations concurrently.
Surgically edit, restyle, and combine visual scenes through plain text instructions without masking or tedious rework