# AI Image Generation Glossary: 31 Terms Explained

Canonical page: https://fuser.studio/articles/ai-image-generation-glossary

Plain-English definitions of 31 AI image and video generation terms, from LoRA, ControlNet and CFG to inpainting, seeds, upscaling and image-to-video.

[All guides](/articles) · [Creative workflows in Fuser](/features/creative-workflows)

**Quick answer:** most AI image vocabulary describes one of three things: what you tell the model (prompt, negative prompt, reference image), how strongly it listens (guidance, strength, LoRA weight, control images) and what you do with the result (upscaling, inpainting, image-to-video). The definitions below are alphabetical, each starting with a one-sentence answer.

### Aspect ratio

The proportion of an image's width to its height, such as 16:9, 1:1 or 9:16. Most generation models take a ratio or a preset size rather than exact pixels; for exact platform dimensions, compose the result on a sized canvas, as in our [social media image sizes guide](/articles/social-media-image-sizes).

### Background removal

Separating the subject from its background and returning it on transparency. In a node workflow it usually sits between generation and layout; Fuser's [Background Remover](/models/background-remover) feeds a cut-out straight into the Compositor.

### Canny edge map

A black image with white outlines, produced by the Canny edge-detection algorithm. It is used as a control image to lock the composition and silhouettes of a new generation; see [Canny Edge Detection](/models/canny-edge-detection).

### CFG scale (guidance scale)

A number that sets how strictly a diffusion model follows your prompt. Higher values stick closer to the text but can look harsh or oversaturated; lower values are looser and more natural. Many image and video nodes expose it as Guidance or CFG, including [Kling 3.0](/models/kling-3-0-video).

### ControlNet

A technique that conditions image generation on a control image such as edges, depth or pose, so the output follows a given structure. [SDXL Canny ControlNet](/models/sdxl-canny-controlnet) takes an edge map plus a prompt.

### Denoising strength

How much an image-to-image or inpainting model is allowed to change the input, usually from 0 to 1. Low values make subtle edits; values near 1 mostly ignore the original.

### Depth map

A grayscale image where brightness encodes distance from the camera, lighter meaning closer. It guides layout and perspective in a new generation; [Depth Anything V2](/models/depth-anything-v2) creates one from any photo.

![A source image beside its depth map and Canny edge map, the two most common control images for guiding AI image generation.](https://statics.fuser.studio/cms/7377c50e-7583-4214-8d27-a5b5267ac000)

_A source image and the two control images most often made from it._

### First and last frame

Image-to-video inputs that fix the opening image and, optionally, the closing image of a clip, with the model generating the motion between them. Kling 3.0, [Veo](/models/veo) and [FLUX.3](/models/flux-3) accept an end frame in Fuser.

### Image captioning

Generating a text description of an image. Captions become prompts, alt text or search metadata; [Florence-2](/models/florence-2-image-captioner) is one captioning node.

### Image-to-3D

Turning one or more images into a 3D mesh. In Fuser, [Hunyuan3D](/models/hunyuan3d-v3) and similar nodes output meshes that can continue into texturing, rigging or a rendered turntable.

### Image-to-image (img2img)

Generating a new image from an existing one plus instructions, keeping some of its structure. Reference-led editing models such as [FLUX.1 Kontext](/models/flux-1-kontext) and [GPT Image](/models/gpt-image) are the modern form of this.

### Image-to-video

Animating a still image into a video clip, with a prompt describing the motion. See [the best image-to-video models](/articles/best-image-to-video-models).

### Inference steps

The number of denoising passes a diffusion model runs. More steps can add detail at the cost of time; many modern models are tuned to need few.

### Inpainting

Regenerating only a masked region of an image while keeping the rest. [SDXL Inpainting](/models/sdxl-inpainting) takes an image, a mask and a prompt; GPT Image also accepts a mask for targeted edits.

### Lip sync

Matching a character's mouth movement in a video to an audio track. [LatentSync](/models/latentsync) is a video-to-video lip-sync node.

### LoRA

A LoRA (low-rank adaptation) is a small add-on file that teaches a base model a specific style, product or character without retraining the whole model. You apply it with a weight; in Fuser, [FLUX.1 Dev](/models/flux-1-dev) and several other nodes take LoRA inputs.

### Mask

A black-and-white image marking which pixels an edit may change. Masks drive inpainting and targeted edits.

### Motion transfer

Applying the movement from a reference video to a different character or image. [Kling 3.0 Motion](/models/kling-3-0-motion) animates a character image with the motion of a reference clip.

### Negative prompt

Text describing what you do not want in the output, such as blur, extra fingers or watermarks. Supported by many diffusion and video models, including Kling 3.0.

### Node

One step in a visual workflow: a model, an input, a utility or an output, connected to other steps by wires. A node canvas makes the whole process visible and repeatable; see [the best node-based AI tools](/articles/best-node-based-ai-tools).

### Outpainting

Extending an image beyond its original borders with newly generated content. Fuser has no dedicated outpaint node; the practical route is to place the image on a larger Compositor canvas and regenerate the frame with a reference-led model.

### Prompt

The text instruction that describes what to generate. Good prompts name the subject, setting, light, composition and style, and change one thing at a time between iterations.

### Prompt template

A reusable prompt with variables filled in by other inputs. Fuser's [Prompt Template](https://docs.fuser.studio/docs/nodes/primitive/prompt-template) node turns {variable} slots into inputs and supports inline choices like {red|blue}.

### Recipe

In Fuser, a Recipe is a saved workflow packaged as a single reusable node, with the inputs you chose exposed. See [Recipes](/features/recipes).

### Reference image

An image given to a model as guidance for subject, style or composition rather than as something to edit directly. Using one shared reference is the simplest way to keep a series consistent; see [consistent characters in AI images and video](/articles/consistent-characters-ai).

### Seed

The random number a generation starts from. The same seed, prompt and settings reproduce a very similar image, which makes it useful for comparing one change at a time.

### Style reference

A reference image or description used to transfer a look, such as palette, lighting or medium, without copying the subject. In Fuser you can keep style wording in a [Style](https://docs.fuser.studio/docs/nodes/primitive/style) node and connect it to every generation.

### Text-to-image

Generating an image from a written prompt alone. See [the best AI image models](/articles/best-ai-image-models).

### Text-to-video

Generating a video clip from a written prompt alone, often with optional audio. Most video models also accept a start image for more control.

### Upscaling

Increasing an image or video's resolution while adding plausible detail. Fuser includes image upscalers such as [ESRGAN](/models/esrgan-upscaler) and [Topaz](/models/topaz-image-upscale), and video upscalers such as [SeedVR](/models/seedvr-upscale).

### Workflow

A chain of steps that turns inputs into a finished asset, for example prompt, image, upscale, video and sound. See [how to chain AI models into one workflow](/articles/chain-ai-models-workflow).

## The ten terms you will meet first.

Quick definitions and where each one appears in a Fuser workflow.

### Essentials

| Term | Plain meaning | In Fuser |
| --- | --- | --- |
| Prompt | What you ask for, in words. | Text or Prompt Template node into any model. |
| Reference image | A picture that guides subject or style. | Connected to reference-led models such as GPT Image or Gemini Image. |
| Seed | The random start point of a generation. | A setting on many image nodes. |
| CFG / guidance | How strictly the prompt is followed. | Guidance or CFG input on image and video nodes. |
| LoRA | A small add-on that teaches a style or subject. | LoRA inputs on FLUX.1 Dev and other nodes. |
| ControlNet | Structure from a control image. | Canny or depth maps into a ControlNet node. |
| Inpainting | Regenerate only a masked area. | SDXL Inpainting; masks on GPT Image. |
| Upscaling | More pixels, more detail. | ESRGAN, Topaz, Recraft and SeedVR nodes. |
| Image-to-video | Animate a still. | Kling, Veo, Seedance and other video nodes. |
| First / last frame | Fix where a clip starts and ends. | End-frame inputs on Kling 3.0, Veo and FLUX.3. |

## Questions, answered.

### What is a LoRA in AI image generation?

A LoRA (low-rank adaptation) is a small add-on file that teaches a base image model a specific style, product or character without retraining it. You load it alongside the model and set a weight for how strongly it applies.

### What does CFG scale do?

CFG, or guidance scale, controls how closely a diffusion model follows the prompt. Higher values follow the text more literally but can look harsh; lower values are looser and more natural.

### What is the difference between inpainting and outpainting?

Inpainting regenerates a masked area inside an image. Outpainting extends the image beyond its borders with new content.

### What is a seed in AI art?

A seed is the random number a generation starts from. Reusing the same seed with the same prompt and settings reproduces a very similar result, which makes controlled comparisons possible.

### What is ControlNet used for?

ControlNet guides a generation with a control image such as a Canny edge map, depth map or pose, so the output keeps a given composition or structure while the prompt changes the look.

## Put the vocabulary to work.

Connect prompts, references, control images and models on one canvas and see each setting change the result.

[Explore creative workflows](https://fuser.studio/features/creative-workflows) · [Explore all guides](https://fuser.studio/articles)

## More articles

- [How to Keep AI Characters Consistent Across Images and Video](https://fuser.studio/articles/consistent-characters-ai.md)
- [AI Storyboard to Animatic: A Node Workflow](https://fuser.studio/articles/ai-storyboard-to-animatic.md)
- [Seedance 2.5 vs Kling 3.0 vs Veo 3.1: Same-Prompt Test](https://fuser.studio/articles/seedance-vs-kling-vs-veo.md)
