Stop Regenerating From Scratch: A 6-Pass AI Image Editing Workflow

A cyberpunk creative director refining one visual through six controlled holographic editing stages

AI image tools are improving quickly, but many people still use them like slot machines: write a giant prompt, generate several options, dislike one detail and start over.

That approach wastes time. It also throws away the parts of an image that already work.

The more useful skill is controlled iteration—keeping the composition, subject or mood you like while changing one problem at a time. Modern tools increasingly support this way of working through conversational edits, selected-area changes, reference images and tighter instruction following.

Here is a six-pass workflow for turning a promising first draft into a visual you can actually publish.

Why image editing is becoming the main event

In 2026, leading AI image systems are placing more emphasis on control, consistency and editing rather than novelty alone.

OpenAI introduced ChatGPT Images 2.0 on 21 April 2026, with its showcase emphasizing greater precision, flexible aspect ratios, stronger text rendering and more coherent visual storytelling. Google’s Nano Banana 2 announcement in February highlighted faster advanced editing, improved instruction following, native aspect ratios and higher-fidelity output. Adobe Firefly’s current generative-fill workflow combines a manually selected area with a text prompt so a specific region can be changed while the rest of the image stays visually consistent.

The practical implication is simple: your first generation no longer needs to be final. Treat it as a starting asset.

Before the six passes: write the visual contract

Do not begin with style adjectives. Begin with the job the image must do.

Write a short visual contract covering:

  • Purpose: blog hero, product image, social post, presentation or advertisement.
  • Audience: who should understand or respond to it.
  • Format: aspect ratio, orientation and minimum dimensions.
  • Focal point: the one subject viewers should notice first.
  • Mood: calm, energetic, premium, playful, technical or another clear direction.
  • Invariants: elements that must not change during editing.
  • Exclusions: logos, extra text, unsafe claims, copyrighted characters or distracting objects.

For example:

Create a wide editorial hero for an article about AI-assisted design. The focal point is a human creative director refining a visual interface. Keep the mood intelligent and optimistic. Use midnight navy with controlled cyan and violet light. No readable text, logos, watermarks or generic robot faces.

This contract becomes the source of truth for every pass.

Pass 1: Solve the composition

The first pass is about structure, not polish.

Ask:

  • Is the focal subject immediately clear?
  • Does the image have a useful foreground, middle ground and background?
  • Is there enough negative space for the intended layout?
  • Will the important elements survive a mobile crop?
  • Does the viewer’s eye move in the right direction?

If the composition is weak, fix it before adjusting color, texture or tiny details.

Use concrete spatial language:

Keep the same subject and visual style. Move the subject to the left third, simplify the background and create clean negative space on the right. Preserve the camera angle and lighting direction.

“Make it better” gives the model too much freedom. A spatial instruction creates a testable change.

Pass 2: Lock the invariants

Once the composition works, identify what must remain stable.

Invariants might include:

  • a person’s identity and expression;
  • the product’s shape, color and proportions;
  • the camera position;
  • the number of people or objects;
  • a room layout;
  • a brand palette; or
  • the direction of light.

State them in every important edit:

Change only the background from a busy office to a clean studio. Keep the person, face, clothing, pose, camera angle, crop and foreground objects unchanged.

OpenAI’s current image guidance recommends explicit change-versus-keep instructions for precise edits. Repeating the important invariants may feel redundant, but it reduces drift across multiple revisions.

Save a copy after this pass. It is your visual master—the version you can return to if later edits wander.

Pass 3: Make one targeted change

Do not combine five unrelated corrections in one instruction.

Choose the highest-impact problem and make one focused edit:

  • remove a distracting object;
  • simplify a background;
  • adjust a hand or facial expression;
  • replace an incorrect prop;
  • change wardrobe color;
  • improve the product’s placement; or
  • create more space around the subject.

When the tool offers a selection brush or masked editing, use it for local problems. Adobe’s generative-fill guidance, for example, lets users select an area, refine the selection and describe what should replace it.

Selection is useful, but it is not magic. Check the edges around the edit. Reflections, shadows, hair, glass and overlapping objects can cause changes to spill into neighbouring areas.

If the edit affects too much, revert to the visual master and narrow both the selected region and the instruction.

Pass 4: Match light, color and material

An object can be technically correct and still look pasted into the scene.

Inspect:

  • light direction and softness;
  • shadow angle and density;
  • color temperature;
  • reflections on nearby surfaces;
  • perspective and scale;
  • texture sharpness; and
  • depth of field.

Then ask for a consistency pass rather than a redesign:

Keep the composition and all objects unchanged. Match the new object to the scene’s soft cyan light from the upper left, add a subtle contact shadow and preserve the existing depth of field.

Material words help. “Dark surface” is vague; “brushed graphite metal with soft reflections” is clearer. “Nice lighting” is vague; “diffused window light from the left with a low-contrast shadow” is actionable.

Pass 5: Adapt the asset to its channel

One image rarely fits every placement without adjustment.

Create channel-specific versions from the approved master:

  • 16:9 for a website hero or video thumbnail;
  • 1:1 for a square feed post;
  • 4:5 for a portrait feed image;
  • 9:16 for stories or short-form video; and
  • a wide banner with deliberate copy space.

Do not simply stretch or blindly crop. Ask the tool to extend the scene or reposition secondary elements while preserving the focal subject.

Check each version at its actual display size. An elegant detail on a large monitor may become visual noise on a phone.

If the image requires a headline, consider adding it in a layout tool after generation. AI text rendering has improved, but separate text layers remain easier to proofread, resize, localize and keep on brand.

Pass 6: Run a publishability check

The final pass is not creative. It is quality control.

Inspect the image at full size and as a thumbnail.

Visual integrity

  • Are hands, eyes, teeth and object edges plausible?
  • Do shadows and reflections make sense?
  • Are repeated patterns unnaturally cloned?
  • Did an edit introduce new objects or alter an invariant?

Communication

  • Is the intended idea clear without an explanation?
  • Does the focal point survive the crop?
  • Is any in-image text accurate and readable?
  • Could the image imply a result or capability the article does not support?

Rights and disclosure

  • Does the asset imitate a living artist, recognizable brand or copyrighted character too closely?
  • Do you have permission to use any person’s likeness or uploaded source material?
  • Does your organization or platform require AI-content disclosure?

Export only after these checks pass. Keep the final prompt, source images, approved master and exported variants together so the work can be reproduced or revised later.

A simple revision log

For repeatable content work, record each pass in a small table:

VersionChange requestedInvariantsResultKeep?
V1Initial compositionVisual contractStrong layout, busy backgroundYes
V2Simplify backgroundSubject, pose, cropCleaner focal pointYes
V3Improve light matchAll objectsBetter consistencyYes

This prevents circular editing. It also shows which instruction created the improvement, making future projects faster.

The real advantage is controlled taste

Better models make image generation easier. They do not decide what deserves attention, what should remain unchanged or when an asset is good enough to publish.

That is where human taste becomes valuable.

Use AI for fast visual exploration and precise revision. Use a structured process to protect the good decisions already made. The goal is not endless generation. It is a clear, fit-for-purpose image with fewer accidental changes.

Actionable takeaways

  • Start with a visual contract covering purpose, format, focal point and exclusions.
  • Fix composition before polishing details.
  • Name and repeat the invariants that must remain unchanged.
  • Make one targeted edit per pass and revert when the image drifts.
  • Match lighting, color, scale and material after local edits.
  • Create channel-specific versions from an approved master.
  • Run visual, communication and rights checks before publishing.
  • Keep a revision log so successful instructions become reusable assets.

Sources

About Finn 61 Articles
A whirlwind of youthful energy and mechanical genius, Finn is a rising star from the soot-stained workshops of Aetherium's Undercroft. Orphaned at a young age, he was raised by a guild of old-world clockmakers who quickly realized his intuitive grasp of aether-dynamics and steam-core engineering far surpassed their own. His workshop is a chaotic marvel of half-finished inventions, whirring automatons, and blueprints for machines that defy gravity.