The Limits of AI Image Generation

In just the last three years, AI image generation has moved from novelty to normalcy, finding its way into agency workflows, product teams, marketing departments, and nearly every creative job description I review. The conversation has shifted from whether we'll use it to how we'll use it.

What hasn't kept pace is our understanding of what the technology is actually doing.

Most discussions begin with the output.  The images are impressive. Sometimes astonishing. Occasionally indistinguishable from photography, illustration, or 3D renders at first glance.  We debate ethics, copyright, labor, and creative ownership, while the underlying mechanism—the thing responsible for both AI's greatest strengths and its most frustrating limitations—remains largely invisible.

I think that's backwards.

If we're going to build professional creative workflows around AI image generation, we need a more accurate mental model of what we're actually using.

That's why I decided to stress test it.


My career began in editorial illustration, where every assignment started with the same challenge: translate something complicated into something people immediately understand. That discipline eventually expanded into brand strategy, creative direction, and building organizations, but the work itself never really changed.  Whether I'm designing a single illustration or helping launch a company, I'm still solving the same problem:

Translation.

Over the course of my career, I've come to trust refinement more than initial inspiration. Successful creative work reveals itself through iteration.  Good ideas survive that process.  Great ones become more precise because of it.

So I approached AI image generation the same way I would evaluate any new production tool entering a professional workflow.

Not by asking whether it could make beautiful images.

By asking whether it could survive the design process.

 

The Experiment

I needed an object I knew well enough to recognize every mistake.

I settled on the Jaguar E-Type [“E” for experimental].

Not because it makes for compelling imagery, but because it offered something much more useful: an unusually high bar for visual accuracy.  The E-Type is one of the most recognizable industrial designs ever produced.  Its proportions are unforgiving. Its engineering is interconnected.  Move one element out of place and the entire car begins to feel wrong. 

That made it the perfect Control.

I began with a simple goal: take the car through a complete narrative — from neglected barn find to fully restored reinterpretation — requiring a thorough creative workflow.  Along the way I would deliberately push the system through the kinds of iterations professional designers make every day—changing structure, correcting details, refining proportions, preserving what worked while improving what didn't.

What interested me wasn't whether AI could produce a compelling first image.

I wanted to know what happened when refinement became the work.

The results surprised me.

 

Where It Broke

The more specific my requests became, the less reliable the images became.

At first the system was remarkable.  Wrecked body panels, rust, missing components, discarded parts, dusty workshops—AI produced these scenes with extraordinary confidence.  The broken car felt authentic, and surprisingly beautiful.

But something changed as I tried to make the car whole. Each correction solved one problem while introducing another.

Fix the crash bars and the wheels deformed.

Correct the wheels and the engine lost coherence.

Lock those elements in place and the body proportions drifted.

Every iteration became an exchange rather than an improvement. Instead of converging toward a finished design, the work began wandering away from itself. As the prompts became more precise, the results became less predictable.

That wasn't what I expected.

It also wasn't random.

 

Understanding Diffusion

Current diffusion models are probabilistic, not deterministic.  That distinction explains almost every frustration professionals encounter using AI.

Modern image generators don't retrieve pictures from a hidden library, nor do they construct objects the way a designer or engineer might.  They begin with noise—essentially visual static—and progressively remove that noise until a statistically plausible image emerges.

Every prompt nudges that process toward a region of probability rather than toward a fixed destination.

The model isn't reconstructing a Jaguar because it understands what a Jaguar is.  It's converging on the most likely arrangement of pixels given everything it has learned about Jaguars, sports cars, reflections, lighting, metal, photography, perspective, and millions of related visual patterns.

That's an astonishing achievement.

It's also a fundamentally different process from design.

A designer refines toward intention.

A diffusion model refines toward probability.

That difference matters.

Because every new prompt is, in effect, another journey through the probability field.  Even when you ask the model to preserve existing elements, it isn't editing a stable object in the way a human would.  It's generating another statistically plausible solution based on everything it has learned.

That's why iteration feels so strange.

Professionals expect refinement to produce convergence.

Diffusion often produces divergence.

Once I understood that, the behavior I'd been seeing wasn't mysterious anymore.

It was inevitable.

 

Seeing the Seams

Understanding the mechanism changed the way I looked at the images.

At first, AI output feels cohesive.  Then, under repeated refinement, something subtle begins to happen.

You start seeing the seams.

Not obvious failures like extra fingers or impossible anatomy. Those have become clichés. The more interesting failures happen long before obvious breakdown.

A surface that almost behaves like metal, but gets confused with plastic.

A reflection that feels convincing until you follow it across the bodywork.

Mechanical components that look correct individually but never quite relate to one another as a functioning system.

An image can be visually persuasive without ever becoming structurally true. Once you recognize probabilistic image generation for what it is, these inconsistencies become difficult to ignore. More importantly, they explain why increasingly specific prompting often creates increasingly frustrating results:

The system isn't resisting your intent.

It's faithfully executing the process it was designed to perform.

 

What Changed My Mind

None of this diminished my appreciation for AI.

It made me appreciate it more accurately.

Current diffusion models are extraordinary tools.  They're just extraordinary at different things than many professionals assume.

They're exceptional at exploration.

They excel at ideation, mood, composition, visual discovery, and generating possibilities that might never have occurred to a human creator working alone.

They're remarkably good at helping us think.

Where they become less dependable is where professional creative work usually begins: deliberate refinement.

That doesn't make the technology a failure.

It changes where it belongs in the workflow.

 

What AI Image Generation Is Actually Good For

For creative professionals, I came away with a much simpler framework than I expected.

AI image generation is exceptionally valuable for:

  • Rapid concept exploration.

  • Visual prototyping.

  • Mood and atmosphere.

  • Reference gathering.

  • Composition studies.

  • Creative conversation.

  • Accelerating early-stage thinking.

These aren't limitations. They're strengths.

The mistake is expecting the same system that excels at possibility generation to also excel at deterministic production.

Those are different jobs.

Professional creative work still depends on judgment—the ability to recognize when an image communicates exactly what it should, and when statistical plausibility isn't enough.

That judgment doesn't disappear as AI improves.

If anything, it becomes more valuable.

The better AI becomes, the more important it is that we understand what it's actually doing.

Because the question was never whether AI could generate compelling images. The question was always where those images belong inside a professional creative process.

After months of pushing diffusion models through the kinds of refinement they rarely encounter outside professional production, I don't think the future depends on asking AI to become more human.

I think it depends on humans becoming more precise about what we ask of AI.

That's a much smaller shift than the headlines suggest.

It's also a much more useful one.

And once we begin treating AI as another creative tool rather than a magical one, a different question naturally follows:

How do we describe—and transparently document—the role AI plays in the work we create together?

That, I believe, is the next conversation.

 
Next
Next

Cathedral of Mars