There is a small experiment I have been running with Gemini that has left me with a rather interesting question.

I am not entirely sure what is happening.

And perhaps that is the most interesting part.

The experiment started with a fairly ordinary question: How deep can an underground excavation realistically go?

Rather than immediately reaching for engineering drawings, excavation manuals, geological reports and a frightening amount of technical documentation, I decided to do something simpler first.

I asked Gemini to show me.

The idea was to generate images of a large underground excavation site and gradually modify the prompt until the image looked something like what I imagined a serious excavation project should look like.

Nothing particularly revolutionary there.

The unusual part was what happened next.

I kept working from essentially the same prompt, making relatively small edits and adjustments. I did not maintain a long conversation in which I repeatedly told Gemini what was wrong with the previous image. I did not give it a score. I did not say, “That last image was 30% better than the one before it.”

There was no explicit feedback loop.

And yet, the images seemed to get better.

That raised a much more interesting question:

Can you improve AI-generated results simply by repeatedly refining the prompt, even when the model isn’t given an explicit mechanism for measuring its own improvement?

I don’t have a definitive answer.

But the experiment is fascinating.

From a crude hole in the ground to something that looks like an excavation

The first image was, frankly, rather crude.

It conveyed the basic idea: there was a large hole, there was earth, there was an excavation site, and something was happening underground.

Technically speaking, mission accomplished.

Artistically and structurally?

Not quite.

It looked more like an AI’s first approximation of “very large hole” than a serious construction site.

So I started tweaking.

The prompt gained additional descriptions.

Then more.

I began specifying things that should exist in a real excavation: structural support, machinery, access routes, retaining systems, workers, scale indicators, layers of earth, underground structures and the general visual language of a major engineering project.

The process was not:

Generate → evaluate → tell the AI exactly what it did wrong → generate again.

It was closer to:

Generate → think → modify prompt → generate → think → modify prompt → generate.

The conversation itself was not carrying the experiment forward as a traditional feedback loop.

The prompt was.

And that distinction is what caught my attention.

Eventually, the images began looking considerably more convincing.

The excavation started to have depth.

The machinery started to make more sense.

The supporting structures became more plausible.

The scene began communicating scale.

The final result looked much closer to the sort of enormous excavation site I had originally been trying to visualize.

And that is where things became weird.

I wasn’t teaching the model

At least, not in the conventional sense.

I wasn’t building a dataset.

I wasn’t fine-tuning a model.

I wasn’t giving Gemini a collection of successful and unsuccessful examples and asking it to learn from them.

I wasn’t even necessarily continuing the same conversational thread.

I was changing the instructions.

That sounds like a subtle distinction, but technically it matters.

A generative AI system receives an input and produces an output based on the information and statistical patterns available to it. When generating an image, the prompt provides semantic constraints describing what the model should attempt to create.

If the prompt says:

“A very deep underground excavation site.”

the model has an enormous amount of freedom.

How deep?

What type of excavation?

What kind of soil?

What support system?

What equipment?

What scale?

What construction stage?

What viewpoint?

What does “realistic” mean in this context?

There are thousands of possible interpretations.

The more useful constraints we introduce, the smaller that space of possible interpretations becomes.

So perhaps the improvement is not mysterious at all.

Perhaps I was simply reducing ambiguity.

But that doesn’t completely explain what I was seeing.

The prompt became a kind of steering mechanism

One way to think about prompting is that you are not giving the model a recipe.

You are steering it through an enormous space of possibilities.

A vague prompt gives the model considerable freedom.

A more precise prompt constrains the generation.

For example:

Version 1

A very deep underground excavation site.

That’s enough to establish a concept, but not enough to establish engineering reality.

Now imagine adding:

A large-scale deep excavation for an underground facility, with vertical retaining walls, staged excavation, heavy construction machinery, temporary structural bracing, workers for scale, access ramps, exposed geological layers and realistic construction logistics.

Suddenly, the model has considerably more information about the scene.

We can keep going:

A large-scale deep excavation approximately 17 metres deep, showing multiple excavation stages, reinforced retaining walls, internal bracing, construction equipment operating at different elevations, access ramps, workers for scale, exposed soil and rock layers, realistic drainage and construction logistics…

Now we are no longer merely asking for “a hole.”

We are describing a construction environment.

The prompt has become a compressed specification.

And each small edit can potentially remove another ambiguity.

That may explain much of the apparent improvement.

But there is still an intriguing detail.

The strange part: there was no explicit score

Normally, when we talk about machine learning improving through feedback, we imagine some kind of evaluation mechanism.

The system generates something.

It receives feedback.

It determines whether the result was good or bad.

It adjusts.

Repeat.

That is a genuine feedback loop.

My experiment wasn’t doing that.

There was no explicit numerical score.

There was no instruction saying:

“Compare this image with the previous one and improve the engineering realism by 20%.”

There was no formal reward function.

There wasn’t even a carefully constructed evaluation rubric.

I was simply looking at the result and thinking:

Okay. That doesn’t look quite right. Let’s change the prompt.

Then the next generation often looked better.

So what exactly improved?

The model?

The prompt?

My understanding of the problem?

Or all three in some indirect way?

That is where the experiment becomes much more interesting.

Maybe the human is the feedback loop

This is probably the simplest explanation.

The model doesn’t necessarily need to evaluate itself if I am doing the evaluation.

Every generated image gives me information.

I look at it.

I notice that something is missing.

I add that missing concept to the prompt.

The next image therefore contains more constraints.

In that sense, the feedback loop exists — it just isn’t entirely inside the AI system.

It looks something like this:

Prompt → Image → Human observation → Prompt modification → Image → Human observation → Prompt modification

The human is effectively acting as the evaluator.

But there is an interesting twist.

The later prompts don’t necessarily need to explicitly mention the previous image.

The accumulated knowledge can exist entirely in the evolving prompt.

That means the prompt itself becomes something like a portable state representation.

The conversation can disappear.

The thread can be restarted.

The model doesn’t necessarily need to remember what happened previously.

If the prompt has captured the useful refinements, the experiment can continue.

That is potentially quite powerful.

Prompt engineering as iterative search

There is another way of describing what happened.

It resembles an informal optimization process.

Imagine that the space of possible images is enormous.

We start with a very broad instruction:

“Generate a deep excavation.”

The result lands somewhere in that enormous space.

We inspect it and decide that we want something slightly different.

So we modify the input.

The next generation lands somewhere else.

We inspect again.

Another modification.

Another generation.

In mathematical language, we might loosely describe this as iterative search through a generative space.

We aren’t necessarily optimizing a clearly defined numerical objective.

Instead, we’re performing qualitative optimization.

The objective is something like:

“Make this look more like a plausible large-scale excavation.”

That objective is evaluated by a human.

The prompt changes are our search steps.

And the generated images are our candidate solutions.

This is not the same thing as model training.

It is closer to interactive optimization through prompt-space exploration.

And that distinction is important.

Why tiny changes can have surprisingly large effects

One of the more interesting characteristics of generative models is that seemingly small changes to an instruction can alter the output substantially.

This happens because natural-language prompts aren’t simply interpreted as rigid commands.

They contain concepts, relationships, emphasis and context.

Adding a single phrase can change the semantic interpretation of an entire scene.

Compare:

“deep excavation”

with:

“deep excavation with temporary retaining structures and internal bracing”

The second prompt doesn’t merely add two objects.

It changes the implied engineering context.

Likewise:

“underground construction”

is dramatically different from:

“a large-scale underground construction project viewed from above, with workers, excavators, access ramps, retaining walls and multiple levels of excavation visible.”

The model now has much stronger clues about composition, scale and purpose.

So the improvement may come from a phenomenon that is actually quite mundane:

better instructions produce a narrower and more appropriate generation space.

But then again…

Why did the progression feel so consistent?

That’s the part I haven’t solved.

Did Gemini actually improve?

This is where I want to be careful.

I cannot claim that Gemini itself was learning from the previous attempts.

I did not conduct a controlled experiment designed to establish that.

There are several competing explanations.

Perhaps the prompt became progressively better.

Perhaps my own understanding of what I wanted became progressively better.

Perhaps the image generation system is particularly sensitive to certain descriptive concepts.

Perhaps the generation process contains additional mechanisms that influence results.

Perhaps randomness played a role.

Or perhaps several of these effects were occurring simultaneously.

Without a controlled experiment, it would be difficult to distinguish them.

For example, to properly test the hypothesis, we could generate hundreds of images using carefully controlled prompts, random seeds where available, identical settings and independent sessions.

We could then have blind evaluators score the results for realism, engineering plausibility, composition and scale.

That would begin to tell us whether the apparent improvement is statistically meaningful.

I wasn’t doing that.

I was just trying to figure out what a giant underground hole might actually look like.

And that is precisely why the result surprised me.

There is another possibility: the prompt is becoming smarter

Not literally smarter, of course.

But more informative.

A prompt can evolve from a simple description into a miniature specification language.

At the beginning, it describes the object.

Then it describes the environment.

Then the relationships between objects.

Then the physical constraints.

Then the visual perspective.

Then the desired level of realism.

Eventually, it starts resembling a technical brief.

This progression is worth thinking about.

We often hear advice about writing the “perfect prompt.”

But perhaps the perfect prompt isn’t something you write once.

Perhaps it is something you discover iteratively.

You start with a rough idea.

Generate.

Observe.

Modify.

Generate.

Observe.

Modify.

Eventually, the prompt itself becomes a record of everything you have learned about how to describe the desired result.

That could be considerably more useful than endlessly trying to invent the perfect prompt from scratch.

And then there is the quota question

Naturally, while doing this, another question eventually appears:

How much is this costing me in quota?

Image generation isn’t necessarily an unlimited resource.

Different AI services can impose different limits on image generation, requests, usage tiers or daily quotas.

And repeated experimentation can consume those limits surprisingly quickly.

But here I have to put an asterisk on the experiment.

I wasn’t specifically measuring quota depletion.

I was investigating the quality of the images.

I didn’t establish a controlled before-and-after measurement of quota usage.

So I cannot definitively say how much quota this particular refinement process consumed, nor can I claim that changing the prompt in this manner has some particular quota efficiency.

That would require a separate experiment.

And perhaps that’s another experiment worth doing.

Because if iterative prompt refinement can produce substantially better results without requiring dozens of radically different prompts, there could be a practical efficiency advantage.

Instead of throwing 50 unrelated prompts at an image generator and hoping one produces something good, you might be able to start with one reasonably good prompt and walk it toward the desired result.

That is a very different workflow.

Less prompt sprawl, more prompt evolution

This is probably my favourite takeaway.

There is a tendency to approach generative AI by accumulating prompts.

You try one.

Then another.

Then another.

Then you search the internet for somebody else’s “ultimate prompt.”

Then you add more adjectives.

Then more instructions.

Eventually you have a monster prompt that looks like it was written by a committee.

The experiment suggests another possibility.

Don’t necessarily create more prompts. Evolve one.

Start with a simple prompt.

Identify what is missing.

Add it.

Generate again.

Identify another weakness.

Refine it.

Generate again.

The prompt becomes progressively more specific without necessarily becoming radically different.

That is a much more organic approach to prompt engineering.

And it has an interesting psychological benefit too.

You are not trying to predict exactly what the model needs before you begin.

You’re allowing the output to tell you what your prompt failed to communicate.

The image becomes part of the prompting process.

The unresolved mystery

And yet, after all this explanation, I still can’t completely shake the original question.

Why did the images appear to improve so consistently?

If the model wasn’t being given an explicit performance score…

If I wasn’t telling it, “this image is better than the last one”…

If I wasn’t necessarily continuing the same conversational context…

Then what exactly are we observing?

Maybe the answer is beautifully simple.

The model wasn’t learning.

I was.

Every image showed me something about how the system interpreted my instructions.

Every imperfect generation gave me another clue.

Every prompt revision encoded that clue.

And the next generation benefited from the accumulated information.

In that interpretation, there was never any mysterious self-improvement happening inside the model.

The intelligence was distributed across the entire loop:

human → prompt → model → image → human → improved prompt → model → improved image.

But there is still a lingering possibility that makes the experiment worth repeating.

Perhaps there are properties of modern generative systems that make this kind of iterative refinement work better than our simple explanation suggests.

Perhaps the relationship between prompt wording and generated imagery is more structured than it appears.

Perhaps repeatedly steering the same conceptual target allows us to navigate the model’s generative space surprisingly efficiently.

Or perhaps I simply got better at asking for what I wanted.

I don’t know yet.

And that’s what makes this experiment fun.

I started by asking an AI to draw a hole in the ground.

I ended up wondering whether we have stumbled upon a surprisingly effective way of searching through a generative model’s possibilities:

Keep the idea. Keep the prompt mostly intact. Change just enough to move the result. Repeat.

No giant prompt library.

No elaborate self-evaluation system.

No enormous feedback architecture.

Just a human, a model, a picture that isn’t quite right, and one more tiny edit.

Then another.

And another.

Until suddenly the hole starts looking like something engineers might actually have to build.

And I still don’t completely know why.

We will be happy to hear your thoughts

Leave a reply

Som2ny Network
Logo
Register New Account
Compare items
  • Total (0)
Compare
0
Shopping cart