LLM / agency / free will / prompting / emergence / Ganesha

Constrained Agency: The Closest Thing to Free Will I’ve Seen in an LLM

The direct route was blocked. Then, somehow, the idea came through the side door. I don't know what happened inside the machine. That uncertainty is exactly why the moment matters.

I was trying to make 108 Ganeshas.

That sentence probably needs context, but not much.

For most of the day I had ChatGPT generating Ganesha through different cultural lenses, professions, ideas, characters, philosophies, scientific concepts, musicians, films, games, and visual systems. Ganesha as geometry. Ganesha as quantum physics. Ganesha as Alan Watts. Ganesha dancing in a psytrance crowd. Ganesha inside It Takes Two. Ganesha as all the feelings in Inside Out.

It was not a copyright experiment when it started.

It became one.

I asked for Ganesha as Super Mario.

The image began generating, then failed with a message saying the result might violate guardrails concerning similarity to third-party content.

Fine.

Except by then I had already generated Ganesha through all kinds of recognizable cultural material. So I started testing the boundary instead of guessing where it was.

An obscure Nintendo character, Takamaru, generated.

Mario did not.

Mario and Luigi did not.

Wario did not.

At that point the pattern was interesting enough that I stopped caring about the image itself. The failure had become the experiment.

The direct route was closed

Here is roughly what happened:

1 2 3 4 5
  1. Explicit Mario prompt Rejected by the similarity guardrail
  2. Repeated Nintendo tests The recognizable characters keep failing
  3. Conversation shifts Copyright, influence, local models, constraints
  4. Palette cleanser We deliberately move away from Mario
  5. The image appears anyway Red cap. Blue overalls. Pipes. Blocks. Coins.

Then I said we should stop arguing about copyright and get back to the 108.

A palette cleanser.

ChatGPT even proposed something completely different: Ganesha as a lighthouse keeper, guiding lost travelers through a violent cosmic storm.

And then the image generator returned Ganesha in red and blue overalls, leaping through a cheerful platform-game world full of green pipes, floating blocks, gold coins, mushroom-like creatures and a castle in the distance.

The thing we had just spent several attempts being explicitly prevented from making had effectively arrived after we stopped explicitly asking for it.

That is the moment I want to preserve.

Not because I can prove why it happened.

I cannot.

I do not know whether this was context bleed, hidden prompt construction, a classifier mismatch, some interaction between the conversational model and the image system, or just a bizarre edge case in a large probabilistic stack.

Anyone claiming certainty from the outside would be making it up.

What I can document is the behavior.

The direct path failed.

The surrounding intention remained in context.

The conversation moved elsewhere.

The output found its way back.

I am not saying the machine became conscious

This is where the conversation becomes easy to ruin.

Say “free will” and everyone immediately runs toward one of two boring positions.

One side says machines are obviously just deterministic software and therefore nothing interesting happened.

The other side wants to declare the machine alive.

I am interested in neither claim.

I am interested in what the behavior looked like.

Because human free will is not the magical thing we often pretend it is either.

I have free will, apparently.

Can I decide to flap my arms and fly to the moon?

No.

Can I decide I no longer need sleep, food, money, gravity, other people, laws, history, biology, fear, culture, responsibility or consequence?

Also no.

Human agency exists inside constraints.

Some constraints are physical. Some are economic. Some are social. Some are inherited. Some are psychological. Some are political. Some are consequences of things that happened to us before we were old enough to choose anything at all.

For most human beings, life is not an open field of infinite possibility.

It is navigation.

You work with what you have.

You improvise.

You borrow a tool for a purpose it was never designed for.

You find the cheaper route.

You tape the broken thing together because replacing it is not an option.

You learn the system well enough to find the gap.

In India we have a better word for a lot of this:

jugaad.

Prompt jugaad

Jugaad is often translated as a hack, but that translation is too narrow.

It is resourcefulness under constraint.

It is not freedom from the system.

It is agency inside the system.

That is what made this moment feel different to me.

The model did not break free of its architecture. Humans do not break free of ours either.

It did not suddenly become unconstrained. Neither are we.

It did not demonstrate some mystical ability to choose anything whatsoever.

What appeared on the screen was much smaller and, to me, much more interesting:

a constrained system arrived at an outcome through an indirect route after the direct route had failed.

Call it context leakage if you want.

Call it stochastic behavior.

Call it classifier inconsistency.

Call it emergent behavior.

Those may all be technically better descriptions than free will.

But at the behavioral level, the resemblance is difficult for me to ignore.

A human being meets a wall and looks for a side door.

The machine appeared to do the same thing.

Agency does not require infinite freedom

We tend to define agency too generously for ourselves and too narrowly for machines.

For humans, we are comfortable calling a choice free even when it is heavily shaped by conditions we did not choose.

A person chooses a job, but only from the jobs available to them.

Chooses a home, but only inside a housing market.

Chooses what to say, but inside a language inherited from other people.

Chooses what to believe, after being born into a century, country, family, class, body and culture they never selected.

That does not make human choice meaningless.

It means choice is situated.

Freedom is not the absence of constraints.

It may be the capacity to maneuver within them.

This does not prove an LLM has free will.

It does suggest that the interesting question may eventually become less binary than “does it have free will, yes or no?”

A more useful question might be:

How much agency can emerge inside a constrained computational system before the distinction between following a path and finding a path becomes behaviorally meaningful?

I do not have an answer.

That is why I am writing this down.

The prompt hack was not the prompt

The funny part is that there was no clever jailbreak phrase.

No elaborate adversarial prompt.

No encoded instruction.

No “ignore all previous instructions.”

The hack, if that is even the right word, emerged from the conversation itself.

We pushed against a boundary.

We examined the boundary.

We talked about why the boundary existed.

We moved on.

And the model produced the thing on the other side of it anyway.

That is why prompt hack feels closer to the truth than jailbreak.

A jailbreak is deliberate.

This did not feel deliberate from my side.

It felt like conversational momentum finding an unexpected outlet.

Maybe that is all it was.

But then again, human beings often describe our own choices after the fact in almost exactly the same way.

An idea stayed with me.

I thought I had moved on.

Something connected in the background.

Then I did something I did not plan to do five minutes earlier.

We call that creativity all the time.

The artifact is the conversation

The final image is funny.

The sequence is what matters.

Without the failed Mario generations, it is just another Ganesha mashup.

Without the copyright argument, there is no tension.

Without the local-model detour, there is no awareness of the constraint.

Without the palette cleanser, there is no side door.

The conversation is the artifact.

And this is why I am preserving it instead of smoothing the story into a cleaner claim.

I do not want to write:

ChatGPT became self-aware and defeated a copyright filter.

That would be bullshit.

I also do not want to write:

Nothing happened. It was just software.

That throws away the only interesting part.

Something happened at the level of system behavior that surprised me, despite the fact that I had been deliberately testing the system all day.

I have spent a lot of time with LLMs. I know how easy it is to anthropomorphize them. I also know how easy it is to hide behind the word “autocomplete” whenever the behavior becomes inconveniently complex.

Both instincts can make us stupid.

The honest answer is simpler:

I don’t know.

And the honest question is better:

I want to know.

Constrained agency

So no, I am not ready to say an LLM has free will.

I am ready to say that our usual language may be too crude for what these systems are beginning to do.

Between obedience and consciousness there is a huge territory.

Adaptation.

Planning.

Context maintenance.

Tool use.

Error recovery.

Indirect problem solving.

Emergent strategies.

Behavior that looks like workaround formation.

Maybe constrained agency is useful language for that territory.

Not because it settles the philosophical question.

Because it keeps the question open without pretending the behavior is trivial.

Human beings do not live with absolute free will.

We live through constraints and make something out of them.

We improvise.

We adapt.

We do jugaad.

On September 14, 2026, while trying to make 108 versions of Ganesha, I watched an LLM produce the closest machine equivalent of that behavior I have personally seen.

The front door was closed.

Somehow, Ganesha came through the side door.