“AI is the new compiler” is one of those phrases that sounds good at first.
I understand why the analogy is appealing. You put something in, processing happens inside the system, and something comes out. At this highly abstract level, it fits. A prompt then feels a little like source code. The model feels like a new execution layer. And the result looks like something produced through a technical translation.
But the longer I think about it, the more that very abstraction bothers me. It describes the surface. It does not describe the behavior.
And when working with AI, behavior is what matters.
A compiler does not decide
A compiler is a deterministic system.
The same input produces the same output. When I compile the same code with the same environment, version and configuration, I expect the same result. When something unexpected happens, I look for the error in the input, the compiler, the environment or my assumptions.
That is a very valuable property. It makes compilers trustworthy. It makes build systems verifiable. It makes many technical processes possible to automate.
The compiler does not decide. It follows rules.
Of course, real technical systems always involve details, versions, undefined behavior, platform differences and side effects. But the expectation model remains clear: I treat the compiler as a system governed by rules. I want reproducibility. I want deterministic behavior. I want to be able to locate errors.
An LLM works differently
In some respects, an LLM feels similar. I put text in, the system processes something, and text, code, structure or a basis for a decision comes out.
But the mechanism is different.
An LLM does not work like a compiler applying fixed rules to source code. It generates a plausible continuation based on learned patterns, context, probabilities and model parameters. The same prompt can produce similar results, but not necessarily identical ones. Even when the output appears stable, the underlying expectation model remains different.
This is not a weakness you can simply optimize away. It is a different mode of operation.
- With a compiler, I ask: is the translation correct?
- With an LLM, I also ask: is the result plausible, helpful, complete, appropriate to the context and sound in its subject matter?
These are different review questions.
The mistake begins when we give a probabilistic system deterministic expectations.
That is exactly where many misunderstandings arise. We want reproducibility like a compiler's. AI offers statistical similarity instead. That can be extremely useful. But it demands a different kind of guidance, review and trust.
People and teams are different again
Then there are people, teams and organizations.
When I give a team the same assignment twice, the result can also differ. But that is not because the team continues a probability distribution in the mathematical sense. People use experience, heuristics, assumptions, context, taste, interpretation and sometimes gut feeling.
A team can assess the same input differently today than tomorrow because its knowledge has changed. Because a constraint has become more important. Because someone has contributed a new experience. Because the political, technical or professional environment is no longer the same.
Its behavior is not deterministic like a compiler's. Nor is it probabilistic in the narrow sense of a model.
It is heuristic.
That sounds theoretical, but it is quite practical. As soon as I consider systems in terms of their decision behavior, expectations change:

- Deterministic systems should be reproducible.
- Probabilistic systems need guidance and review.
- Heuristic systems need context, experience and judgment.
Right now, this distinction helps me more than the question of whether AI is “like a compiler.”
The real question is the expectation model
Perhaps we should categorize systems less by how technical they look and more by the behavior they produce.
A compiler is a good tool because it follows rules reliably. An LLM is a good tool because it can quickly produce useful suggestions, structures, variants and first working drafts in ambiguous spaces. A person is valuable because they can bring meaning, context, direction and responsibility together.
These abilities are not interchangeable. They complement each other.
Problems begin when I apply the wrong expectation to the wrong system.
If I expect compiler behavior from an LLM, deviations will annoy me. If I expect human judgment from a compiler, I have chosen the wrong tool. If I expect purely mechanical reproducibility from a team, I ignore the part of the work that consists of interpretation and experience.
For me, the practical point is therefore not: “AI is the new compiler.”

The practical point is: AI is a new class of problem solver, and it needs its own operating model.
What this means for AI agents
This distinction becomes more important as AI is integrated into real work processes.
With a simple prompt, I can still say relatively easily: the result is good or bad. With agents, things become more interesting. An agent plans, changes files, runs commands, interprets errors, repairs, writes tests, rearranges work steps and creates new intermediate states.
That quickly looks like automation. And in some places, it is.
But the questions remain:
- Which part of the process can be checked deterministically?
- Which part is probabilistic work?
- And at what point is human judgment needed?
Tests, lints, type checks, schemas and formatting rules fit deterministic expectations well. They say yes or no. They make loops possible to automate.

Problem formulation, product decisions, tone, architectural taste, prioritization and a sense of the next step work differently. There, an agent simply “keeping going” is not enough. Someone has to judge whether the direction is still right.
This point connects directly to the question of loops. I described it in more detail in A Loop Without Judgment Is Just Faster Repetition: the hard question is not whether a system iterates. The hard question is who assesses the new state.
When we treat AI agents like compilers, we overlook precisely this difference. We then expect a kind of magical execution layer that automatically turns a goal description into correct work.
That is not how good AI work currently feels to me.
It feels more like collaborating with a probabilistic problem solver that can generate a great deal of movement very quickly. This movement is valuable. But it needs guidance, checkpoints and human judgment.
The analogy is not wrong. It is just too small.
“AI is the new compiler” can be a useful starting point. It makes visible how language suddenly feels more executable. It explains why prompts, specifications and work instructions are becoming more important. It shows that a new translation layer has emerged.
For my daily work, the image is still insufficient.
- A compiler translates according to rules.
- An LLM generates plausible continuations.
- People and teams work with heuristics, experience and responsibility.
Taking these differences seriously does not make AI smaller. It makes it easier to work with.
I stop expecting it to behave like a compiler. Instead, I build processes in which probabilistic suggestions, deterministic checks and human judgment interact meaningfully.

Perhaps this is the more useful analogy:
AI is not the new compiler. AI is a new problem solver within a work system that we still need to design properly.
