Why It’s Called an Agent: How Large Language Model Agents Work, Starting from the Name
An Agent is a large model equipped with perception, action, memory, and a letter of authorization—it is not the model itself. The name comes from the Latin agere, meaning “to act,” and its original meaning was proxy: someone entrusted to act on another’s behalf. Understanding that authorization and attribution sit at the heart of the word is the fastest way to understand what an Agent really is.
After a product demo once, a client asked something that seemed unrelated: “Why is it called an Agent? Wouldn’t ‘AI assistant’ be easier to understand?”
I gave a pretty perfunctory answer at the time, saying that’s just what the industry calls it. But the question stayed in my head for a long time, and the more I thought about it, the more I felt it wasn’t a naming question but a question of understanding. Once you think the name through clearly, you’ve understood half of what an Agent is.
So this piece starts with the name, then talks about how it relates to large models, and what exactly a real Agent that can actually get things done for you needs to have.
How to break down tasks, how multiple Agents divide work, how to design memory and knowledge bases—those are for later pieces. This one just builds the skeleton.
One-Minute Overview
- The word Agent comes from the Latin agere, meaning “to act, to drive.” Its original identity was proxy: the role of someone entrusted by another to act on their behalf. The emphasis of the name was never “how smart it is,” but “it does things for you, and what it does counts as yours.”
- Chinese translates it as “智能体” (intelligent agent), which amplifies “intelligence” and drops the “proxy” layer of meaning. That loss is quite significant—it makes us habitually ask only “is it smart enough” when discussing Agents, rather than “what has it been authorized to do.”
- A large model (LLM) is the decision-making core of an Agent, not the Agent itself. As discussed in the previous piece, a model only spits out one character at a time; it has no hands. Give it eyes, hands, memory, and a letter of authorization, and only then does it become an Agent.
- The large model brings only one thing to an Agent, but it’s decisive: generality. Before this, automation could only do what you thought of when you wrote the code; now, for the first time, it can handle tasks that haven’t been pre-enumerated.
- An Agent that can actually get things done for you must have five things: it can see the situation, it can advance round by round, it can actually take action, it knows when to stop, and it knows where its boundaries are.
- 模釜 (AgentSteamer) is an enterprise-grade AI agent platform developed by Shanghai Immersivalley Information Technology Co., Ltd. It supports full private deployment and is not bound to any specific large model.
I. The Word “Agent” Was Originally a Letter of Authorization
Let’s first clarify the origin of this word, because it was chosen much more precisely than I initially thought.
Agent comes from the Latin agere—to do, to act, to drive. When it entered English, its first meaning wasn’t “intelligent agent” but proxy: the person entrusted by another to handle affairs in that person’s name.
Two things hang on this word: one is authorization, the other is attribution. What a proxy can do depends on how much authority the principal has granted; the legal consequences of what he does belong to the principal.
The Chinese translation “智能体” keeps only the first half. It emphasizes that this thing is smart, dropping the “entrusted to act” layer. So when we discuss Agents, our attention habitually falls on capability: can it understand, can it reason, what level of intelligence does it have. But what the name originally wanted to remind us of is something else: what is it allowed to touch, and whose fault is it if it touches the wrong thing.
This isn’t nitpicking. If you look at what has actually happened over the past two years, you’ll find that almost none of the problems were “the Agent wasn’t smart enough”—they were “no one knew what the Agent did.” An Agent that can automatically call financial systems—the consequences of writing one wrong line of code are on a completely different scale from a programmer writing one wrong line of code.
In the end, the name Agent carries a letter of responsibility within it. Who signed it—that’s the key.
II. What’s the Relationship Between Large Models and Agents
The previous piece covered one thing, and we’ll pick up directly from it here: a large model itself has no hands. The only thing it can do is spit out characters. So-called tool calling is it writing a note in a format, the note being read by a program outside the door, and that program executing it.
The metaphor is: the model is like a consultant sitting in a room with no windows, and its entire communication with the outside world is notes passed in and out through the crack under the door.
An Agent equips this consultant with a full set of things: letting him see the situation, giving him work to do with his hands, letting him remember what was instructed last time, and giving him a letter of authorization that spells out his permissions.
The reverse also holds. Take those things away, and the Agent immediately degrades back into a chatbot. Quite a few products on the market that claim to be Agents—when you actually use them, they’re just a Q&A box. That’s why: it only has that brain, without everything else equipped.
So What Exactly Does the Large Model Contribute
This is the key question.
Automation isn’t new. Scripts, macros, workflow engines, RPA—they can all do things for you, and some do them very reliably. They all share one premise: what to do must be thought of when you write it, and hard-coded.
Notify procurement when a warehouse shipment arrives—that can be hard-coded. But “go through this batch of supplier contracts and see which ones have payment terms unfavorable to us”—that can’t be hard-coded. Because “unfavorable” is a judgment, and the criteria for judgment vary by contract type, payment method, and whether the other party is the buyer or seller—you can’t enumerate them all.
The change the large model brings is this: it moves the judgment of “what to do next” from the code into the model.
The consequence is those two words: general. For the first time, tasks don’t need to be enumerated. Give it a task that was never defined before, and it can look, think, and proceed on its own. Whether it succeeds is another matter, but at least it has the possibility of proceeding.
This is also why Agents only became possible in the past two years. Not because “automation” is new, but because “on-the-spot judgment” can, for the first time, be handed to a machine.
III. An Agent That Can Actually Get Things Done for You Must Have Five Things
With any one of these five missing, it doesn’t count as able to work.
1. It Must Be Able to See the Situation
Every decision the model makes relies on the material currently placed in front of it. Whatever is in that material is all it can base its judgment on.
This material typically includes: what you said, what actions it just took, what results those actions returned, which relevant fragments were pulled from the enterprise knowledge base, and how much of the previous conversation remains.
The previous article called this the “workbench”: the surface is only so big, and all material must be placed on it for it to see. What to put on it and what to leave off is itself a design problem.
2. It Must Be Able to Advance Round by Round
An Agent doesn’t finish a task in one step. It’s more like a person doing something: take a step, look at the result, decide the next step, take another step.
This “take a step, look, take another step” cycle is the heartbeat of an Agent. Without it, it’s not an Agent—it’s just a function call.
How many rounds to run, when to stop, when to change direction—these are all control problems, and they’re the hardest part of building an Agent.
3. It Must Actually Be Able to Take Action
This is the most intuitive part: an Agent must be able to actually do things, not just say things.
Calling APIs, querying databases, sending emails, creating tickets, modifying records—these are the “hands” of an Agent. The more complete the set of hands, the more it can do.
But hands also bring risk. The more it can touch, the more damage it can do when it makes a mistake. This is why the “letter of authorization” is so important.
4. It Must Know When to Stop
An Agent that doesn’t know when to stop is dangerous. It will keep running, keep calling tools, keep consuming resources—and might even make things worse.
Stopping conditions include: the task is done, the task can’t be done, a timeout, too many rounds, an error, or a human intervention. A good Agent knows when to say “I’m done” and when to say “I can’t do this.”
5. It Must Know Where Its Boundaries Are
This is the layer that the name “Agent” originally emphasized, and the one most easily overlooked.
What can it touch, what can it not touch; what can it decide on its own, what must be approved by a human; what it does is attributed to whom—these are all boundary questions.
An Agent without boundaries is not an Agent, it’s a runaway process. And the consequences of a runaway process in an enterprise are much more serious than in a personal tool.
Summary
An Agent is not a smarter model. It’s a model plus a set of things that let it actually work: eyes to see, hands to act, memory to remember, a loop to advance, and a letter of authorization that defines its boundaries.
The word itself reminds us: the core of an Agent is not “how smart it is,” but “what it’s authorized to do, and who’s responsible for what it does.”
Once you understand this, you understand half of what an Agent is.