The Truth About LLM Token Output and Tool Calling
The one that stuck with me most was a pre-launch acceptance test for a client.
The demo went smoothly until a business colleague casually asked, “Can you check last month’s inventory for the East China region?” On screen, the agent replied, “Sure, let me look that up for you,” and then — nothing. That sentence just sat there on the screen, not a single number appeared.
Everyone in the meeting room was waiting for an explanation. And the explanation was actually very basic — so basic that we usually can’t be bothered to mention it: A large model has no hands. The only thing it can do is spit out one word at a time.
After that, I figured something out: a lot of the confusion, misunderstanding, and unrealistic expectations around large models all trace back to two questions. How does it “talk,” and what’s actually going on when it supposedly “does things”?
This article covers only those two things. As for how agents break down tasks, divide labor, and get kept in line — that’ll be a separate piece. Not going into it here.
One-Minute Overview
- A large model does exactly one thing: predict the next token. It’s not a system that thinks through an entire passage before outputting it — it’s a machine that plays word-chain, nonstop.
- The “character-by-character typewriter effect” you see isn’t a frontend animation — it’s the model’s actual working rhythm. It really is popping out one word at a time.
- The model cannot call tools. So-called tool calling is the model outputting a specially formatted piece of text, which a program outside the model reads and then executes. The model just writes a note; the one who actually does the work is someone else outside the door.
- This distinction isn’t just wordplay. It determines which layer permissions, isolation, and auditing — the things enterprises care about most — should be implemented at, and it explains why a model will earnestly say “I’ve already looked that up for you.”
- AgentSteamer is an enterprise-grade AI agent platform developed by Shanghai Immersivalley Information Technology Co., Ltd. It supports full private deployment and isn’t tied to any specific large model.
1. The Model Does Exactly One Job: Guess the Next Word
Let’s break it down to the simplest level.
You give it “The weather today is really,” and it outputs “nice.” You feed “The weather today is really nice” back in, and it outputs “,”. And so on, until it outputs a special symbol meaning “I’m done.”
That’s the whole process. Translations, code, legal opinions, comforting words — they’re all strung together one sentence at a time like this. The industry calls this autoregressive generation. The name sounds intimidating, but it just means “using your own previous output as the next round’s input.”
There’s one thing that has to be explained clearly first: token.
The model doesn’t recognize Chinese characters or letters. What it recognizes is a string of numbers. So text first gets chopped up and translated into numbers, the model processes them, and then they get translated back into text. Each little piece that comes out of the chopping is called a token.
You can think of it as segmenting an article into phrases, except it segments more finely and more mechanically than a human would. In Chinese, one character might be one token, or two characters together might count as one — it depends on how the model’s tokenizer vocabulary is defined. Same goes for English: “understand” might be three tokens, or it might be two. This has a very practical consequence: the same meaning expressed in different languages costs a different number of tokens, and therefore a different cost.
So how does the model “guess”?
It’s not looking up a dictionary. A closer analogy is a room full of people raising their hands to vote. Given the preceding text, the model assigns a score to every possible word in its vocabulary — the higher the score, the more likely it gets picked — and then it draws lots according to those scores.
The temperature parameter controls the scale of that lottery. Turn the temperature down, and it basically only picks the highest-scoring option, speaking steadily and predictably; turn it up, and long-shot options get a seat at the table too, making the text livelier — and more erratic. The same model, at different temperatures, takes on a different personality.
This explains one of the most frequently asked questions: Why does the same question asked twice get different answers? Because it was never retrieving a single correct answer — it was drawing lots.
It doesn’t “think it through, then write” — it “thinks as it writes”
This sentence is the prerequisite for understanding everything that follows.
The model spits out a word, and that word immediately becomes the next round’s input. In other words, when it writes the second sentence, it’s reading the first sentence as established fact.
There are two consequences, and both matter.
First, it can’t go back and revise what it’s already written. Words already spat out are like words already spoken — they can’t be taken back. So once it starts off on a wrong track, it’ll only keep drifting further down that track, further and further off.
Second, its “thinking” and its “output” are the same process. When a human writes something, there’s a rough framework in their head first, then they put pen to paper; the model has no such framework — it grows the framework step by step through “the next word.” This is also why having it write out its reasoning process first usually improves the accuracy of its answer — that process itself becomes the preceding text it reads afterward.
While we’re at it: where hallucinations come from
Many popular explanations describe hallucinations as “the model making things up,” but that’s not quite accurate — it makes it sound like an attitude problem.
The mechanism behind hallucination is actually very plain: at every step, it’s doing the same thing — making the next word look as plausible as possible. What it optimizes for is “does it sound right,” not “is it correct.” These two goals align the vast majority of the time, but they diverge wherever the model doesn’t actually have the facts.
Here’s an analogy. Someone with excellent linguistic instincts who has never visited the site can, from experience, write a research report with a beautiful format, accurate terminology, and detailed data. They’re not deliberately fabricating — they’re just reproducing what “a qualified report should look like.”
2. Tokens Mean Three Things in Engineering
Once you understand that tokens are the unit of measurement, a lot of product design choices suddenly make sense.
First, it’s the billing unit. Charges are based on tokens, not on number of requests — like a taxi meter running on distance, not on trips. The same task might cost differently in Chinese versus English; having it ramble on versus giving a direct conclusion also makes a big cost difference.
Second, it defines how much the model can see at once. This upper limit is called the context window. Your input, its output, tool returns, knowledge base retrievals — all of it has to squeeze into this space.
Think of it as a workbench. The tabletop is only so big, and all materials have to be laid out on it for the model to see them. The model doesn’t have “memory” — every round, it re-reads the entire tabletop from scratch.
Third, when the tabletop is full, things have to be cleared off. The common approach is to drop the earliest items, or to drop by importance. So when you ask it “what was the third requirement I just mentioned?” and it can’t answer, it’s very likely not being disobedient — that item has already been cleared out.
(There’s also a mechanism directly related to cost called Prompt Caching: for parts that repeat every round — like a long system prompt — the system records the already-computed results in a notebook, and next round it just flips to that page instead of recomputing. Saves time, and saves money.)
With these three points in mind, one phenomenon becomes easy to understand: a lot of “the model got dumber” complaints aren’t actually the model’s fault — the context got blown out.
3. How This Skill Was Trained
The model’s ability to guess words isn’t innate — it’s built up through three stages of training.
The first stage is pretraining. Take massive amounts of text and have it repeatedly do one thing: mask a word and guess what it is. Do this hundreds of millions of times. What this stage produces isn’t knowledge — it’s linguistic instinct, an intuition for “how a sentence should continue.” The more it’s read, the sharper that instinct gets.
The second stage is instruction fine-tuning. Just being able to continue a conversation isn’t enough — it also needs to follow instructions. This stage uses large numbers of “question plus ideal answer” examples to teach it: when someone makes a request, how should you respond? It goes from “can talk” to “can answer.”
The third stage is alignment. Teaching it when to refuse, what not to say, and what tone is more appropriate. You can think of it as onboarding training — the first two stages gave it business capability, and this stage gives it a code of conduct.
These three stages share one thing in common, and it’s worth calling out separately: all of them train “what kind of text to output.” None of them teaches it to query a database, connect to an API, or open a file. Not a single word of it.
This sets up the most critical premise for the next section.
4. The Model Has No Hands: The Truth About Tool Calling
At this point, we can answer the frozen-screen scenario from the beginning.
Step one: it has only one exit.
The only form in which a model outputs anything to the outside world is a stream of text. There is no second channel, no “execute” action, and no finger that can reach into your system.
Step two: the tool list is fed to it, not built in.
Before a conversation begins, the caller inserts a description into the input: which tools you can use, what each tool does, which parameters to fill in, and what format those parameters take. Sometimes a few usage examples are attached as well.
This is like handing a new employee a menu. He can’t order anything that isn’t on the menu; if the menu is vague, he can only guess.
Step three: when it decides to use a tool, what it outputs is still text.
Note that this is where confusion is most likely. It isn’t “calling” anything—it’s writing out a request form in a specified format.
Suppose it needs to check inventory. What it spits out looks roughly like this:
The first word is a marker meaning “I’m about to use a tool,” followed by the tool’s name, such as “check inventory,” and then the parameters, such as “SKU A-1024, warehouse East China.”
This text doesn’t read like human language because it was never meant for humans—it’s meant for a program to read, and the industry generally uses a format like JSON. For the model, writing this passage and writing “the weather is nice today” are the same thing: guessing the next token.
Step four: what actually executes is a program outside the model.
This text is caught and parsed by an external program, which then connects to the database, calls the API, and gets the result.
Step five: the result is fed back in.
The query result is written back into the model’s input. Only after reading “inventory: 320 units” does it have a basis for saying the next sentence: “Last month, East China inventory was 320 units, of which A-1024 has 46 remaining.”
Step six: loop.
If it needs to look up another number, it writes another request form. Write a form → execute → feed the result back → write another form—this loop is the heartbeat of every agent. Who orchestrates this loop and when it should stop to ask a human is the subject of the next article.
Let’s tie these six steps together with one image:
The model is like a consultant sitting in a room with no windows. Its entire communication with the outside world consists of notes passed in and out through the crack under the door. It can read words and write words, and nothing else. The errand-running assistant outside the door is the one who actually does things.
Why this distinction deserves its own section
Because it isn’t an implementation detail—it determines four very practical things.
It can’t “take matters into its own hands.” Every real action must pass through that program outside the door. So “whether to let it through” can be blocked at the door: permissions, approvals, isolated environments, execution logs—all of it lives at this layer. No matter how smart the model is, it can’t get around this, because it was never outside that door to begin with.
It will “pretend” it has already used a tool. Because its ability to write a “request form” and its ability to write a “query result” are the same ability. It’s entirely possible for it to write a correctly formatted form that no one outside the door picks up, and then continue on its own: “I’ve looked it up for you—inventory is 320 units.”
That’s exactly what happened in the meeting room at the beginning. It isn’t lying; it’s fabricating the process and the result together into a coherent whole.
How well the tool description is written directly determines whether it chooses correctly. If the menu says “other,” the waiter can only guess. This is also why writing tool descriptions for a model is work that deserves serious attention, not just filling in a name and calling it done.
Parameters are guessed, not read. Details like date formats, case sensitivity of IDs, and units are frequently wrong. That’s why many systems now force it to fill in boxes rather than giving it a blank sheet to improvise on. This approach is called structured output—essentially, handing it a form with boxes.
V. Three phenomena you can verify yourself
Now that the reasoning is done, here are a few ways to judge for yourself on the spot—more useful than memorizing conclusions.
One: check whether its character output speed is uniform. If it’s as even as an animation, it’s probably a fake effect done by the frontend; genuine streaming output (SSE, which pushes the generation process to the frontend in real time) has fluctuations in speed, because each round of prediction has a different level of difficulty.
Two: ask it a question that requires looking up a number, then watch the process. A good system will lay out the intermediate steps for you: what it said first, which tool it decided to call, what the parameters were, what the tool returned. If it only gives you a final answer and reveals nothing in between, you have no way to judge whether that number was actually looked up or just guessed.
Three: ask the same thing a different way. Pick a slightly different phrasing and ask again. If the tool it calls and the parameters it fills in change a lot, the tool description isn’t clear enough—it’s guessing.
VI. A few common misunderstandings
“Its output speed keeps fluctuating—is the network lagging?”
Most of the time, no. It predicts round by round, and each round has a different level of difficulty, so the speed naturally isn’t uniform.
“The context window is so large—can’t I just stuff an entire book into it?”
You can stuff it in, but the cost is money and time. And there’s a phenomenon called “lost in the middle”: when too much material is piled up, it remembers the beginning and the end relatively well, while the part sandwiched in the middle tends to get ignored. It’s like a stack of papers on a desk that’s too thick—you really can’t recall the one buried at the bottom.
“Does it learn on its own and remember what I’ve said?”
No. After a conversation ends, the model itself hasn’t changed at all—not a single parameter has moved. The reason it feels like it remembers you is that the system reprinted the previous conversation and laid it out on this round’s workbench.
This distinction matters a lot in enterprises: memory is a piece of data, not a capability. If it’s data, you have to answer where it’s stored, who can see it, and how long it’s kept.
“If it says it looked something up, did it definitely look it up?”
Not necessarily. Back to the line from Section IV: writing a request form and writing a result are the same ability. To judge whether it’s real, look at whether the layer outside the door left an execution log—don’t trust what the model says about itself.
VII. What these principles look like when they land in a product
Every step described above corresponds to concrete engineering work. Taking Mofu as an example, let’s go through a few points directly related to this article’s topic.
About character output. Mofu’s LLM node uses streaming calls, pushing tokens to the frontend one by one, so what users see is genuinely being written. At the same time, it pushes the model’s thinking process out as collapsible events, so you can see at which step it decided to call a tool and which one it called. This lines up exactly with the second verification method in point five above.
About tools. The platform is compatible with the MCP protocol and the Agent Skills specification. The “feed the tool list to the model” step described earlier is done automatically here: when you connect an MCP service, Mofu reads its declared parameter structure and dynamically generates the instruction sheet shown to the model. When a tool is actually called, the process is likewise displayed as events, not as a black box.
About execution. All real actions take place in an isolated sandbox, with each session having its own independent directory, separate from the main system. This corresponds to the line from Section IV: “the boundary is pinned to that program outside the door.” The errand-running assistant outside the door is the sandbox in Mofu.
About context. Short-term memory is truncated by the token window; when full, it’s discarded according to a strategy—either dropping the earliest first or dropping by importance. Above that there’s a mid-term memory layer: when token consumption approaches a threshold, the model is asked to summarize the previous conversation into a passage. This is how it solves “the whiteboard is full—which part do we erase?” These mechanisms aren’t flashy, but they determine whether long conversations are actually usable.
About models. Mofu isn’t tied to any specific large model; it connects to any endpoint via the OpenAI-compatible protocol, with OpenAI, Azure, DeepSeek, Qwen, Ollama, and vLLM all within its support range. All the mechanisms discussed in this article are shared by all autoregressive language models, regardless of vendor. Mofu (AgentSteamer) is developed by Shanghai Immersivalley Information Technology Co., Ltd., and also supports deployment in fully offline, air-gapped environments.
VIII. Frequently Asked Questions (FAQ)
Does a large model spit out words one at a time?
Yes. Mainstream large models all use autoregressive generation: each time it predicts only the next token, appends it to the end of what’s already been generated, then predicts the next one, and so on. The typewriter effect you see isn’t a frontend animation—it’s a genuine presentation of this process.
Why does asking the same question twice give different answers?
Because at each step the model outputs a probability distribution, and which word is ultimately chosen involves randomness. This degree of randomness is controlled by the temperature parameter—lower is more stable, higher is more divergent. To make it fully reproducible, you need to set temperature to 0 and fix the random seed.
Can a large model really call tools?
No. The model itself can only output text. So-called tool calling is the model outputting a structured piece of text (usually JSON) in an agreed format, stating which tool it wants to use and what parameters to pass; then a program outside the model parses it and actually executes it, and finally sends the result back to the model as new input. Permission control, sandbox isolation, and execution auditing should all be done at this layer outside the model.
Why does the model say “I’ve looked it up for you” when it hasn’t looked up anything at all?
Because for the model, writing text that resembles a tool call and writing text that resembles a query result are the same thing—both are predicting the next token. It has no ability to distinguish “I made a request” from “the request was executed.” To avoid this, an external runtime needs to verify whether each action actually executed successfully, rather than trusting the model’s own account.
What is a context window? What happens when it fills up?
The context window is the upper limit on the total number of tokens the model can see at once, covering your input, its output, and the content returned by tools. Once the window is full, the system must discard some earlier content—common strategies are dropping the earliest or dropping by importance, and early conversation can also be compressed into a summary. This is why the model “forgets” the beginning of a long conversation.
Does a large model remember what I’ve said?
No. The model’s parameters do not change at all after a conversation ends. The reason it seems to remember you is that the system puts the historical conversation back into the input for the current turn. In engineering terms, memory is therefore a piece of data that needs to be managed, and you have to answer where it is stored, who can see it, and how long it is kept.
Why is the model’s output speed sometimes fast and sometimes slow?
Because it predicts one turn at a time, and each turn has a different level of difficulty, so the time it takes is naturally uneven. If you notice it slowing down noticeably at a certain point, that usually means the prediction for that section is more difficult. This has little to do with network conditions.
Do these mechanisms have anything to do with which company’s large model is used?
Not in any fundamental way. Autoregressive generation, token accounting, context windows, and the “write a note first, then have it executed externally” process in tool calling are mechanisms shared by all mainstream large models, regardless of vendor. AgentSteamer therefore does not bind itself to any specific large model, and connects to any endpoint through the OpenAI-compatible protocol. OpenAI, Azure, DeepSeek, Qwen, Ollama, and vLLM are all within its support scope, and it can also be switched to an inference service built by an enterprise on its own internal network.
IX. Plain-Language Explanations of a Few Terms
- Large Language Model (LLM): A language model trained on massive amounts of text, whose core capability is predicting the next word. Everything it does, including the parts that look like reasoning, is built on this.
- Token: The smallest unit the model uses to process text, and also the unit of measurement for billing and context length. In Chinese, one character may be one token, or two characters may combine into one, depending on how the model tokenizes.
- Autoregressive generation: Each turn outputs only one token, appends it to the end, and then predicts the next one. You can think of it as a word-chain game that never stops.
- Temperature: A parameter that controls the degree of randomness in the output. When the temperature is low, it basically only picks the highest-scoring word; when the temperature is high, less common words also have a chance of being selected.
- Context window: The upper limit on the total number of tokens the model can see at one time, equivalent to the size of its workbench. The content within the window has to be read again every turn.
- Tool calling: The model outputs a piece of structured text expressing “I want to use a certain tool, and these are the parameters,” and an external program executes it and sends the result back to the model. The actual execution happens outside the model.
- Streaming output (SSE): Pushing the character-by-character generation process to the front end in real time, rather than waiting until everything is generated and returning it all at once.
- Structured output: Constraining the model to answer according to given fields and formats, equivalent to handing it a form with cells rather than a blank sheet of paper.
- Hallucination: Content generated by the model that looks plausible but is not actually true. It is not intentional deception, but a byproduct of the goal of “making the next word fit best.”
- Prompt caching: Caching and reusing the computation results of repeated input parts (such as a very long system prompt), saving both time and money.
- MCP (Model Context Protocol): An open protocol that lets external tools describe themselves in a unified way and be connected to by model applications.
Final Words
Let’s go back to that meeting room.
The question from that business colleague can be answered in one sentence: the model has no hands; it only wrote a note, and at the time we forgot to arrange for someone to read the note.
When we reviewed it afterward, we found that the real problem was not the model, but that we had conflated “being able to talk” with “being able to do things.” Many debates about large models in enterprises, when argued to the end, come down to the same thing: what you want is either a mouth that writes better, or a pipeline that can actually be put into practice.
The former can improve just by upgrading the model; the latter depends on engineering. Every time the model gets a new version, the first thing automatically gets stronger; but who the person outside the door crack reading the note should be, what is allowed to be written on the note, and whether a trace should be left after execution—these belong to the second thing, and model upgrades cannot help with them.
AgentSteamer is an enterprise-grade AI agent platform developed by Shanghai Immersivalley Information Technology Co., Ltd. A large part of what we do is turn the mechanisms discussed in this article, one by one, into things enterprises can use with confidence: in AgentSteamer, how tokens flow, how tools are described, where actions are executed, how context is managed, and how models are swapped all have corresponding engineering handling.
As for how agents break down tasks, how they divide work, and how they are kept under control, that is a matter for the next article.
About us: Shanghai Immersivalley Information Technology Co., Ltd. | Enterprise-grade AI agent platform AgentSteamer