knowledge > tech
Almost Everything That Happens When You Ask an LLM a Question
The epic adventure of getting just one token out, through API requests, embeddings, tensors, transformers, GPUs, KV cache, HBM, and token generation.
This article follows what happens after sending the question Are tomatoes fruit? to an LLM. It looks at the inputs and outputs at each step, why particular operations are needed, and what data and metadata each step uses.[1]1The server-processing flow in this article is based on the previously read Reference: How an AI Token Travels Through a Data Center and its original source. This does not mean every service must implement the process in this order or with this structure. ↩ ..
My main reason for writing this is to understand the process myself. I also want to think about the fact that an LLM is not so much a being with a personality or preferences as a machine that takes language, turns it into numbers, and produces the most likely next piece of language.
As models handle more kinds of data beyond language, it becomes easier to mistake them for beings with personalities or subjective views. I think it is important to learn how they work and remember that they are machines, especially for people without much computer science literacy.
- The application client sends a question to a server.
- The application server builds a request for the vendor API.
- The vendor server accepts the request and prepares the model input.
- The tokenizer converts text into token IDs.
- Requests to be processed together are grouped into a batch on the GPU.
- The token array is converted into a GPU input tensor.
- The GPU looks up embedding vectors using the token IDs.
- Transformer layers update the input vectors.
- K and V from each layer are stored in a KV cache for reuse when generating the next token.
- The final vector is used to score next-token candidates and choose one.
- The selected token is restored to text and sent to the screen.
- The newly generated token is fed back in to calculate the next token.
- The request ends at a termination token, and its resources are cleaned up.
0. Background: An LLM predicts the next token from the tokens before it
A large language model (LLM) is an autoregressive language model that predicts one next token based on the tokens before it. For example, when continuing I ate an apple, the model scores possible next tokens and selects one. Suppose it selects . in this example. A chat model extends an answer one token at a time in the same way, using the question as context.[2]2Hugging Face, Causal language modeling. Describes how autoregressive language models are trained to predict the next token from preceding tokens. ↩ huggingface.co
The model does not have a completed answer waiting in advance. It selects the first token, includes that token in the input for the next iteration, and selects the next token. It repeats this process to build the answer.
The following terms and concepts appear in the sequence below.
| Term | Meaning | Role |
|---|---|---|
| Token | A unit of text processed by an LLM. It may be a whole word, part of a word, or punctuation. Special tokens can mark roles or the end of a message. | The model receives tokens in order and predicts the next one. Input and output lengths are usually counted in tokens. |
| Tokenizer | A program that converts text into token IDs. It uses the model’s vocabulary and splitting rules, and can also turn generated IDs back into text. | It turns input text into an integer array the model can process, then converts selected output IDs into readable text. |
| Token ID | A number identifying a token: an integer pointing to one token in the tokenizer’s vocabulary. | The model uses this number to look up an embedding. After choosing an output number, it uses the tokenizer to turn it into text. |
| Embedding | A predefined bundle of numbers for each token ID. In this article’s hypothetical model, ID 7342 corresponds to [0.8, 0.3, 0.1, -0.2]. | Given a token ID, the model retrieves its associated bundle from the embedding table and uses those numbers as the input to its first computation. |
| Tensor | An array in which the model stores and computes numbers. shape tells how many values there are along each dimension; dtype tells how each number is stored.[3]3PyTorch, Tensors. Describes tensor shape, data type (dtype), storage devices, and array operations. ↩ docs.pytorch.org | It groups token IDs or token vectors into a specified form. CPUs and GPUs compute many numbers together according to that form. |
| Vector | A kind of tensor. A one-dimensional array of numbers is called a vector; tensor refers to an array of numbers regardless of its number of dimensions. For example, [0.2, -0.1, 0.7, 0.4] is a vector of length 4. | It represents one token’s state using several values. Model layers compute these values and turn them into a representation for the next layer. |
| Weight | A number in the model that determines how much, and in what direction, input values affect a computation’s result. It is applied to input vectors during inference. | It is used in computations that transform input vectors. In ordinary inference, weights stay fixed while the model computes results from the input. |
| Transformer layer | A computation step that updates each token’s bundle of numbers. Attention combines information from other tokens, and a feed-forward network (FFN) computes each token’s numbers again. | It receives one bundle per token and passes one updated bundle per token to the next layer. Across multiple layers, relationships in the question are reflected in the computation for the last token. |
| Attention | A computation that determines how much information from itself and preceding tokens to include when computing the current token’s vector. | It incorporates information from preceding tokens into the current token’s computation. It does not look ahead at future tokens. |
| Hidden state | The bundle of numbers the model has computed so far for each token. It begins with the embedding, and its values change every time it passes through a transformer layer. | It passes an intermediate result that includes information from earlier tokens to the next layer. The final bundle for the last input token is used to score possible next tokens. |
| Logits | A score assigned to each possible next token. If the model can output 10,000 kinds of tokens, it produces 10,000 scores. These are scores before conversion to probabilities. | The model uses the last token’s hidden state to produce a score for each token, then selects one next output token according to a selection rule. |
| Key-value cache (KV cache) | Memory that stores the key and value bundles needed for attention, computed from tokens that have already been processed. Each transformer layer has its own cache. | When generating the next token, the model retrieves earlier keys and values instead of recomputing them. When it processes a new token, it stores that token’s keys and values too. |
Using these terms, the flow below shows how a language query becomes data a computer can calculate with.
flowchart LR
Q["Conversation input<br/>system · prior turns · question"]
I["Tokens → token IDs<br/>Are tomatoes → 7342"]
E["Embedding vector<br/>[0.8, 0.3, 0.1, -0.2]"]
H["Layer-by-layer hidden states<br/>shape [1, 16, 4]"]
L["Logits → output IDs (repeated)<br/>4512, 27, 9311, 4"]
R["Response text<br/>Yes, tomatoes are fruit."]
Q --> I --> E --> H --> L --> R
Putting the layers responsible for processing each piece of data into this data-flow diagram gives us the following diagram.
flowchart TD
subgraph FIRST["Request and model-input preparation"]
direction LR
A["Application client"] -->|"question · model · conversation ID"| B["Application server"]
B -->|"messages · options"| C["Vendor API<br/>conversation format · tokenization"]
C -->|"token ID tensor [1, 16]"| D["Embedding lookup"]
end
subgraph SECOND["Model computation and response"]
direction LR
E["Transformer · attention"] -->|"final hidden state [1, 4]"| F["Logit calculation · token selection<br/>logits [1, 10000]"]
F -->|"output ID → response text"| G["Text restoration · display"]
end
FIRST -->|"embedding vectors [1, 16, 4]"| SECOND
The model name, token splits, IDs, vector values, and server names are hypothetical values for explanation, not an actual API or execution trace from a particular model. The input and output data at each step are also shown as illustrative JSONC.
1. The application client sends a question to a server
The user types Are tomatoes fruit? into the input box and selects a model. In this example, the client sends the question and model name it received from the screen to the application server.
{
"input": "토마토는 과일이야?",
"model": "example-llm"
}
2. The application server builds a vendor API request
The application server reads the system instructions and previous conversation, then adds the newly received user question at the end. It sends these messages along with API settings, such as a generation-length limit and streaming option, in the format required by the selected vendor. Generation settings may include temperature and max_output_tokens.
Here we assume a stateless chat API that requires the conversation history for each request. With this approach, the vendor API does not automatically append the conversation from a previous request, so the application server must assemble the context needed for this computation.[4]4Anthropic, Messages guide. Describes the Claude Messages API pattern of sending conversation content in a messages array with each request. The article uses this as an example of a stateless API. ↩ platform.claude.com Some APIs can continue server-side state using a previous response ID or conversation ID, so not every vendor API is stateless.[5]5OpenAI, Conversation state guide. Describes continuing conversation state through a previous response ID or conversation ID. ↩ developers.openai.com
Below is a hypothetical chat API request with system instructions in a system message. Actual field names and arrangements differ by vendor. For example, the Claude Messages API puts system instructions in a top-level system field rather than in messages with a system role.[6]6Anthropic, Messages API documentation. Describes passing system instructions in the top-level system field in the Claude Messages API. ↩ platform.claude.com
{
"model": "example-llm",
"messages": [
{
"role": "system",
"content": "짧게 답하세요."
},
{
"role": "user",
"content": "안녕"
},
{
"role": "assistant",
"content": "안녕하세요"
},
{
"role": "user",
"content": "토마토는 과일이야?"
}
],
"max_output_tokens": 16,
"stream": true
}
3. The vendor server accepts the API request and prepares model input
The vendor API server checks the request format, authentication and permissions, usage limits, and other requirements. For this article, assume it assigns the ID req_7f3a to an accepted request and stores its processing state. The ID lets the server match the request with its response after computation.
For this explanation, assume the vendor converts messages into the conversation format used to train the model. The prompt_text below is a hypothetical string that makes the result easy to read. An actual implementation may convert directly to token IDs without creating a separate string.
{
// A hypothetical store where the vendor API looks up request state by ID
"request_states": {
// Use this key to look up this request's state
"req_7f3a": {
// Waiting for model computation
"status": "queued",
"max_output_tokens": 16,
// Illustrative string with the model-specific conversation format applied
"prompt_text": "<system>짧게 답하세요.</system><user>안녕</user><assistant>안녕하세요</assistant><user>토마토는 과일이야?</user><assistant>"
}
}
}
4. The tokenizer converts text into token IDs
The tokenizer splits the prepared input into tokens and returns the ID corresponding to each one. Assume the hypothetical vocabulary splits the system instructions, prior conversation, current question, and response-start marker into the following sixteen tokens.
{
"request_id": "req_7f3a",
// Includes role markers for system instructions and prior conversation
"tokens": ["<system>", "짧게 답하세요.", "</system>", "<user>", "안녕", "</user>", "<assistant>", "안녕하세요", "</assistant>", "<user>", "토마토는", " 과일", "이야", "?", "</user>", "<assistant>"],
// Integer IDs in the same order as the tokens
"token_ids": [10, 201, 11, 1, 610, 2, 3, 611, 12, 1, 7342, 55, 188, 9, 2, 3],
// Input length, including special tokens
"input_length": 16
}
5. The GPU batch is selected
If several replicas of the model can handle the request, the system must decide which server will compute it. A load balancer and inference router assign the request based on factors such as available servers, load, model, and cache state. A scheduler chooses which assigned requests to include in the next GPU execution.
Multiple requests can be calculated together by grouping them into a batch instead of sending one request at a time to the GPU. Users can ask questions at different times while sharing the same GPU.
{
// Hypothetical server selected for execution
"replica": "llm-07",
// Requests selected for this GPU computation
"batch": [
{
// Key the serving program uses to connect request state and input data
"request_id": "req_7f3a",
// This request's item number in the current batch
"batch_index": 0,
// The full input has not been processed yet
"phase": "prefill",
// Number of input tokens in this request
"token_count": 16
}
]
}
To make the numbers easier to follow, the rest of the example assumes that this is the only request in the batch.
6. The token array is converted into a GPU input tensor
The token IDs from step 4 are prepared as an input tensor the GPU can compute with. The input has 16 integers: [10, 201, 11, 1, 610, 2, 3, 611, 12, 1, 7342, 55, 188, 9, 2, 3]. The model input also marks token order and which tokens to include in the computation.
The tokenizer may return a tensor directly, or the program that runs the model on the vendor server may create it.
{
// Sixteen tokens in one request
"input_ids": [[10, 201, 11, 1, 610, 2, 3, 611, 12, 1, 7342, 55, 188, 9, 2, 3]],
// Position of each token within the input
"position_ids": [[0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15]],
// Real tokens are 1, padding is 0. There is no padding in this input
"attention_mask": [[1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1]],
// Illustrative metadata: one inner array containing 16 IDs
"shape": [1, 16],
// Storage format for each number in this example
"dtype": "int64"
}
Token IDs alone do not say which position a token occupies, so position information is also included in the computation. position_ids explicitly represents the positions in this example. The way position information is applied depends on the model architecture.
An attention_mask distinguishes actual tokens from padding. For example, if an input has three tokens and two extra slots are filled with padding, the mask is [1, 1, 1, 0, 0]. There is no padding in this input, so every value is 1. This mask marks whether each slot contains an actual token.
shape and dtype are metadata describing a tensor’s shape and number-storage format. For example, the shape of [[0.8, 0.3, 0.1, -0.2], [0.2, 0.0, 0.4, 0.1]] is [2, 4]: there are two rows, each containing four values. dtype: float32 means that each number is stored as a 32-bit floating-point value.
7. The GPU looks up embedding vectors using the IDs
The serving program sends the input tensor to the GPU where the model runs. The model’s weights must already be available in GPU memory. PyTorch’s embedding operation uses a token ID as an index to look up the corresponding vector in the embedding table.[7]7PyTorch, Embedding documentation. Describes using an integer token ID as an index to retrieve its vector from an embedding table. ↩ docs.pytorch.org The Transformer paper describes the relationship between input embeddings and the output layer.[8]8Vaswani et al., Attention Is All You Need, §3.4. Describes the relationship between input embeddings and output-layer weights. The original paper covers an encoder–decoder architecture; this does not mean all modern LLMs use that architecture unchanged. ↩ arxiv.org
From this point on, most model computation takes place on the GPU.
{
// Embedding vectors corresponding to the 16 input tokens
"embeddings": [[
[0.1, 0.0, 0.2, 0.1],
[0.4, 0.1, 0.0, 0.3],
[0.1, 0.1, 0.1, 0.0],
[0.2, 0.0, 0.1, 0.2],
[0.5, 0.2, 0.0, 0.1],
[0.1, 0.1, 0.0, 0.7],
[0.3, 0.2, 0.0, 0.1],
[0.2, 0.4, 0.1, 0.3],
[0.1, 0.2, 0.1, 0.0],
[0.2, 0.0, 0.1, 0.2],
[0.8, 0.3, 0.1, -0.2],
[0.6, 0.4, 0.2, 0.1],
[0.0, 0.2, 0.5, 0.3],
[0.1, 0.1, 0.0, 0.6],
[0.1, 0.1, 0.0, 0.7],
[0.3, 0.2, 0.0, 0.1]
]],
// One request, 16 tokens, four numbers per token
"shape": [1, 16, 4],
"dtype": "float32"
}
The model’s embedding layer looks up the vector associated with each token ID. These sixteen vectors are the transformer’s first input. No answer token has been produced yet, and the vectors do not yet reflect the preceding conversation. In this example, 토마토는 at position 10 corresponds to [0.8, 0.3, 0.1, -0.2], and <assistant> at the final position, 15, corresponds to [0.3, 0.2, 0.0, 0.1].
8. Input vectors are updated by each transformer layer
In step 7, the token IDs were used to look up an embedding table, producing a vector of four numbers for each token. The transformer now computes with these vectors. It does a great many complicated computations.
To state the obvious once: the model has already been trained. When the model was prepared in advance, it used the beginnings of sentences to predict the next token and adjusted the embedding table and weights to give higher scores to the tokens that actually followed. For example, if 과일 followed 토마토는 in a training sentence, the values were adjusted to give 과일 a higher score in that context.
Each value adjusted during training is called a parameter. If looking up a token ID in the embedding table returns [0.2, -0.5, 0.8, 0.1], each of these four values is one parameter. Each value stored in a weight matrix is also a parameter. A 10B model has about 10 billion parameters like these. During inference, when answering this question, the values are held fixed; the model uses them to produce intermediate results and next-token scores from the input.[9]9Hugging Face, Causal language modeling. Describes comparing the predicted next token with the correct token and adjusting model parameters during training. ↩ huggingface.co
Prefill is the step in which the entire input is passed through all transformer layers. The vector for each token is updated and passed to the next layer. Here we assume a hypothetical model with two layers to make the explanation easier. After both layers have run, the vector for the final input token is used to choose the first answer token.
The diagram below shows how the input vectors for one layer become result vectors, which are passed to the next layer.
flowchart LR
A["Vectors for 16 input tokens"] --> B["Q · K · V in layer one"]
B -->|"attention · FFN and other computations"| C["16 result vectors"]
C --> D["Q · K · V in layer two"]
D -->|"attention · FFN and other computations"| E["16 final vectors"]
Each layer applies learned weight matrices to its input vectors to make intermediate vectors called queries (Q), keys (K), and values (V). It compares the current token’s Q with the K of itself and preceding tokens to get scores for the tokens. Softmax turns these scores into ratios that add up to 1. The model uses those ratios to incorporate each token’s V and produce the attention result. In other words, V itself is not the result vector.
Within a layer, the model uses Q, K, and V to produce an attention-result vector through scaled dot-product attention, combining head results, a feed-forward network, residual connections, and normalization. That is hard to understand. I tried to make sense of it while reading too, but gave up. Just remember that various computations happen and, when they are done, one result vector comes out for each input token.[10]10Vaswani et al., Attention Is All You Need, §§3.1–3.4. Describes Q, K, V, scaled dot-product attention, position-wise feed-forward networks, residual connections, and normalization. The original paper describes an encoder–decoder architecture; this article refers to it for the computational principles. ↩ arxiv.org
These result vectors are the layer’s hidden states. The second layer receives the sixteen vectors from the first layer and computes them again using its own weights. Passing through one layer does not produce one answer token. The final vector for the last input token, after passing through all layers, is used to choose the next output token.
Having multiple transformer layers lets a later layer use the results of earlier computations.[11]11Elhage et al. (2021), A Mathematical Framework for Transformer Circuits, “Three Kinds of Composition.” Analyzes how computations from earlier layers can be used in later layers’ Q, K, and V computations. The Minsu-and-apple example in the text illustrates this principle. ↩ transformer-circuits.pub This is central to the autoregressive behavior described above. For example, in Minsu bought an apple. He ate it., if an earlier layer reflected the relationship between it and apple, a later layer can use that result to compute more about the relationships in the sentence. Each layer uses different weights. Adding more layers does not guarantee a better result. It is a bit like solving a fill-in-the-blank problem by looking at the surrounding sentence several times, reasoning about what the answer might be.
Large LLMs compute much longer vectors across many more layers than this example. They also pass each new token through every layer to select the next token as they continue an answer. They repeatedly read many weights and perform matrix operations, which is why running a model consumes enormous computing power and memory.
Below is a hypothetical excerpt showing only the vector for the final input token among the final results from each layer. No answer token has been selected yet. Step 9 looks at how each layer stores K and V; step 10 uses the final layer’s vector to select the first answer token.
{
// Full result from each layer: 16 vectors of length 4
"hidden_state_shape": [1, 16, 4],
// Only the result at position 15, the final input token, is shown below
"sample_position": 15,
// Result from layer one and input to layer two
"layer_1_hidden_state": [0.5, 0.4, 0.2, 0.1],
// Result from layer two, used to calculate output scores in step 10
"layer_2_hidden_state": [0.9, 0.1, 0.5, 0.2]
}
9. K and V are stored in the KV cache for reuse when generating the next token
The KV cache stores K and V previously computed from input tokens. Each transformer layer has its own cache because each layer has different weights and results.[12]12Hugging Face, How caching works. Describes the structure of K and V caches for each transformer layer. ↩ huggingface.co
{
// Key that connects the cache to this request
"request_id": "req_7f3a",
// Processing of the initial input tokens is complete
"phase": "prefill_complete",
// Each layer stores its own computation separately
// The actual K and V arrays are omitted; only the number of stored tokens is shown
"cache_by_layer": {
// K and V computed from the 16 input tokens in layer one
"layer_1": {
// Positions 0–15, 16 tokens total
"cached_token_count": 16
},
// K and V computed from the 16 input tokens in layer two
"layer_2": {
// The same 16 positions' K and V are stored in this layer too
"cached_token_count": 16
}
}
}
When prefill finishes, each layer’s cache contains the K and V for the 16 tokens in the initial input. After selecting the first answer token 네, the next iteration passes 네 through each layer as a new input. Each layer reuses the K and V for the previous 16 tokens and adds the K and V for 네. The cache now contains values for 17 tokens.[13]13Hugging Face, How caching works. Describes reusing past K and V and adding new K and V to the cache when processing a new token. ↩ huggingface.co The model still needs to read the stored K and V to calculate attention for the new token.
Prefix caching reuses K and V computed for one request in another request. If the beginning of a new request matches a cached prefix under the same model and reuse conditions, the system treats it as a hit and retrieves it. Otherwise it treats it as a miss, computes that prefix, and stores it.[14]14vLLM, Automatic Prefix Caching. Describes reusing KV blocks when the beginning tokens of requests match. ↩ docs.vllm.ai
For example, suppose an earlier request was Answer politely. Are tomatoes fruit? and a new request is Answer politely. Are apples fruit?. The K and V corresponding to Answer politely. can be reused. If that cache is still available, it is a hit; otherwise, this part has to be computed again. The question differs between the requests, so it is processed anew.
The presence of a single matching word or a sentence with a similar meaning is not enough for reuse. The matching portion must begin at the start of the request, with the same token IDs. This is important when optimizing costs in AI applications: if frequently changing dates or user information are placed near the end of a message, the fixed instructions at the beginning are more likely to be reusable from cache.
The KV cache helps improve response speed and throughput by avoiding recomputation of K and V that have already been calculated.[15]15NVIDIA, Optimizing Inference for Long Context and Large Batch Sizes with NVFP4 KV Cache. Describes how KV-cache memory use and read/write demands affect inference performance with long contexts and large batches. ↩ developer.nvidia.com It uses memory in return. As context gets longer or more requests are processed at once, each layer has more K and V to store and read. That makes both the capacity and read/write bandwidth of a GPU’s high-bandwidth memory (HBM) important.[16]16How an AI Token Travels Through a Data Center. Describes how GPU memory capacity and bandwidth affect inference performance. ↩ www.datagravity.dev This is why HBM is considered a key part of running LLMs, and why SK hynix is making a lot of money.
10. The final result vector is converted into scores
Finishing prefill does not immediately produce text. The model still has only vector values, which only the computer can understand. It still needs to choose an actual text token using the final vector for the last input token.
Assume the example model’s output list contains 10,000 items. The list includes whole words, word pieces, punctuation, and so on, with each item associated with a token ID. In this example, suppose ID 4512 corresponds to 네.
The transformer’s final input vector has length 4 and shape [1, 4]. The language-model output layer (LM head) calculates a score for each item in the list, producing an array of scores with shape [1, 10000]. For example, if the value for ID 4512 is 4.2, then the logit for 네 is 4.2. Each slot corresponds to the score for one token ID. This 4.2 does not mean there is a 4.2% chance of 네 appearing.
Softmax converts these logits into probabilities. The probabilities represent the likelihood that each token will be output next in the current context, and add up to 100%. For example, if softmax gives 네 a probability of 70% and 아니 30%, sampling draws one next token according to those proportions. 네 is drawn more often, but it is not guaranteed to be selected every time. Depending on the generation settings, the model may choose the highest-scoring token or sample from the probability distribution. Temperature and top-p, a cumulative-probability threshold for selecting candidates, affect the distribution or selection. Here we assume ID 4512, 네, is selected as the first token.[17]17Hugging Face, Generation strategies. Describes greedy selection, sampling, temperature, and generation settings such as top-p. ↩ huggingface.co
{
// Shape of the model's output score array
// The actual 10,000 scores are omitted; only the array shape is shown
"logits_shape": [1, 10000],
// ID of the first output token selected in this example
"selected_token_id": 4512,
// Value written as a human-readable string
"selected_token_text": "네"
}
One first output token has been selected based on the sixteen input tokens. The first token of the final response 네, 과일입니다. is 네.
11. The first token is restored to text and sent to the screen
The serving program passes the selected ID to a detokenizer to restore it as text. The vendor’s response server sends the text fragment to the application server associated with the request ID. The application server passes it to the client for conv_42. With streaming, the client can receive this fragment before the full answer is complete.
{
// Illustrative streaming event name
"event": "text_delta",
// ID connecting request and response
"request_id": "req_7f3a",
// Text to add to the screen this time
"delta": "네"
}
The client displays 네 in the previously empty answer area. This is the moment when the user first feels the wait is over. The time it takes to output this first token after the request is called time to first token (TTFT). Naturally, it is affected not only by the model’s prefill but also by network transfer, server queueing, and preprocessing.
12. Decode feeds the newly generated token back in
After sending the first token, 네, to the screen, the model calculates the next token. At the moment it selected 네, the cache did not yet contain this token’s K and V. In the next iteration, it processes ID 4512 as a new input and reuses the K and V for the previous sixteen input tokens from the cache.
{
// New token for this iteration: the just-selected "네"
"input_ids": [[4512]],
// Positions 0–15 came before; the new position is 16
"position_ids": [[16]],
// Attend to the 16 past tokens in the cache together with the new token
"attention_mask": [[1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1]],
// Shape of the new input ID tensor to be calculated
"input_shape": [1, 1],
// Number of tokens in the cache before this computation
"past_token_count": 16
}
The new input contains one token, but attention can refer to the sixteen past tokens in the cache plus the new token, for a total of seventeen. That is why this example’s attention_mask has seventeen values, all 1 because they are actual tokens. This mask indicates padding; the causal mask is responsible for preventing the model from looking at future tokens.
The model converts the newly selected 네 into an embedding and passes it through every layer. Each layer updates the new token’s vector and refers to past tokens’ K and V in the cache. The vector from the final layer is used to select the next token.
| Iteration | New input being processed | Past tokens’ K and V reused | Cached token count after processing | Response accumulated on screen |
|---|---|---|---|---|
| Prefill | The original 16 input tokens | No previous cache. The 16 input tokens are processed for the first time | 16 | 네 |
| Decode 1 | 네 / 4512 | K and V for the original 16 input tokens | 17 | 네, |
| Decode 2 | , / 27 | K and V for the original 16 inputs plus 네 | 18 | 네, 과일입니다. |
| Decode 3 | 과일입니다. / 9311 | K and V for the original 16 inputs plus 네 and , | 19 | 네, 과일입니다. |
Treating the long string 과일입니다. as one token is also a setting of the hypothetical vocabulary. An actual tokenizer may split it into several tokens, so we assume one token here for ease of explanation.
The selected token is added to the cache when it is processed as new input in the next iteration. So even after 네 is selected during prefill, the cache remains at 16 tokens; it grows to 17 only after 네 is processed in Decode 1.
13. The end is signaled and request resources are cleaned up
When the model selects the termination token <EOS> in the final iteration, the server treats it as a signal to stop generating rather than displaying it as text. Generation can also stop when the length limit is reached or when the user cancels.
In this example, the model selected [4512, 27, 9311, 4]. The final ID, 4, is the termination marker. The string shown to the user is 네, 과일입니다..
{
// Hypothetical event indicating that generation has finished
"event": "completed",
// Same ID as the one first assigned to the request
"request_id": "req_7f3a",
// Generation stopped after selecting the termination token
"finish_reason": "eos",
// Final response shown on the screen
"text": "네, 과일입니다.",
// All selected IDs in this example; separate from service-specific billing fields
"generated_token_ids": [4512, 27, 9311, 4]
}
The vendor server updates the request state to completed and records usage and latency. The application server stores the user’s question and the completed assistant response in the conversation history for conv_42, so it can construct the next request.
The scheduler excludes the completed request from the next batch. Per-request caches and execution resources that are no longer needed are released or managed according to the cache policy. The API server also cleans up its connection and request state according to the service’s retention policy.
Notes
- 1.The server-processing flow in this article is based on the previously read Reference: How an AI Token Travels Through a Data Center and its original source. This does not mean every service must implement the process in this order or with this structure. ↩..
- 2.Hugging Face, Causal language modeling. Describes how autoregressive language models are trained to predict the next token from preceding tokens. ↩huggingface.co
- 3.PyTorch, Tensors. Describes tensor shape, data type (dtype), storage devices, and array operations. ↩docs.pytorch.org
- 4.Anthropic, Messages guide. Describes the Claude Messages API pattern of sending conversation content in a messages array with each request. The article uses this as an example of a stateless API. ↩platform.claude.com
- 5.OpenAI, Conversation state guide. Describes continuing conversation state through a previous response ID or conversation ID. ↩developers.openai.com
- 6.Anthropic, Messages API documentation. Describes passing system instructions in the top-level system field in the Claude Messages API. ↩platform.claude.com
- 7.PyTorch, Embedding documentation. Describes using an integer token ID as an index to retrieve its vector from an embedding table. ↩docs.pytorch.org
- 8.Vaswani et al., Attention Is All You Need, §3.4. Describes the relationship between input embeddings and output-layer weights. The original paper covers an encoder–decoder architecture; this does not mean all modern LLMs use that architecture unchanged. ↩arxiv.org
- 9.Hugging Face, Causal language modeling. Describes comparing the predicted next token with the correct token and adjusting model parameters during training. ↩huggingface.co
- 10.Vaswani et al., Attention Is All You Need, §§3.1–3.4. Describes Q, K, V, scaled dot-product attention, position-wise feed-forward networks, residual connections, and normalization. The original paper describes an encoder–decoder architecture; this article refers to it for the computational principles. ↩arxiv.org
- 11.Elhage et al. (2021), A Mathematical Framework for Transformer Circuits, “Three Kinds of Composition.” Analyzes how computations from earlier layers can be used in later layers’ Q, K, and V computations. The Minsu-and-apple example in the text illustrates this principle. ↩transformer-circuits.pub
- 12.Hugging Face, How caching works. Describes the structure of K and V caches for each transformer layer. ↩huggingface.co
- 13.Hugging Face, How caching works. Describes reusing past K and V and adding new K and V to the cache when processing a new token. ↩huggingface.co
- 14.vLLM, Automatic Prefix Caching. Describes reusing KV blocks when the beginning tokens of requests match. ↩docs.vllm.ai
- 15.NVIDIA, Optimizing Inference for Long Context and Large Batch Sizes with NVFP4 KV Cache. Describes how KV-cache memory use and read/write demands affect inference performance with long contexts and large batches. ↩developer.nvidia.com
- 16.How an AI Token Travels Through a Data Center. Describes how GPU memory capacity and bandwidth affect inference performance. ↩www.datagravity.dev
- 17.Hugging Face, Generation strategies. Describes greedy selection, sampling, temperature, and generation settings such as top-p. ↩huggingface.co