Making Tokens, Pt. 5: The Forward Pass

April 30, 2026·Spencer SaldanaUpdated October 5, 2026

The final piece of the Making Tokens series. We have built the chain. Sand to wafer to chip to cluster. The chain ends here, with the user pressing enter and the model producing tokens.

What "one token" actually is

A token is a unit of text the model operates on. It is not a word. The English tokenizer used by most large models breaks text into subword units, so common words ("the", "and") are single tokens while uncommon words ("photolithography") might be split into three or four.

Useful rules of thumb:

  • 1 token ≈ 4 English characters ≈ 0.75 words
  • A 500-word page is about 666 tokens
  • A typical conversational chat response is 200-1000 tokens
  • A long-form essay this length is around 3000 tokens

When a model "produces a token," what is happening at the silicon level is one forward pass through the trained neural network. The prompt is processed first. During standard autoregressive decoding, the model predicts successive tokens while reusing cached attention state from earlier tokens. It does not need to recompute the entire prompt from scratch for each output token. Serving optimizations such as speculative decoding can change how that loop executes.

The arithmetic

For a transformer of N parameters, generating one token requires roughly 2N floating-point operations. The factor of 2 is the standard back-of-envelope: one multiply and one add per parameter, per token, for a dense model. The arithmetic gets more nuanced for mixture-of-experts (MoE) models and for prompt processing (which is more parallelizable than generation), but the 2N rule is good enough for most purposes.

For some concrete numbers:

  • 70B-parameter dense model: ~140 GFLOPs per token (2 × 70 × 10⁹)
  • 400B-parameter dense model: ~800 GFLOPs per token
  • MoE model with 8 experts but 2 active per token: roughly the active-parameter compute, so a "1T total, 100B active" MoE is about 200 GFLOPs per token

Memory comes before the throughput claim

A dense 70-billion-parameter model needs roughly 140 GB just for its weights at two bytes per parameter. One 80 GB H100 cannot hold those weights, let alone the KV cache and runtime overhead. Quantization or distributing the model across GPUs changes the memory requirement and the performance problem.

There is no useful generic claim that "a 70B model on one H100 produces 5,000 tokens per second." A benchmark needs a named model, precision, GPU count, serving engine, prompt length, output length, concurrency, and latency target. Batching can increase aggregate throughput while making the individual user wait longer.

NVIDIA's H100 specifications are hardware limits. They are not measured application throughput. That distinction decides whether the rest of the economics means anything.

An energy calculation with visible assumptions

To demonstrate the arithmetic, choose a hypothetical serving setup drawing 700 W and delivering 5,000 output tokens per second in aggregate. This is not a measured 70B deployment.

QuantityCalculationResult
Output-token energy700 J/s ÷ 5,000 tok/s0.14 J/token
Energy in Wh0.14 ÷ 3,6000.0000389 Wh/token
500 output tokens500 × 0.00003890.0194 Wh
With assumed PUE0.0194 × 1.30.0253 Wh

That calculation counts the chosen power draw and an overhead multiplier. It does not establish the whole request's energy: input processing, reasoning tokens, other hardware, idle capacity, and the actual serving configuration matter. Training energy is a separate ledger.

Using the rounded 0.025 Wh per response assumption, a billion responses a day would consume:

0.025 Wh × 1,000,000,000 × 365 = 9.125 billion Wh = 9.125 GWh per year.

The earlier 12.5 GWh figure did not follow from these inputs. More fundamentally, the hypothetical per-response value cannot be generalized to every AI service. A comparison with search needs measurements using compatible workloads and accounting boundaries. These assumptions do not establish that an AI response is ten times cheaper, or ten times more expensive, than a search.

Price is not cost

Providers publish API prices. They generally do not publish enough of their fleet economics to calculate a model's gross margin from those prices. The relevant variables include hardware utilization, model architecture, caching, prompt and output length, reasoning, and service commitments.

Check current model-specific prices rather than a class-wide table: OpenAI, Anthropic, and Google. Those pages can change. Reasoning can add billable tokens under a provider's rules; it does not imply a universal multiplier relative to visible output.

Generating output sequentially and processing input in parallel helps explain why input and output have different economics. It does not require every provider to charge the same ratio, or even to structure every product's pricing the same way.

Keep the cost ledger from double-counting

A useful serving-cost model separates:

  • Hardware and facility allocation: owned-asset depreciation or rental costs, assigned to delivered work with utilization made explicit.
  • Operating costs: electricity, networking, staff, and other service expenses.
  • Training and development recovery: an allocation that depends on how much lifetime usage you assume.
  • Commercial price: what the customer pays under the provider's published terms.

The bare silicon is already part of the accelerator purchase. Adding a separate silicon allowance on top of complete hardware depreciation would double-count it. Training amortization is an allocated cost, not the marginal electricity cost of the next token.

A full per-token reconciliation needs actual inputs for those categories. Without them, a table of tiny dollar amounts creates more confidence than information.

The point of all of this

A token is a physical computation running on equipment with a supply chain. The silicon, the wafer, the fabrication, the cluster, and the serving system all matter. But an engineering explanation only helps if its arithmetic, assumptions, and measurement boundaries survive inspection.

That is the useful move from "AI is magic" to something you can reason about. Pick a workload. Name the equipment. Measure the output. Keep the units straight. Then decide what the result actually supports.