Skip to main content
Interactive 3D LLM Engine

3D LLM Visualization

Interactive visual LLM and 3D transformer visualization. Explore how GPT Transformer models process tokens, calculate self-attention, project query/key/value vectors, and compute next-token probability distributions step-by-step.

Direct Answer: What is this 3D LLM Visualization & Diagram?

The 3D LLM Visualization engine provides an interactive visual LLM graphic and 3D transformer visualization. Operating as a step-by-step 3D LLM diagram and architectural LLM graphic, it renders the internal layer-by-layer forward pass of a GPT Transformer. Step inside the neural network to inspect token embedding lookups, multi-head self-attention QKV vector matrix math (Softmax(QK^T / sqrt(d_k))V), GELU feedforward expansions, layer norm, and output logit probabilities.

Mathematical Deep Dive

3D Transformer Self-Attention Breakdown

A step-by-step linear algebra derivation of how Large Language Models transform discrete token IDs into contextualized vector representations via scaled dot-product attention and multi-head projections.

Tensor Dimension Cheat Sheet

Standard GPT Architecture
Sequence Length
n
Number of Tokens
Model Dimension
dmodel
e.g., 768 or 4096
Attention Heads
h
e.g., 12 or 32 heads
Head Dimension
dk = dv
dmodel / h
Attention Matrix
n × n
Token affinities
FFN Hidden Dim
dff ≈ 4dmodel
MLP Expansion
01

Token Embedding & Positional Encoding

Input Tensor: X ∈ ℝn × dmodel

When text enters a Transformer, each token string is mapped to an integer token ID from a vocabulary matrix V of size |V| (e.g., 50,257 for GPT-2 or 100,277 for GPT-4o). An embedding lookup matrix WE ∈ ℝ|V| × dmodel translates each discrete token ID into a continuous dense embedding vector E ∈ ℝn × dmodel.

Because self-attention operations are inherently permutation-invariant (order-agnostic), Transformers inject sequential position information via a positional encoding tensor P ∈ ℝn × dmodel. In traditional architectures, this is computed via sinusoidal harmonic frequencies or learned absolute tables:

  • Sinusoidal Evens: PE(pos, 2i) = sin(pos / 100002i / dmodel)
  • Sinusoidal Odds: PE(pos, 2i+1) = cos(pos / 100002i / dmodel)
  • Modern LLMs: RoPE (Rotary Position Embeddings) apply a 2D rotation matrix directly to Q and K vectors.
Formula Step 1
X = E + P
E (Embeddings): [n × dmodel]
P (Positional): [n × dmodel]
X (Input Tensor): [n × dmodel]
02

Query, Key, and Value Projections (QKV)

Projections: Q, K ∈ ℝn × dk, V ∈ ℝn × dv

At each attention layer, the input tensor X is linearly projected into three distinct semantic vector spaces using learned weight matrices:

Queries (Q)

Represents what the current token is seeking or asking about surrounding context.

Keys (K)

Represents what characteristics or semantic content this token offers to others.

Values (V)

The actual feature information extracted and passed forward if matched.

The projections are computed via matrix multiplication: Q = X · WQ, K = X · WK, and V = X · WV, where WQ, WK ∈ ℝdmodel × dk and WV ∈ ℝdmodel × dv.

Projection Equations Step 2
Q = X · WQ  → [n × dk]
K = X · WK  → [n × dk]
V = X · WV  → [n × dv]
In standard multi-head setups: dk = dv = dmodel / h
03

Scaled Dot-Product Attention Mechanism

Core Equation: Attention(Q, K, V)

The mathematical heart of the Transformer is the Scaled Dot-Product Attention operator. It determines how much attention token i pays to token j across the context window:

1. Raw Affinity Scores (Q · KT):

Multiplying the Query matrix by the transposed Key matrix yields an n × n matrix where each cell Sij = qi · kj measures directional alignment in embedding space.

2. Variance Normalization (√dk):

Under the assumption that components of Q and K are independent random variables with mean 0 and variance 1, their dot product has variance dk. For large dk (e.g. 64 or 128), dot products grow extremely large in magnitude, pushing the Softmax function into regions with near-zero gradients. Dividing by √dk stabilizes variance back to 1.0.

3. Softmax Probability & Value Weighting:

Applying row-wise Softmax transforms raw logits into a strictly normalized probability distribution summing to 1.0 per row. Finally, multiplying by Value matrix V produces the attention-weighted context vector representation.

Canonical Equation Attention Engine
Attention(Q, K, V) = softmax(Q · KT / √dk) · V
Q · KT Shape: [n × dk] × [dk × n] = [n × n]
Softmax(A) Shape: [n × n] (Normalized weights)
Final Output Tensor: [n × n] × [n × dv] = [n × dv]
04

Multi-Head Attention & Output Projection

h Heads → WO ∈ ℝ(h · dv) × dmodel

Rather than calculating a single attention distribution, Multi-Head Attention linearly projects the queries, keys, and values h times with different learned parameter matrices.

This enables the neural network to jointly attend to information from different representation subspaces at different positions. For instance, one head might specialize in tracking syntactic verb-object agreements, while another head attends to long-range coreference resolution (pronoun antecedent matching).

Each head produces an output tensor of shape n × dv. The h heads are concatenated along the channel dimension into a unified tensor of size n × (h · dv) = n × dmodel, then multiplied by output weight matrix WO:

Multi-Head Formula Step 4
headi = Attention(Q · WQ,i, K · WK,i, V · WV,i)
MultiHead(Q, K, V) = Concat(head1, ..., headh) · WO
Each Head Tensor: [n × (dmodel / h)]
Concatenated: [n × dmodel]
After Output Projection WO: [n × dmodel]
05

Feed-Forward Network (FFN) & Residual LayerNorm

Residual: LayerNorm(x + Sublayer(x))

Following multi-head self-attention, each token vector passes identically and independently through a Position-wise Feed-Forward Network (FFN). The FFN consists of two linear transformations with a non-linear activation function (such as GELU or SwiGLU in contemporary LLMs):

FFN(x) = GELU(x · W1 + b1) · W2 + b2

Here, the hidden inner dimension expands to dff ≈ 4 × dmodel before projecting back to dmodel. This expansion provides the parameter capacity where factual knowledge and associative memories are stored.

Every sublayer is wrapped in a residual skip connection followed by Layer Normalization (Pre-LN in modern decoders or Post-LN in original Transformers), which prevents vanishing or exploding gradients across deep stacks:

xnorm = LayerNorm(x) = γ · ((x - μ) / √(σ2 + ε)) + β
Layer Pipeline Step 5
1. Self-Attention Output:
x1 = LayerNorm(x + MultiHead(x))
2. MLP Expansion & Contraction:
x2 = LayerNorm(x1 + FFN(x1))
The resulting tensor x2 ∈ ℝn × dmodel feeds into subsequent transformer layers or output vocabulary unembedding (Softmax over logits).
Key Architectural Takeaway

Self-attention performs pairwise token communication across the sequence length n, while the Feed-Forward Network operates independently per position. Together, repeated across dozens of stacked transformer layers, these mathematical operations give Large Language Models the ability to grasp syntax, reason over multi-step instructions, and generate coherent text.