AI Diagram Tool

Transformer Architecture Diagram Generator encoder-decoder, GPT, and ViT

Generate Transformer architecture diagrams from a text prompt: encoder-decoder, GPT-style decoder-only, and ViT. Includes exact block diagrams.

Encoder-decoder, decoder-only, and ViT diagramsAttention, Add & Norm, and N× stacking labeledExact reference diagrams with attention masksKnown flaws noted on every AI example

Create Your Transformer Diagram

Describe your Transformer architecture
741 / 20,000 characters
Start from an example
AI-generated diagram of the original encoder-decoder Transformer with Add and Norm, cross-attention, and Linear plus Softmax outputExample: Encoder-decoderView full size

Preview is free on this page ·

Preview

Your Transformer diagram will appear here

AI diagrams can misdraw arrows or labels. Check every block against your model before use.

Transformer Architecture Diagram Examples

Two exact reference diagrams drawn in code, plus two AI diagrams with their known flaws noted. AI diagrams need human review.

View:

Original encoder-decoder Transformer (exact diagram)

Drawn in code, not by AI, with the post-LN layout of the 2017 paper: residual skips around every sublayer, the encoder output feeding the decoder cross-attention, and Linear plus Softmax on top. Schematic: one head, dimensions, and dropout are not drawn.

encoder-decoderAdd & Normexact

Encoder-only vs decoder-only vs encoder-decoder (exact diagram)

Drawn in code: the three families side by side with their attention masks (full for encoders, causal for decoders, plus the cross-attention grid). Example models are BERT-style, GPT-style, and the original Transformer, T5, and BART. To keep the masks readable it omits the residual connections and Add & Norm blocks, which the first diagram shows. This is also where to see the correct decoder-only structure: causal self-attention and no cross-attention.

encoder-onlydecoder-onlyattention mask

Encoder-decoder (AI diagram)

AI-generated from the prompt shown. All labels came out spelled correctly and every required block is present. Known flaws: the arrow between the two stacks has an arrowhead at both ends and lands on the encoder top Add & Norm, so the direction of the encoder-to-decoder link is ambiguous (it should only point into the cross-attention); the residual loops are routed loosely.

encoder-decodercross-attentionAI

Vision Transformer, ViT (AI diagram)

AI-generated from the prompt shown. The flow is right: patches, linear projection, [CLS] token, position embeddings, Layer Norm before attention and MLP, MLP head. Known flaw: the labeled 3 x 3 grid is drawn as 8 visible patch tiles in a 2 x 4 layout in the next step, so the patch count does not match the "9 patches" label.

ViTpatch embeddingAI

What a Transformer architecture diagram shows

A Transformer architecture diagram shows how a sequence of tokens becomes a prediction: embedding and position information go in, a stack of identical blocks mixes the tokens with attention and transforms each one with a feed-forward network, and a final layer turns the result into output probabilities. Figures in papers, lectures, and documentation fall into three types: the full encoder-decoder block diagram, the single-stack diagrams for encoder-only and decoder-only models, and the Vision Transformer (ViT), which swaps tokens for image patches.

The parts every Transformer diagram must get right

  • Token embedding plus positional encoding: attention has no sense of order by itself, so position information is added to the embeddings before the first block (sinusoidal in the original paper; learned in many later models).
  • Multi-head attention: several scaled dot-product attention heads run in parallel and are concatenated. The original base model uses 8 heads with a model width of 512.
  • Add & Norm: a residual connection that adds the sublayer input to its output, followed by layer normalization. Every attention and feed-forward sublayer has one.
  • Feed forward: the position-wise two-layer network applied to each token (2048 hidden units in the original base model).
  • Masked self-attention: in a decoder, each position may only attend to earlier positions, which is what makes next-token generation possible.
  • Cross-attention: only in encoder-decoder models. Queries come from the decoder, keys and values from the encoder output.
  • N× stacking: the block is repeated N times (6 in the original base model). Draw it once with a bracket rather than copying it.
  • Output Linear + Softmax: the final projection to vocabulary size and the probability distribution over the next token.

Encoder-only, decoder-only, and encoder-decoder

  • Encoder-only (BERT-style): self-attention with no mask, so every token sees every other token. Used for classification, tagging, and embeddings.
  • Decoder-only (GPT-style): causal self-attention, so each token sees only earlier tokens. Used for text generation and chat models; no cross-attention.
  • Encoder-decoder (the original Transformer, T5, BART): an encoder reads the input, a decoder writes the output and consults the encoder through cross-attention. Used for translation and summarization.
  • Vision Transformer (ViT): an encoder-only design in which the input image is cut into patches, each patch is linearly projected to an embedding, a learnable [CLS] token is prepended, and position embeddings are added before the encoder.

Post-LN or pre-LN: a detail diagrams disagree on

The original 2017 diagram applies layer normalization after the residual addition (post-LN), which is why its boxes are labeled "Add & Norm". Many later models, including GPT-2 style decoders and ViT, normalize before each sublayer (pre-LN) and add a final LayerNorm before the output. Both are legitimate; the mistake is mixing them in one figure without saying so. Pick the one your model uses and state it in the caption.

Exact diagrams vs AI illustration on this page

  • Exact diagrams: the two reference diagrams in the gallery are drawn in code, so the sublayer order, residual loops, cross-attention link, and mask patterns are fixed. Use them when a figure has to show the real structure.
  • AI illustration: the generator is best for quick concept diagrams and for variants, such as a ViT, a decoder-only model, or your own block layout. AI can still misdraw arrows, drop a residual connection, or mislabel a block.
  • Always check an AI diagram against the paper or the model code before it goes into a manuscript, thesis, or slide deck, and cite the original architecture paper.

How to prompt for a correct Transformer diagram

  • Name the family: encoder-decoder, decoder-only, encoder-only, or ViT. Each draws a different set of blocks.
  • List the blocks from bottom to top in the order you want them, with the exact label text in quotes.
  • State the connections you need: residual skip arrows around each sublayer, the encoder output into cross-attention, the N× bracket.
  • Say what to leave out, for example "no encoder, no cross-attention" for a GPT-style figure.
  • Compare every arrow with the exact diagrams above before using the result, and regenerate if a label or an arrow direction is wrong.

Frequently Asked Questions

Need an MLP, CNN, or recurrent network instead? Use the Neural Network Diagram Generator. This page covers the Transformer family only.

Sources for the architectures drawn here: Vaswani et al., Attention Is All You Need (2017), and Dosovitskiy et al., An Image is Worth 16x16 Words (2020).