NoteMachine LearningRoboticsEvergreenupdated 2026.0811 min read

Which sequence-modelling paradigms exist, what exactly separates BERT from GPT, and why robot VLAs (such as π₀-FAST) use a hybrid of the two called Prefix-LM. This page unfolds the "autoregressive + cross-entropy" and "bidirectional prefix / causal action mask" remarks made in Action Tokenization.

§1Preliminaries

The main text uses the following concepts over and over. Each is given as "English term = 中文名 = minimal definition"; the Chinese name is kept alongside so that English readers can trace the term back into the Chinese literature:

  • self-attention = 自注意力 = the mechanism by which every token in a sequence uses its own query to match the keys of the other tokens and aggregates their values by the resulting weights, thereby extracting information from context[1].
  • attention mask = 注意力掩码 = the boolean matrix specifying "which key positions each query position may attend to"; the difference between the three paradigms of this page (bidirectional / causal / Prefix-LM) lives entirely in this matrix[1].
  • cross-attention = 交叉注意力 = the form of attention in which decoder-side queries attend to the keys/values produced by the encoder, injecting the input condition into the generating side[1].
  • autoregressive = 自回归 = the modelling style that factorises the joint probability into per-token conditionals by the chain rule and predicts the next token from the already-generated history[2].
  • masked language modeling (MLM) = 掩码语言建模 = the training objective that hides part of the input tokens and reconstructs them from the unmasked left and right context[3].
  • cross-entropy = 交叉熵 = the loss function measuring the discrepancy between the predicted and the target distribution; next-token training is exactly the negative log-likelihood of the correct token[2].

§2What is a sequence model actually estimating

Given a string of tokens x1,x2,…,xnx_1,x_2,\dots,x_n, a language model is essentially estimating their joint probability p(x1,…,xn)p(x_1,\dots,x_n). How you factorise that joint probability determines the paradigm:

ParadigmFactorisation / training objectiveCan it generate?
Autoregressive (AR)p(x)=∏tp(xt∣x<t)p(x)=\prod_t p(x_t\mid x_{<t}), predict the next token✅ Built for generation
Autoencoding / masked (AE/MLM)Hide a portion x~\tilde x, reconstruct p(xmasked∣x~)p(x_{\text{masked}}\mid \tilde x)❌ Mainly for understanding/representation
Sequence-to-sequence (seq2seq)Encode the input, decode the output p(y∣x)p(y\mid x)✅ Translation/summarisation

§3The three architectural paradigms (one set of bricks, three ways to stack them)

The Transformer brick is "multi-head self-attention + feed-forward layer"[1]. The three paradigms differ only in how you stack them and which mask you use:

Encoder-only (BERT)Bidirectional attn ×NFeed-forwardAll inputs visible↑ Per-position representationsDecoder-only (GPT)Causal attn ×NFeed-forwardLeft context only↑ Predicts the next tokenEncoder-Decoder (T5)Encoder (bidir.)Decoder (causal)Linked by cross-attention
Figure 1 · The three architectural paradigms side by side. The same Transformer block; the difference is only "which direction the mask runs" and "how the blocks are stacked". Just read off the type of the attention layer in each box: green = bidirectional attention (encoder side), blue = causal attention (decoder side), purple arrow = the cross-attention that links the two. Takeaway: the paradigms differ not in the bricks but in the mask and the stacking.

§4The core mechanism: the attention mask (flip it yourself and see)

Self-attention lets every token "query" the other tokens. The mask decides which queries are permitted. This is the only essential difference between BERT, GPT and Prefix-LM. In the figure below: row = the token being computed (query), column = the token being attended to (key), a lit cell = attention allowed.

Figure 2 · Demo: the three attention masks compared
Click the buttons to switch. The sequence is assumed to be [CLS] task prompt ; A1 A2 A3 |, the first 4 (including the separator ;) being the prompt prefix and the last 4 the actions to be generated. What to look at: row = query (who is looking), column = key (being attended to); token labels in green = prompt prefix, in blue = action suffix; lit cell = attention allowed (in Prefix-LM mode the bidirectional region inside the prefix is drawn in green), grey cell = forbidden by the mask. Takeaway: causal = lower triangle, bidirectional = fully lit, Prefix-LM = a bidirectional block at the upper left + a causal triangle at the lower right.
  • Causal mask (lower triangle): token tt may attend only to positions ≤t\le t. This guarantees "no peeking at the future while predicting the next token", and is the precondition for autoregressive generation. GPT uses it throughout.
  • Bidirectional mask (fully lit): every token sees the whole sentence. Understanding tasks (classification, extraction) need global context, but you cannot generate token by token directly (you would peek at the answer). BERT uses it.

§5BERT — encoder / bidirectional / masked language model

Bidirectional Encoder Representations from Transformers[3]. It uses only the encoder stack of the Transformer, with bidirectional attention throughout.

Training objective: MLM (Masked Language Modeling, a fill-in-the-blanks exercise)

text
Input:   机器人 [MASK] 可爱      ← randomly mask out 15% of the tokens
Target:  predict [MASK] = "很"    ← guess it from the context on both sides

Because you have to "guess the middle from both sides", the attention must be bidirectional. Another classic objective is NSP (deciding whether two sentences are adjacent)[3], which later work (RoBERTa) found could be dropped[4].

PropertyBERT
StructureEncoder-only
AttentionBidirectional
Training objectiveMLM (+NSP)
StrengthUnderstanding / representation (classification, extraction, retrieval)
Can it generate❌
Notable descendantsRoBERTa, ALBERT, DeBERTa, ELECTRA

§6GPT — decoder / causal / autoregressive

Generative Pre-trained Transformer[5]. It uses only the decoder stack of the Transformer (with the cross-attention to an encoder removed), under a causal mask throughout.

Training objective: next-token prediction

L=−∑tlog⁡pθ(xt∣x<t)\mathcal{L}=-\sum_t \log p_\theta(x_t\mid x_{<t})

This is precisely the objective by which π₀-FAST treats action tokens as language and applies next-token cross-entropy. The demo below shows how autoregression "spits the sequence out one token at a time":

Figure 3 · Demo: autoregressive generation (token-by-token sampling)
Click "Generate next token": the model produces a probability distribution over the next token from the history generated so far, samples from it, and feeds the sample back into the input — this is the loop of GPT and of token-by-token action sampling. What to look at: the boxes along the top = the generated history (each new token is appended on the right), the purple bar chart below = the probability distribution over the next token (each bar is labelled with its probability); drag the temperature slider and you can see the distribution flatten or sharpen, making the sampling correspondingly more random or more deterministic. Takeaway: generation = the loop "predict a distribution → sample → feed back into the input".
PropertyGPT
StructureDecoder-only
AttentionCausal (unidirectional)
Training objectivenext-token prediction
StrengthGeneration (dialogue, writing, code)
Can it generate✅ (by construction)
Notable familyGPT-2/3/4[6][7], LLaMA, Mistral, Qwen, Gemma, Claude

§7T5 / BART — encoder-decoder (seq2seq)

Put the two branches together: the encoder reads the input bidirectionally (say, an English sentence), the decoder generates the output causally (say, a Chinese sentence), and cross-attention in between lets the decoder read the encoder's representations. This is naturally suited to "input → output" conversion tasks (translation, summarisation).

§8Prefix-LM — bidirectional prefix + causal suffix (exactly what π₀-FAST uses)

This is the hinge connecting this page to action tokenization. Prefix-LM (the prefix language model) is a variant of decoder-only[8][10]: split the sequence into a prefix and a suffix, and use a single hybrid mask:

  • Prefix (prompt + images + discretised state): attends bidirectionally within itself (like an encoder, so the condition is fully understood).
  • Suffix (the action tokens to be generated): attends causally (like a decoder, preserving autoregression). The suffix can see the entire prefix.

The "Prefix-LM" button in the demo above (Figure 2) shows you the shape of this mask directly: a fully lit square at the upper left (bidirectional prefix) and a lower triangle at the lower right (causal suffix).

§9The full model-family table (quick reference)

ModelParadigmAttentionTraining objectiveTypical use
BERT / RoBERTaEncoder-onlyBidirectionalMLMUnderstanding, classification, retrieval
GPT / LLaMA / Gemma / ClaudeDecoder-onlyCausalnext-tokenGeneration, dialogue
T5 / BARTEncoder-DecoderBidirectional encoding + causal decodingDenoising / span reconstructionTranslation, summarisation
PaLM / UL2 / PaliGemmaPrefix-LM / hybridBidirectional prefix + causal suffix(prefix) next-tokenMultimodal, conditional generation
π₀-FASTPrefix-LM (VLA)Bidirectional prefix + causal actionsCross-entropy on action tokensRobot action generation

§10Back to action tokenization

Threading it together (see Action Tokenization for details):

  1. Actions are turned into discrete tokens by a tokenizer (FAST/FSQ/Binning).
  2. π₀-FAST uses Prefix-LM: the prefix (observations + instruction + state) is bidirectional, the suffix (actions) is causal[12].
  3. Training uses GPT-style next-token cross-entropy, with the loss computed only over the action segment[12].
  4. Inference uses autoregressive sampling (the demo on this page, Figure 3), relying on temperature to draw out the different modes of a multimodal solution set.

You should now be able to answer: Why not a pure BERT for a VLA? (It cannot generate.) Why not a purely causal GPT instead of Prefix-LM? (So that the observation condition is understood fully and bidirectionally.) Why does training look like GPT? (next-token cross-entropy + multimodality for free.)


§11Further reading

  • Action Tokenization — the main thread on discretising robot actions (FAST / FSQ / Binning + autoregression + multimodality)
  • Fourier Transform and the DCT — frequency-domain transforms and energy compaction, the mathematical core of FAST's first step

§12References

  1. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., Polosukhin, I. Attention Is All You Need. NeurIPS 2017. arXiv:1706.03762.
  2. Goodfellow, I., Bengio, Y., Courville, A. Deep Learning. MIT Press, 2016.
  3. Devlin, J., Chang, M.-W., Lee, K., Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL-HLT 2019. arXiv:1810.04805.
  4. Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692.
  5. Radford, A., Narasimhan, K., Salimans, T., Sutskever, I. Improving Language Understanding by Generative Pre-Training. OpenAI technical report, 2018.
  6. Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I. Language Models are Unsupervised Multitask Learners. OpenAI technical report, 2019.
  7. Brown, T. B., Mann, B., Ryder, N., et al. Language Models are Few-Shot Learners. NeurIPS 2020. arXiv:2005.14165.
  8. Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. JMLR 21, 2020. arXiv:1910.10683.
  9. Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., Zettlemoyer, L. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. ACL 2020. arXiv:1910.13461.
  10. Tay, Y., Dehghani, M., Tran, V. Q., et al. UL2: Unifying Language Learning Paradigms. arXiv:2205.05131.
  11. Beyer, L., Steiner, A., et al. PaliGemma: A versatile 3B VLM for transfer. arXiv:2407.07726.
  12. Pertsch, K., Stachowicz, K., Ichter, B., Driess, D., Nair, S., Vuong, Q., Mees, O., Finn, C., Levine, S. FAST: Efficient Action Tokenization for Vision-Language-Action Models. arXiv:2501.09747.
Titles, sections and body text, in this language.
    ↑↓ · Enter · Escastro-inkstone