Transformer architecture

Transformer architecture is a neural network design that processes sequences of data, like text, by weighing the importance of different parts of the input.

What it is

Transformer architecture is a type of neural network introduced in 2017, specifically designed to handle sequential data, such as natural language or time series. Unlike previous architectures, it processes entire sequences in parallel, using a mechanism called "self-attention" to weigh the relevance of different input elements to each other. This parallel processing and attention mechanism significantly improved efficiency and performance for tasks like translation and text generation, making it a cornerstone for large language models.

The transformer architecture is central to the development of large language models (LLMs) and generative AI, frequently mentioned in news about new frontier models. Its efficiency in processing vast amounts of data during model training and inference has driven the demand for specialized compute resources like GPUs. Investors track companies leveraging or developing transformer-based models, as they are key to innovation in AI applications from chatbots to advanced reasoning models.

Why it matters

This architecture underpins most state-of-the-art AI, influencing the capabilities and market value of leading AI companies and their products.

Reviewed under editorial standardsUpdated September 26, 2026Not investment advice