Skip to main content

TPU

Tensor Processing Unit

Pronunciation
tee-pee-YOO
Updated 3 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/tpu

In short

A TPU (Tensor Processing Unit) is Google's custom chip for the matrix math of neural networks, used to train and run AI models, mainly on Google Cloud.

What is a TPU?

A TPU is an application-specific integrated circuit (ASIC), a chip built for one kind of job, that Google designed to speed up neural networks. A tensor is simply a multi-dimensional array of numbers, and training or running a neural network is mostly multiplying large tensors together. Google began using TPUs in its data centers in 2015 and announced them in 2016; the first generation could only run trained models, and training arrived with the second generation in 2017.

The heart of a TPU is its matrix multiply unit, a grid of thousands of small multiply-and-add circuits arranged as a systolic array: numbers pulse from one cell to the next, so each value read from memory is reused many times instead of being fetched again. TPUs mostly compute in low-precision formats such as bfloat16, which is accurate enough for neural networks and much cheaper in silicon. Chips are joined by a fast, dedicated network into pods; a pod of the seventh generation, Ironwood, links 9,216 chips.

Google has offered TPUs mainly as a cloud service rather than as cards for your own servers: developers rent them on Google Cloud, and Google uses them to train and serve its own models, including Gemini. Code reaches a TPU through the XLA compiler, which JAX, TensorFlow and PyTorch use to turn a model's operations into TPU instructions. JAX, which Google developed with TPUs in mind, is the most common choice on them.

TPU vs GPU is the comparison most people mean. A GPU is a general parallel processor that started with graphics, runs almost any parallel code and is sold by NVIDIA, AMD and others, with support in nearly every AI library. A TPU gives up that flexibility to do tensor math more efficiently, like a bakery line built for one kind of loaf next to a kitchen that can cook anything, and it is mainly available from one company's cloud. Models with many custom operations usually run more easily on GPUs, while large standard models, especially ones written in JAX, fit TPUs well.

Key takeaways

  • A TPU is a chip Google designed specifically for the matrix math in neural networks.
  • Its systolic array reuses each value many times, saving slow memory reads.
  • TPUs are rented through Google Cloud and power Google's own models.
  • JAX, TensorFlow and PyTorch reach TPUs through the XLA compiler.
  • GPUs are more flexible and widely available; TPUs trade flexibility for efficiency.

Example

Running a computation on a TPU with JAXpython
import jax
import jax.numpy as jnp

print(jax.devices())  # on a Cloud TPU VM: a list of TpuDevice objects

@jax.jit  # XLA compiles this function for the TPU's matrix units
def predict(weights, inputs):
    return jnp.tanh(inputs @ weights)

key = jax.random.key(0)
weights = jax.random.normal(key, (4096, 4096), dtype=jnp.bfloat16)
inputs = jnp.ones((1024, 4096), dtype=jnp.bfloat16)
print(predict(weights, inputs).shape)  # (1024, 4096)

Readers ask

What is the difference between a TPU and a GPU?

A GPU is a general-purpose parallel processor that handles graphics and almost any parallel code, sold by several vendors. A TPU is Google's chip built specifically for neural network math: less flexible, but very efficient for large models, and available mainly through Google Cloud.

Can I buy a TPU?

Generally not: Google offers its data center TPUs as a cloud service that you pay for as you use it. Notebook services such as Google Colab and Kaggle give limited free TPU time for experiments.

What is an NPU?

An NPU, or neural processing unit, is an AI accelerator built into phones and laptops to run small models efficiently on the device. It follows the same idea as a TPU, dedicated hardware for neural network math, at a much smaller scale.

See also

Sources

Spotted a mistake or something missing on this page?Suggest an edit

More

Settings