Skip to main content

Side by side

GPUvsTPU

What is the difference between a GPU and a TPU?

Updated 3 min read8 differences

In short

A GPU is a general parallel processor for graphics and AI, sold by several vendors; a TPU is Google's chip for neural network math, rented on Google Cloud.

GPU

Graphics Processing Unit

A GPU (graphics processing unit) is a processor whose thousands of small cores run the same calculation on lots of data at once, for graphics and AI.

Read the page on GPU

TPU

Tensor Processing Unit

A TPU (Tensor Processing Unit) is Google's custom chip for the matrix math of neural networks, used to train and run AI models, mainly on Google Cloud.

Read the page on TPU

GPU and TPU compared

AspectGPUTPU
What it isA general-purpose parallel processorAn application-specific chip (ASIC) for neural network math
Designed forGraphics first, now also AI and scientific computingTraining and running neural networks
Core designThousands of small cores working in parallelMatrix units built as systolic arrays
FlexibilityRuns almost any parallel code, including custom operationsBest at standard tensor operations
Made byNVIDIA, AMD and othersGoogle
How to get oneBuy cards, or rent them from any major cloudRent them, mainly through Google Cloud
SoftwareCUDA and nearly every AI frameworkJAX, TensorFlow and PyTorch through the XLA compiler
Best forGraphics, research, custom models and most AI workLarge, standard models at scale, especially in JAX

The difference, explained

A GPU (graphics processing unit) is a processor with thousands of small cores that run the same calculation on lots of data at once. Built for graphics, it became the main hardware for AI after NVIDIA's CUDA platform, released in 2007, made it programmable for general computing. A TPU (Tensor Processing Unit) is an application-specific chip that Google designed only for the tensor math of neural networks; Google began using TPUs in its data centers in 2015 and announced them in 2016.

The difference is flexibility versus specialization. A GPU runs almost any parallel code, from games and video to simulations and custom AI operations, and nearly every AI library supports it. A TPU gives up that flexibility: its matrix unit is a systolic array that passes numbers from cell to cell, so each value read from memory is reused many times, and it mostly computes in low-precision formats such as bfloat16, which makes it very efficient for large, standard models.

Availability differs as much as design. GPUs are sold by NVIDIA, AMD and others as cards for desktops, laptops and servers, and every major cloud rents them. TPUs are offered mainly as a Google Cloud service, linked by a dedicated network into pods of thousands of chips, and programs reach them through the XLA compiler from JAX, TensorFlow or PyTorch, with JAX the most common choice.

A common misconception is that a TPU is simply a faster GPU. It is faster or cheaper for some workloads, such as large transformer models written in JAX, while models with many custom operations usually run more easily on GPUs, and a TPU can't render graphics at all. Another is that AI always needs one or the other: CPUs still run small models well, and phones and laptops include NPUs, small accelerators that follow the same idea as a TPU.

Which one should you use?

Choose GPU when…

  • You need graphics, video or general parallel computing as well as AI.
  • Your model uses custom operations or libraries that expect CUDA.
  • You want to own the hardware, or run on any cloud or your own servers.
  • You are experimenting, and broad tool support matters most.

Choose TPU when…

  • You train or serve large, standard models on Google Cloud.
  • Your code is written in JAX, or already compiles well with XLA.
  • Cost per training run at large scale matters more than flexibility.
  • You need to scale across many chips joined by a fast dedicated network.

Readers ask

Is a TPU faster than a GPU?

It depends on the workload. TPUs can be faster or more cost-efficient for large, standard neural networks, especially in JAX, while GPUs are usually easier for custom models and handle many workloads that TPUs can't run at all.

Can I use PyTorch on a TPU?

Yes, through PyTorch/XLA, which compiles PyTorch operations with the XLA compiler for the TPU. JAX remains the most common choice on TPUs, and some PyTorch code needs changes to run efficiently.

Can I buy a TPU?

Generally not: Google offers its data center TPUs as a cloud service you pay for as you use it. GPUs, by contrast, can be bought as cards or rented from almost any cloud.

More

Settings