Protocol Buffers
- Pronunciation
- PROH-toh-buf
In short
Protocol Buffers (protobuf) is Google's language-neutral format for defining structured data in a schema and serializing it into small, fast binary messages.
What are Protocol Buffers?
Protocol Buffers, usually shortened to protobuf, is a data format created at Google for its own services and released as open source in 2008. You describe your data once in a .proto file, where each message lists typed fields such as string name = 1;. The protoc compiler then generates classes for languages such as Java, Python, Go, C++, C#, Kotlin and Dart, with methods that turn objects into bytes and back.
On the wire, a protobuf message is binary, not text. Each field is stored as its field number and a compact encoding of its value, so names like customer_email are never sent, and a small integer value fits in a single byte. Messages are therefore much smaller than the same data in JSON and faster to read and write, but you need the schema to make sense of them. The format is designed to evolve: you can add fields with new numbers, and old programs simply skip fields they don't know, as long as numbers are never reused.
Protobuf is the default message format of gRPC, and it is also used to store data, to send events through Kafka and other message queues, and for configuration inside large systems. Think of it as a form with numbered boxes that both sides printed in advance: only the box numbers and their contents travel, not the labels. Since 2024, the syntax = "proto3" line has been giving way to editions, such as edition = "2023", which control the same behaviors through individual feature settings.
Protobuf is often compared with JSON and confused with gRPC. JSON is text that any program can read without a schema, which makes it easy to debug and the usual choice for public web APIs; protobuf trades that readability for smaller messages, faster parsing and strict types. gRPC, in turn, is a framework for calling remote methods, and protobuf is only the format it uses to describe services and encode messages, so protobuf can be used without gRPC.
Key takeaways
- Protobuf defines data in
.protoschema files and serializes it to compact binary. - The
protoccompiler generates code for many languages from one schema. - Fields are identified by numbers, so new fields can be added without breaking old readers.
- Messages are smaller and faster than JSON but not human-readable.
- gRPC uses protobuf by default, but protobuf works on its own too.
Example
syntax = "proto3";
package shop.v1;
message Order {
int64 id = 1; // field numbers, not names, go on the wire
string customer_email = 2;
repeated LineItem items = 3; // a list of LineItem messages
reserved 4; // a removed field: never reuse its number
optional string coupon = 5; // added later; older readers skip it
}
message LineItem { string sku = 1; int32 quantity = 2; }
// Generate classes: protoc --python_out=gen --java_out=gen order.protoReaders ask
Is protobuf better than JSON?
It depends on the use. Protobuf messages are smaller, faster to process and come with a strict schema, which suits internal traffic between services and high volumes. JSON is human-readable and needs no generated code, so it remains the usual choice for public and browser-facing APIs.
What is the difference between Protocol Buffers and gRPC?
Protocol Buffers is a format for describing and serializing data. gRPC is a remote procedure call framework that uses protobuf by default to define its services and to encode every request and response.
What is the difference between proto2 and proto3?
They are two versions of the .proto language. proto3, released in 2016, simplified proto2 by dropping required fields and custom default values, and both are now being succeeded by editions, which express the same choices as feature settings.
See also
- gRPCBackend & APIs, p. 23gRPC is an open-source framework for calling functions on a remote server as if they were local, using Protocol Buffers and HTTP/2 for fast, typed messages.
- JSONBackend & APIs, p. 28JSON is a lightweight, text-based format for storing and exchanging structured data as key-value pairs and lists, readable by both humans and machines.
- SerializationBackend & APIs, p. 45Serialization is the process of converting in-memory data structures into a format such as JSON or bytes, so they can be stored or sent over a network.
- XMLBackend & APIs, p. 52XML (Extensible Markup Language) is a text format for structured data that uses nested tags you define yourself, readable by both people and machines.
- Database SchemaDatabases, p. 12A database schema is the blueprint of a database that defines its tables, columns, data types, relationships, and the rules that stored data must follow.
- API VersioningBackend & APIs, p. 4API versioning is the practice of labeling and managing changes to an API so existing clients keep working while new versions add or change features.
Sources
Spotted a mistake or something missing on this page?Suggest an edit