# Kafka (Apache Kafka)

URL: https://softwaredictionary.org/terms/kafka
Category: Backend & APIs
Last updated: 2026-10-03
Pronunciation: KAHF-kuh

In short: Apache Kafka is a distributed event streaming platform that stores events in durable, ordered logs for many services to publish, read in real time or replay.

## What is Kafka?

Kafka was built at LinkedIn to move huge volumes of activity data and was open-sourced in 2011; it is now an Apache Software Foundation project. Producers write events, such as "order placed" or "page viewed", to named topics. Kafka appends each event to the end of a log and keeps it for a configured time, days or even forever, whether or not anyone has read it yet.

Each topic is split into partitions spread across a cluster of servers called brokers. Events within one partition keep their order, and each has a numbered position called an offset. Consumers read at their own pace and remember their offset, so a slow or restarted consumer simply continues where it left off, and a new one can replay history from the beginning.

Consumers that share a group name split the partitions between them, so adding consumers spreads the work, while separate groups each receive every event. This lets one stream of orders feed billing, shipping and analytics independently. Kafka is used for event-driven microservices, activity tracking, log and metric pipelines, and change data capture from databases.

A common misconception is that Kafka is just a message queue. A queue usually deletes a message once it is handled, while Kafka keeps the log and lets any number of readers go back in time. That power comes with operational weight: partitions, replication and retention need planning, so smaller systems often start with a simpler broker such as RabbitMQ.

## Key takeaways

- Kafka stores events in durable, ordered, append-only logs called topics.
- Topics are split into partitions spread across a cluster of brokers.
- Consumers track their own offset, so they can resume or replay history.
- Consumer groups share partitions; separate groups each get every event.
- It is more powerful, and heavier to run, than a simple message queue.

## Example: Producing and consuming events (Node.js with kafkajs)

```javascript
import { Kafka } from "kafkajs";

const kafka = new Kafka({ brokers: ["localhost:9092"] });

// Producer: append an event to the "orders" topic
const producer = kafka.producer();
await producer.connect();
await producer.send({ topic: "orders", messages: [{ key: "1001", value: '{"total": 49}' }] });

// Consumer: members of "billing" share the topic's partitions
const consumer = kafka.consumer({ groupId: "billing" });
await consumer.connect();
await consumer.subscribe({ topic: "orders", fromBeginning: true });
await consumer.run({ eachMessage: async ({ message }) => console.log(message.value.toString()) });
```

## Frequently asked questions

**Is Kafka a message queue?**

Not exactly. It can be used like one, but Kafka keeps events in a log after they are read, so many consumers can read the same events independently and replay them later. A classic queue removes a message once it is handled.

**What is a Kafka topic?**

A topic is a named stream of events, such as "orders". It is split into partitions, which are ordered logs stored across the brokers of the cluster.

**Does Kafka still need ZooKeeper?**

No. Newer versions manage the cluster themselves with a built-in mode called KRaft, and Kafka 4.0 removed ZooKeeper support entirely.

## Sources

- [Apache Kafka documentation](https://kafka.apache.org/documentation/)

---

Software Dictionary: https://softwaredictionary.org/ · https://softwaredictionary.org/llms.txt
