10 chapters are open
First lesson: Batch report or stream: how long an answer waits
Interactive course · Apache Kafka
The broker answered "accepted".
The message is gone.
It starts from zero: what Kafka is, where it is used and how it works. From there every chapter opens with a failure you must REPRODUCE before you fix it: an acknowledgement that meant nothing, a consumer that did not advance all night, a stranger’s group stalled on a renamed field, duplicates on top of enabled transactions. All in the browser, in Python, with no installation and no cluster.
- Grading checks what the code did, not its text
- Zero installation: no Docker, no cluster, no card
- The whole first chapter is free — 7 lessons
- Python, not Java: every bit of code is a cell in the browser
What hurts
A failure from work — and the chapter that takes it apart
On the left, what people arrive with. On the right, the method from the program that closes it — and the chapter where the failure is first reproduced and then fixed.
The report is ready by morning; they ask about "right now"
a log instead of a reportWhat Kafka is and where it is used, the log on disk, partitions and keys, compaction down to a lookup. Plus one honest lesson on when you do not need Kafka.
The client got an "accepted" and the message is not in the log
what acks promiseAn acknowledgement says who managed to write, not that it is safe. You lose a message three ways by hand — twice one already called accepted.
A consumer runs all night and does not advance one message
the consumer loop and rebalancesCommitting before processing loses messages, committing after gives duplicates. A consumer is thrown out of its group when a batch takes too long, and a rebalance is priced in seconds of downtime and duplicates.
A field was renamed and someone else’s group stalled
a schema in the registryThe schema moves out of both codebases into a registry. A schema checks the shape; only your own data check covers the meaning.
Transactions are on and the mart still has duplicates
where exactly-once stopsExactly-once is assembled from four parts, and beyond Kafka’s boundary the sink assembles it — not the transaction.
The panel reset on restart — and grows while the source is silent
processor state and windowsA counter that survives a restart; the moment a slice of the stream can be published; event time against processing time.
The board is green and the data is half an hour behind
a diagnosis from a pair of numbersOne number gives no diagnosis, and you must hear the one that went quiet. A rack failure, a DLQ you can return from, a stranger on your bandwidth.
This course is for you if
- A data engineer or analyst whose input became continuous while the tooling stayed batch
- Anyone new to Kafka: the course starts with what it is and where it is used, and goes all the way to replicas, transactions, stream processing and operations
- Anyone who writes Python: the Kafka course market is almost entirely Java and Scala, and there is none of either here
- Anyone whose interviews ask not "what is a topic" but "why Kafka here rather than a queue table in the database"
And NOT for you if
- Anyone looking for a cluster administrator’s course: installation, standing up racks, TLS configuration and hardware monitoring are not here, and that is stated in a section of its own below
- Anyone who does not write Python and SQL at all. The cells and tasks are Python and SQL, and their basics are not taught here
- Anyone expecting Java clients, Spring and Scala. All the code here is Python cells in the browser
- Anyone who needs a real cluster at hand: the broker here is a teaching one and deterministic, and we say so before you buy rather than after
How it works
Break it first, fix it second
Every lesson has an executable cell and every lesson has a check. What counts is not the text of your code but what it DID: what landed in the log, what was read, where the committed offset stopped.
You reproduce the failure yourself
A chapter starts not with a definition but with something broken: an acknowledgement with nothing behind it; a group that does not advance; a summary showing twice what arrived. First you reproduce the failure in a cell, and only then learn its name and its cure.
A teaching broker — but a deterministic one
The broker is emulated by the platform’s Python layer: no Docker, no cluster, no cloud account. The same run yields the same failure, so a breakdown can be reproduced on demand rather than "sometimes". Where the model is simpler than real Kafka a notice says so — a platform rule, not an author’s promise.
The result lands in PostgreSQL and ClickHouse
Landing the stream goes into the same PostgreSQL and ClickHouse sandboxes as every other course on the platform. The mart that showed twice what arrived, and the sink that killed it by inserting one message at a time, are real SQL over real data.
The seventh shift of one story
2 November 2184, 04:40. You switch the receiver on.
For half a year you ran the night intake belt. Then an antenna of Vault-9 caught a carrier that does not end: you cannot take it in batches, and while the belt waits for midnight half the signal is already worthless. The station has a mothballed comms level — the dispatch room, which once fed the live signal to the reading room, the intake level and the engine deck. You bring it back and take the watch.
Twenty-one days, twelve antennas, forty-eight frames a second. Every chapter adds one more participant to the feed — a consumer or a supplier — and each newcomer breaks what worked yesterday. The mentor is the same: QUERY, the old archive intelligence with the temper of a cat. But here he has a thread of his own: the dispatch room is the one level of the station where he worked BEFORE you, and when you ask why it was shut down, his first-chapter answer is: "Ask me again the first time an acknowledged write disappears on you."
The route
Four stretches of one shift
The course is not a list of broker settings. Every chapter plugs one more participant into the feed, and it breaks exactly what worked yesterday.
- I
The basics
What Kafka is and how it works: the log, partitions and keys — and an honest look at when you do not need it
- kf1Why Kafka: a log instead of a report
- II
Both ends of the wire
A write that was not lost, and a read that neither doubled nor stalled
- kf2The producer: writing without loss
- kf3The consumer: reading without duplicates or stalls
- III
Strangers on the line
The message contract, exactly-once, and your own processor between the feed and the consumer
- kf4Data format and Schema Registry
- kf5Delivery guarantees and transactions
- kf6Processing the stream
- IV
The shift handed over
The edges of the stream, a night under failure, and a capstone that must survive a second run
- kf7Integration: Kafka Connect, CDC and outbox
- kf8Operations: monitoring, failures, access
- kf9The shift handed over: the capstone
- kf10Going deeper
The program
10 chapters, 59 lessons, 19 h 5 min
The contents are approved and match the canon — a test enforces that, not a promise. A chapter whose lessons are already in the database expands into links and progress; the rest are waiting for their wave.
Why Kafka: a log instead of a report
Your first shift on this side of the dispatch room: the curator asks about "right now" and all you have is yesterday’s report. The chapter explains what Kafka is, what parts it is made of and where it is used, and honestly works through when you do not need it at all. Then it shows how a topic works inside: a log of segments on disk, retention, compaction into a lookup, and a partition count worked out by arithmetic.
The producer: writing without loss
What happens to a message between "sent" and "in the log" — and why those are not the same thing. You lose messages three different ways with your own hands, and twice those are messages the broker had already called accepted. By the end you can state the price of each durability setting, not just its name, and prove with numbers that "Kafka is slow" is usually "our own gateway is waiting".
The consumer: reading without duplicates or stalls
It is not the feed that breaks, it is the reading of it. First, how a consumer reads: the poll loop, position and commit, and why committing before processing loses messages while committing after gives duplicates. Then a live consumer works all night without advancing by a single message, and a rebalance costs seconds of downtime and duplicates. It ends with a lag board that shows not "bad" but what exactly to fix, and a re-read window that does not double a single row in the mart.
Data format and Schema Registry
The shape of a message stops being a private arrangement between two ends you own the moment a stranger joins the line. This chapter teaches you to make the arrangement machine-checkable: a message is taken apart into five pieces, the schema moves out of both codebases into Schema Registry, and a compatibility mode is read as the answer to "who ships first". It ends at the boundary no registry covers: a schema checks the shape, only your own data check covers the meaning.
Delivery guarantees and transactions
"Exactly once" is not a checkbox in a config but a construction of four parts. You learn to see delivery as a chain of links and to measure it with a four-number reconciliation, for as long as that reconciliation still proves anything. Then producer transactions: what they cover and why the committed offset belongs inside them. And then the boundary: outside Kafka, exactly-once is assembled by the sink itself.
Processing the stream
For five chapters Kafka was the thing that broke. Here it behaves perfectly in every lesson, and every time the culprit is the code between the feed and the consumer — yours. A processor that does not stall on one unparseable message; a counter that survives a restart; windows over an endless stream and knowing when a window can be closed; event time against processing time and what to do with latecomers; and choosing the tool — Kafka Streams, Flink, Spark or ksqlDB.
Integration: Kafka Connect, CDC and outbox
The edges of a stream are where it stops being a stream. On one side someone else’s database arrives in a topic through Kafka Connect rather than your own script: CDC, a key, updates and deletes instead of a nightly full reload. On the other side the same stream goes into a mart, and you find out why a database write and a topic send do not glue into one action, why a per-minute summary shows twice what arrived, and why "insert immediately" kills the mart by evening.
Operations: monitoring, failures, access
A night where everything is visible before the customer calls. A diagnosis comes from a pair of numbers, not one, and you have to hear the number that went quiet as well as the loud one. A rack goes down without a single action from you, and you have exactly one control. A DLQ is built so that a message can actually be brought BACK out of it. And access: who you are to the cluster and what you are allowed to do — SASL, TLS and ACLs.
The shift handed over: the capstone
The Academy sends a five-line letter and a two-day deadline. This chapter teaches you to characterise a stream before writing any code, and to assemble intake, processing, the mart and observability so that they work AT THE SAME TIME rather than one after another. A second run over the same data must produce the same result — and that is proved by a number, not by a claim. It closes with an exam on a different window of the feed.
Going deeper
Deep dives for people who already work with Kafka or want to go further than the main course: producer performance, heartbeats and static membership, reading only committed transactions, stream and table, joining two streams, Kafka Connect SMTs and converters, quotas, KRaft and tiered storage, share groups. The main course does not depend on these lessons — take them if you need them.
The honest boundary
What is not in the course, and why
You should learn the edge of your skill before a recruiter finds it for you. So it is stated here, before you buy, and repeated in the course’s second-to-last lesson as a map of the neighbouring territories.
A real cluster and its administration
A Python cell runs in the browser: it has neither sockets nor a network. Installation, standing up racks, configuring TLS and Kerberos, hardware monitoring and chaos testing all need machines the reader does not have.
The broker is a teaching one and deterministic: replicas, acknowledgements, silence timeouts, rebalances and transactions are simulated so that the same run yields the same failure. Every lesson where the model departs from real Kafka carries a notice saying what is real and what is simplified — a platform rule, not an author’s promise.
Kafka Streams, Flink and Spark as tools
All three live on the JVM, and there is no JVM in a browser. They cannot be configured inside a lesson, and pretending otherwise would mean selling slides as practice.
You write stream processing by hand on the Python layer — building the mechanism rather than configuring a ready-made one. The chapter’s last lesson pays that back: it shows which of the things you wrote were already written for you, and by what criterion people choose between a library and a cluster runtime.
Cluster security as a whole
Issuing certificates, Kerberos and configuring transport encryption are work on machines and in other people’s directories. There is no honest way to show that in a browser.
Only the part a data engineer actually does is taken from the topic: the access control list for a topic — who may do what — and the case where a stranger’s client ate the bandwidth and was given a quota. The rest is named by name in the map of neighbouring territories, so that you know what you do not know.
FAQ
Frequently asked
No — and that is not "we made installation easier", there is none. The teaching broker runs as a Python cell right in your tab: no Docker, no cluster, no cloud account, no credit card. Courses on the market spend anywhere from a module to half the program on installing and running a cluster; here that time goes into semantics.
The shift starts at 04:40
The first chapter is open in full — 7 lessons, no subscription and no card. You decide after that.