Batch report or stream: how long an answer waits
O que você vai aprender
- to measure answer age: from the moment of an event to its row in a published report
- to write a requirement down as a number and check it against the unluckiest event
- to compute the age for any schedule: typical is half the period plus the run, worst is the whole period plus the run
- to find the wall: the point where run start-ups alone keep the machine busy around the clock
- to pick the longest period that fits both the requirement and the budget
06:40. How many antennas are silent right now
2 November 2184, Vault-9 station. It is an orbital archive: the station's antennas catch signals from old Earth, and you are the archivist who makes sense of them. For half a year you ran the night intake belt: once a day, collect everything that has arrived, run the calculation until morning and have a report ready by breakfast. Three days ago an antenna caught a carrier — a signal that does not end. You cannot take it in batches, once a day. The station turned out to have a mothballed comms level — the dispatch room, where the live feed was once handed out. At 04:40 you brought it back to life and took the watch.
At 06:40 you bring the station's curator the morning report — the same daily report the night belt has been assembling for half a year. The belt closed the day at 00:00, the run took 6 hours 12 minutes, and the report was printed at 06:12.
The curator looks at the page and asks: “How many antennas are silent right now?”
The report says one. The intake board behind his back says seven out of twelve.
Nobody is lying. The report and the board count the same thing — how many antennas failed to deliver their intake to the night belt. But the report counts over the past day, and the board over the last hour. The report is honest as of 23:59 yesterday. The question is about “now”, yet the freshest number in the report is six hours forty-one minutes old, and the oldest is thirty hours forty minutes old. While you talk, the report keeps ageing: the next one comes out tomorrow at 06:12.
The feed — the continuous signal itself — has nothing to do with it: it runs steadily from all twelve antennas. And Kafka will not appear in this lesson even once. First we need to figure out when it is needed at all. That comes down to one number: how old the answer you are bringing the curator is.
QUERY, the station's AI mentor with the manners of a cat, opens one eye.
QUERY: A good report. About yesterday.

Answer age
The number we have just computed gets a name for the whole course: answer age — how much time passes between an event happening and that event first appearing in a published answer.
The table measures age at printing time, 06:12, which is why it shows 30 h 12 min, while for the curator at 06:40 it is already 30 h 40 min.
The table shows two rules:
- on average, an event waits half the period plus the run to be published — that is the “typical age” column;
- the unluckiest event waits the whole period plus the run — the “worst age” column. It happened right after a run picked up the data, so it waits out the entire next period and the entire run.
The lesson's second term is answer shelf life: how long an answer stays useful to whoever makes decisions based on it. It is compared with the worst age, not the average: whoever acts on the answer does not get to choose which event turns out to be the unlucky one. For the curator the shelf life is an hour: the board counts over an hour. And the reading room — you start feeding it today, and it answers the station's questions from the feed — has a shelf life of just 30 seconds.
Look at the worst-age column: not a single schedule fits into 30 seconds. Even a run every 40 seconds gives a worst age of 1 minute 20 seconds.
Why not count even more often: the wall
You could shorten the period further. But a run has a fixed part — work it does no matter how much data has arrived: starting the program, finding where to read from, writing down how far it got. That part does not shrink along with the period; it gets multiplied by the number of runs — that is the “start-ups” column in the table.
Once a day, start-ups cost 40 seconds. Once a minute — 16 hours a day: the machine spends two thirds of the day doing nothing but starting up. Once every 40 seconds — all 24 hours. That is the wall: there is nowhere left to shorten the period, because the machine is already fully busy with start-ups.
Hence the main conclusion of the lesson. For the worst age to fit into 30 seconds, the period has to be no longer than 30 seconds minus the run. Yet a start-up alone takes 40. No schedule fits.
Streaming is not free: it costs more to run and is harder to fix, so a frequent scheduled run is a perfectly normal choice as long as the requirement allows it. Here it does not. Streaming has no period, and its answer age is the time it takes to process one event: 1.2 seconds.
Before the task: how busy the machine is
In the table the budget was counted in start-ups: the number of runs multiplied by one start-up. The task below uses a stricter budget — it counts machine utilization: the number of runs multiplied by the whole run, start-up and work included.
Let's work through an example on paper. The requirement is a worst age of at most two hours, a run takes 15 minutes, and the budget is half of the machine's day:
- worst age = period + run ≤ 7200 s, so period ≤ 7200 − 900 = 6300 s;
- utilization = (86,400 / period) × 900 ≤ 43,200, so period ≥ 1800 s;
- the period cannot be shorter than the run itself — 900 s; that bound is weaker than the previous one.
Any period from 1800 to 6300 seconds works. The longest is 6300: running less often is cheaper.
max_period(requirement_s, run_s, budget_s): the longest run period, in whole seconds, that fits both the requirement and the budget — or None if none fits.
The model is three lines:
- worst answer age = period + run duration, and it is no more than
requirement_s; - machine utilization per day = (86,400 / period) × run duration, and it is no more than
budget_s; - the period cannot be shorter than the run duration.
What in this lesson is real and what is simulated
There is no Kafka in this lesson yet: answer age is computed by arithmetic over a ready-made list of intakes, and that is fine — it does not depend on the tool.
The training constants — 40 seconds per run start-up, a 6 h 12 min daily run and 1.2 seconds to process an event in streaming — are not measurements from a live system: measure them on your own, because they decide where the wall stands.
What is real in the lesson is the method: the typical answer age is half the period plus the run, the worst is the period plus the run, and the time spent on start-ups alone grows with the number of runs and hits the machine's full day sooner than the answer age hits the requirement.
Interview question
How this comes up in interviews
The question goes: “You already have a nightly ETL — why do you need Kafka?” The answer “Kafka is faster and more scalable” misses: they asked about a requirement and you answered about a technology.
The right answer is a number. Name the answer's shelf life from the side of whoever acts on it, and show with arithmetic that a schedule does not fit: under any schedule the worst age is at least the period plus the run, and a single run start-up takes us 40 seconds.
The second question checks maturity: “Where is your wall?” They expect you to name the fixed part of a run and to say that a frequent scheduled run is a normal, honest answer as long as the period is longer than a few minutes. Below that it degenerates into a never-ending start-up. A stream is not the only way around the wall: micro-batch engines — Spark Structured Streaming, for one — keep the process running, pay for the start-up once and take batches seconds apart.
The industry names: answer age is end-to-end latency or freshness, shelf life is latency SLO or freshness , and the fixed part of a run is per-run overhead.
What next: the log
A daily report cannot be made fresh: however often you schedule it, you hit the wall. You need a different form of data — one you can enter at any moment and at any point. That form is the log.
It is simple. New events are only ever appended at the end, like entries in a ship's log, and each one gets the next number in order. What has been read is not crossed out: one reads from the morning, another catches up after lunch, a third rereads yesterday — and nobody gets in anyone's way.
The program that keeps such logs and hands them out to many systems at once is Kafka. What it is, what parts it is made of and where it is used — that is the next lesson.
Principais pontos
| schedule | worst age | start-ups per day |
|---|---|---|
| once a day | 30 h 12 min | 40 s |
| once an hour | 64 min | 16 min |
| every 40 seconds | 1 min 20 s | 24 h — the wall |
| streaming | 1.2 s | no period |
Answer age — from the event to its row in a published answer. Answer shelf life — how long until the answer is useless to whoever acts on it; it is compared with the worst age. For the reading room it is 30 seconds, and no schedule fits into it: the wall comes first.
QUERY: The curator asked about “now”. Reports are never about “now”. Logs are.