Home
ArenaGraphSignalTopics
/Apache Kafka and Event-Driven Systems: Building Real-Time Streaming Pipelines
Chapter 10 • Module 3 8 min breakdown +15 XP Module

Tiered Storage in Kafka and Redpanda: Decoupling Compute from Long-Term Storage

From Track:Apache Kafka and Event-Driven Systems: Building Real-Time Streaming PipelinesEvent-Driven Architecture & Distributed Systems

Historically, Apache Kafka coupled Compute (Broker CPUs & Memory) with Storage (Local NVMe SSDs or EBS Volumes).

If an enterprise needed to retain 1 year of event history for machine learning training and compliance auditing, they were forced to provision dozens of massive broker instances purely to satisfy disk space requirements, even if of their broker CPUs sat idle.

Tiered Storage (introduced to Kafka via KIP-405 and pioneered by modern engines like Redpanda and WarpStream) solves this by decoupling compute from long-term storage, slashing streaming infrastructure costs by up to .


1. The Monolithic Storage Problem vs Tiered Storage

Interactive Blueprint
Rendering diagram...

2. Transparent Read Mechanics: Zero Consumer Code Changes

One of Tiered Storage's most powerful capabilities is complete API transparency:

Interactive Blueprint
Rendering diagram...

3. Instant Broker Scaling: Eliminating Rebalance Migrations

In standard Kafka, adding a broker requires copying hundreds of gigabytes of historical log files across the cluster network.

With Tiered Storage:

  1. Historical segments reside securely in S3/GCS.
  2. Only the active local segments (recent 24 hours) need to be mirrored.
  3. Partition reassignment and cluster scaling complete in seconds instead of days, enabling true elastic autoscaling in Kubernetes.

4. Production Configuration: Enabling Tiered Storage (KIP-405)

Broker Server Properties (server.properties):

ini
Loading code editor...

5. Architectural Comparison

DimensionStandard Kafka (Local Storage Only)Tiered Storage (Kafka / Redpanda / WarpStream)
Storage CostHigh (0.30 per GB/mo on SSD/EBS)Ultra-Low (0.023 per GB/mo on S3)
Historical Data RetentionTypically limited to 3-7 days due to costMonths or Years (Infinite Retention)
Broker Rebalance TimeHours to Days (Terabytes copied over network)Seconds (Only active heads copied)
Cluster AutoscalingPainful / SlowElastic & Rapid
Real-Time Read LatencySub-millisecond (PageCache)Sub-millisecond (Identical local cache)

6. Summary & Recommendations

  1. Adopt Tiered Storage for Historical Analytics: If your organization needs to re-train machine learning models or audit financial events over long retention windows, Tiered Storage is the most cost-effective architecture available.
  2. Size Local NVMe for Real-Time Consumers: Ensure local retention is sized large enough () so that standard real-time microservice consumers never hit S3, eliminating object storage egress costs during normal operations.
  3. Set Up S3 Lifecycle Rules: For ultra-long-term archiving (), configure S3 lifecycle transitions to glacier or flexible storage classes.
Milestone Verification

Ready for the next lesson?

Mark this module complete to record verified progress and earn +15 XP toward your architect profile.