Project Overview
A digital media company with over 30 million monthly active users was running on a legacy batch ETL system that could not keep up. Data was arriving hours late, dashboards were stale by the time editors checked them, and the engineering team spent more time firefighting pipeline failures than building features. They needed a modern, real-time data architecture but could not afford downtime during the migration.
We ran a two-week architecture sprint to map every data source, transformation, and downstream consumer. From there, we designed a streaming-first architecture using Apache Kafka for event ingestion, Apache Flink for real-time transformations, and a lakehouse layer for historical analytics. The migration was planned as a parallel-run strategy with new pipelines running alongside legacy systems with automated data quality checks.
Over 10 weeks, we migrated 47 data pipelines from batch to streaming, implemented schema registry and data contracts, and built a self-serve analytics layer for the editorial and product teams. The new system processes over 2 billion events per day with sub-second latency.
Pipeline failures dropped by 90%, and the editorial team now has access to real-time content performance metrics. The engineering team reclaimed 30% of their sprint capacity that was previously spent on data infrastructure maintenance.
Key Takeaways
- Parallel-run migration strategies minimize downtime risk
- Schema registries and data contracts prevent downstream breakage
- Real-time analytics unlock editorial and product agility
- Modern streaming architectures dramatically reduce maintenance burden
Challenge
A legacy batch ETL system could not keep up with 30M+ monthly active users. Data arrived hours late, dashboards were stale, and engineers spent more time on pipeline failures than feature work.
Strategy
We mapped every data source and downstream consumer in a two-week sprint, then designed a parallel-run migration strategy to avoid downtime while transitioning from batch to streaming.
Solution
We migrated 47 pipelines to a streaming-first architecture using Apache Kafka and Apache Flink, implemented schema registry and data contracts, and built a self-serve analytics layer. The system now processes 2B+ events daily with sub-second latency.
Digital Media Engagement
(Confidential)
