01 · Handed over
Multi-Cloud VM Activity Log Platform
One ingestion framework for Azure, AWS and GCP VM logs — new fields merge themselves, type conflicts wait for a human.
- Period
- Mar 2026 – May 2026
- Role
- Freelance Data Engineer · Data Insight Consulting
- Stack
- Python · PySpark · Databricks · Delta Lake · Azure · AWS
- Bronze Delta tables
- 60
- of runtime spent on validation
- <5%
- clouds, one framework
- 3
The problem
Virtual machines run on three clouds, and each cloud writes its activity logs in its own shape. The schemas change without notice. Before this project, every new field meant a manual schema update — and a broken pipeline whenever someone forgot.
What I built
A single, reusable ingestion framework on Databricks that reads raw logs from Azure Blob Storage and AWS S3 (GCP uses the same framework) and lands them in a Medallion architecture. At its core is a four-layer schema registry:
- Discover — one
spark.read.jsonpass infers the widest schema; type conflicts fall back to strings. - Diff — compare against the last known snapshot and sort every change into three tiers: new scalar fields merge silently, new nested structures merge with an alert, type conflicts block and wait for human review.
- Flatten — nested fields become flat columns with safe names; the raw row is always kept in
raw_json. - Validate — null rates for every column are computed in one aggregation; anomalies stop the write.
Key design decisions
- 1Skip discovery when stable — the registry snapshot is used to flatten directly.
- 2Block, don’t guess — type conflicts or anomalous null rates stop the write for a human.
- 3Keep the raw row —
raw_jsonlets flattening be replayed at any time. - 4Validate in one pass — every null rate in a single
agg(), under 5% of runtime.
Azure and AWS logs (GCP planned) flow through discover, diff, a drift decision, flatten and validate. New fields merge and are written back to the schema registry; type conflicts or anomalous null rates stop the write for human review; when nothing changed, discovery is skipped and the registry snapshot is used directly. Validated data lands in Bronze, then Silver, Gold and Power BI.
The impact
- 60 Bronze Delta tables from 40+ Azure and AWS log sources, all through one framework.
- No manual schema maintenance: safe changes merge automatically, risky ones are stopped before they reach a table.
- Schema checks cost less than 5% of pipeline runtime, thanks to single-pass aggregation and skipping discovery when nothing has changed.
Key decisions
- Keep the raw row. Every Bronze row carries its original JSON, so a bug in flattening can be replayed without re-ingesting.
- Block, don’t guess. A type conflict is an audit question, not an engineering one: the pipeline stops and asks a human.
- Validate in one pass. All null-rate checks run in a single
agg()call instead of one query per column.