← All work

01 · Handed over

Multi-Cloud VM Activity Log Platform

One ingestion framework for Azure, AWS and GCP VM logs — new fields merge themselves, type conflicts wait for a human.

Period
Mar 2026 – May 2026
Role
Freelance Data Engineer · Data Insight Consulting
Stack
Python · PySpark · Databricks · Delta Lake · Azure · AWS
Bronze Delta tables
60
of runtime spent on validation
<5%
clouds, one framework
3

The problem

Virtual machines run on three clouds, and each cloud writes its activity logs in its own shape. The schemas change without notice. Before this project, every new field meant a manual schema update — and a broken pipeline whenever someone forgot.

What I built

A single, reusable ingestion framework on Databricks that reads raw logs from Azure Blob Storage and AWS S3 (GCP uses the same framework) and lands them in a Medallion architecture. At its core is a four-layer schema registry:

  1. Discover — one spark.read.json pass infers the widest schema; type conflicts fall back to strings.
  2. Diff — compare against the last known snapshot and sort every change into three tiers: new scalar fields merge silently, new nested structures merge with an alert, type conflicts block and wait for human review.
  3. Flatten — nested fields become flat columns with safe names; the raw row is always kept in raw_json.
  4. Validate — null rates for every column are computed in one aggregation; anomalies stop the write.

Key design decisions

  1. Skip discovery when stable — the registry snapshot is used to flatten directly.
  2. Block, don’t guess — type conflicts or anomalous null rates stop the write for a human.
  3. Keep the raw row — raw_json lets flattening be replayed at any time.
  4. Validate in one pass — every null rate in a single agg(), under 5% of runtime.
Architecture — multi-cloud log ingestion with a four-layer schema registry

Azure and AWS logs (GCP planned) flow through discover, diff, a drift decision, flatten and validate. New fields merge and are written back to the schema registry; type conflicts or anomalous null rates stop the write for human review; when nothing changed, discovery is skipped and the registry snapshot is used directly. Validated data lands in Bronze, then Silver, Gold and Power BI.

The impact

  • 60 Bronze Delta tables from 40+ Azure and AWS log sources, all through one framework.
  • No manual schema maintenance: safe changes merge automatically, risky ones are stopped before they reach a table.
  • Schema checks cost less than 5% of pipeline runtime, thanks to single-pass aggregation and skipping discovery when nothing has changed.

Key decisions

  • Keep the raw row. Every Bronze row carries its original JSON, so a bug in flattening can be replayed without re-ingesting.
  • Block, don’t guess. A type conflict is an audit question, not an engineering one: the pipeline stops and asks a human.
  • Validate in one pass. All null-rate checks run in a single agg() call instead of one query per column.