3+ years building scalable batch and real-time ETL/ELT pipelines with Python, SQL, PySpark, and Apache Airflow — processing 100M+ daily events and 2TB+ of data for analytics, risk, compliance, billing, and Customer 360 use cases across telecom and finance.
SV / Data Engineer
I build pipelines that data teams can trust — the kind that hold up under regulatory scrutiny and peak-hour load alike. My work sits at the intersection of telecom-scale batch processing and finance-grade compliance analytics, turning raw usage logs and transaction records into datasets that risk, compliance, and reporting teams depend on every day, developed and optimized to process 100M+ daily events and 2TB+ of data.
Built a batch ETL pipeline extracting 1M+ records/day from a public API into an AWS S3 data lake using PySpark on EMR. Modeled a star-schema warehouse in Redshift, orchestrated with Airflow (99%+ DAG success), and visualized in Tableau.
Built a real-time Kafka pipeline streaming 100K+ clickstream events/hour into Spark Structured Streaming with sub-second latency. Applied watermarking and 5-min windowed aggregations, persisting results to S3 and Redshift for near-real-time analytics.