Hi, I’m Nikhil Aggarwal
I’m an ML engineer with 15+ years of experience in data engineering, cloud computing and distributed systems. Today I build ML and Generative AI capabilities for enterprise data platforms. I’ve spent most of my career building the pipelines those systems depend on, so I understand both sides.
What I work on
- ML and GenAI platforms: RAG-based operational intelligence, AI-assisted pipeline onboarding, and reusable ML/AI capabilities such as AI guardrails.
- Data pipelines at scale: automated ETL workflows with Airflow, Jenkins and CI/CD, bringing many sources into AWS and Snowflake.
- Performance and architecture: system design, data modelling and tuning Spark and Snowflake workloads in production.
Tools I use daily: Python, PySpark, SQL, Spark, Snowflake, Hive, Impala, Teradata, PostgreSQL, Redshift, DynamoDB, and AWS (Glue, EMR, Kinesis, Lambda, S3, SQS, SNS, IAM, CloudWatch).
Why DataForGeeks
I write about the problems I’ve actually hit in production, with real screenshots, working code and the reasoning behind each decision. The goal is to make complex topics simple enough to use at work the next day, or to explain clearly in an interview.
Where to start
- Spark performance: How to Debug a Slow Spark Job and Fix Data Skew · Spark Performance Tuning
- Snowflake: Running Snowpipe in Production at Scale · Snowflake Performance Tuning
- Lakehouse: Apache Iceberg Explained · Iceberg Hilbert Curve · Medallion Architecture
- System-design interviews: Ride-Hailing System Design · Third-Party Ingestion with Sub-Minute Dashboards
- AI for engineers: Running Multiple AI Coding Agents in Your Terminal
- Practice: SQL Problems with Solutions
Background
I’ve received several awards for innovation and teamwork at the organisations I’ve worked with. I also placed in the top 1% of GATE 2017 (All India Rank 3,253).
Get in touch
Connect with me on LinkedIn, or subscribe to the newsletter to get new articles by email.