Naveed Mohiuddin
Data Engineer
Hands-on experience building batch and incremental pipelines across AWS, Azure, and Databricks with an emphasis on reliable reruns, data quality, production troubleshooting, and analytics-ready datasets. Two AWS certifications and a Master's in Computer Science.
About Me
Engineer, not just a resume.
I'm a Data Engineer with experience building cloud-native data platforms that serve real business needs from relational and API source systems in insurance/financial data domains to serverless lakehouse architectures on AWS. My work sits at the intersection of software engineering and data infrastructure: I design ETL systems, optimize distributed processing, model dimensional data, and automate everything I can.
I hold a Master's in Computer Science from Illinois Institute of Technology and a Bachelor's from Osmania University. I'm AWS certified in both Solutions Architecture and Data Engineering, and I've worked across AWS and Azure stacks in production settings. I care about building systems that are reliable, cost-efficient, and maintainable not just technically interesting.
Tech Stack
Tools I work with daily.
Production-tested across AWS, Azure, and open-source data ecosystems.
Certifications
AWS Certified, production validated.
Industry-recognized credentials demonstrating cloud architecture and data engineering depth.
AWS Certified Solutions Architect – Associate
Demonstrates ability to design secure, scalable, and cost-optimized cloud architectures using AWS services.
AWS Certified Data Engineer – Associate
Validates expertise in designing and maintaining data pipelines, data stores, data processing, security, and governance on AWS.
Microsoft Certified: Azure Fundamentals
Covers core Azure services, cloud concepts, security, governance, and pricing fundamentals.
Experience
Where I've built things that matter.
Production systems, real data, measurable outcomes.
Data Engineer
Jan 2025 – PresentBenda Infotech — Remote, US
- Build cloud data pipelines across AWS, Azure, and Databricks covering batch ingestion, transformation, validation, and reporting use cases — working across both stacks depending on where the source systems and consumers live.
- On AWS, develop Lambda and EventBridge-based ingestion into S3 with clearly separated raw and curated zones, designed so any load can be traced back to its source file and reprocessed in isolation without touching downstream data.
- On Azure, build ADF and Databricks pipelines using Delta Lake medallion layers (Bronze/Silver/Gold) for incremental ingestion, transformation, and curated reporting tables that analysts query directly.
- Write PySpark and Spark SQL logic for data cleaning, validation, joins, reconciliations, and curated table creation — standardizing types and applying business rules before data reaches reporting layers.
- Use Delta MERGE INTO alongside deterministic S3 object naming so that reruns are safe by design, which cut down the duplicate-load issues that previously required manual cleanup.
- Tune Athena and Delta workloads using columnar Parquet, partitioning, partition pruning, OPTIMIZE, and Z-ORDER to reduce both query latency and bytes scanned per query.
- Add record-count, load-completion, and quarantine checks at layer boundaries so failed or incomplete loads are caught in the pipeline rather than discovered by an analyst looking at a broken dashboard.
- Manage orchestration through Airflow and MWAA alongside ADF and Databricks Workflows, configuring retries, alerting, task dependencies, and parameterized backfill for reprocessing historical windows.
- Build Athena, Databricks SQL, QuickSight, and Power BI-ready datasets for operational and business reporting, and partner with analysts to translate reporting requirements into source-to-target mappings, validation rules, and reusable data models.
Data Engineer
Feb 2021 – Jul 2023Applied Information Sciences — Hyderabad, India
- Built and supported Azure-based data pipelines for policy, claims, and operational data spanning ADF, Databricks, Snowflake, and SQL-based source systems.
- Developed ADF ingestion pipelines pulling from SQL Server, Oracle, REST APIs, and flat files into ADLS Gen2 and Snowflake, handling the schema and connectivity differences each source type brought with it.
- Used control-table driven pipeline patterns so onboarding a new source became largely a configuration change rather than net-new ADF development, which cut repeated build work across the team.
- Implemented full-load and incremental extraction using watermark columns, runtime parameters, and validation checkpoints so pipelines could run reliably on a schedule without manual date handling.
- Wrote PySpark, Spark SQL, and SQL transformations to standardize schemas, apply validation rules, and prepare curated datasets for the Snowflake warehouse.
- Built reconciliation queries comparing source and target counts, checking load timestamps, and surfacing duplicate or missing records ahead of reporting cycles — catching issues before business users did.
- Investigated production failures across ADF, Databricks, Snowflake, and Splunk, tracing root cause through logs and dashboards, then executing controlled reruns and verifying downstream tables were correct.
- Maintained runbooks covering recurring failure patterns, validation steps, rerun procedures, and escalation paths so on-call handoffs didn't depend on tribal knowledge.
- Worked with business analysts, QA, and support teams to confirm reported data issues, validate fixes, and keep reporting cycles moving during month-end crunch periods.
Projects
Engineering projects, not just exercises.
Each project solves a real data engineering problem end-to-end.
Chicago Crime Analytics Data Lake
Scheduled serverless pipeline pulling public Chicago crime data into a queryable, partitioned S3 data lake with BI dashboards on top.
- Lambda and EventBridge pull records from the Socrata API on a schedule and store raw responses in S3
- Cleaned output written as date-partitioned Parquet for efficient Athena querying
- Curated tables registered in the Glue Data Catalog for schema management
- QuickSight views covering category, location, and time-of-day trends
- Raw layer preserved so the full history can be replayed when transformation logic changes
Azure Databricks Lakehouse Pipeline
Governed medallion lakehouse on Databricks combining batch and streaming ingestion with data-quality enforcement and access control.
- Bronze/Silver/Gold architecture with ADF for ingestion and PySpark for validation and transformation
- Delta Lake providing versioned, replayable storage with time travel and schema evolution
- Delta Live Tables expectations enforcing data quality declaratively within the pipeline
- Kafka producers with Structured Streaming consumers for near-real-time ingestion
- Unity Catalog access controls covering governance, lineage, and column-level permissions
Big Data Processing with Spark
Multi-format distributed processing on a Hadoop/HDFS cluster with hands-on Spark performance tuning.
- Processed CSV, JSON, and XML datasets on a GCP Dataproc Hadoop/HDFS cluster
- PySpark and Spark SQL for schema checks, cleansing, joins, and aggregations against a Hive metastore
- Traced a shuffle-heavy join through the Spark UI to identify the bottleneck
- Switched the smaller side to a broadcast join and retuned shuffle partitions, cutting data movement and job runtime
Why Hire Me
What I bring to your team.
Not just skills on paper, here's why it translates to real value.
Multi-Cloud, Not Just One Stack
Production work across AWS, Azure, and Databricks comfortable picking the right tool rather than forcing everything into one ecosystem.
AWS Certified
Solutions Architect and Data Engineer Associate certifications, applied to real pipelines rather than kept on paper.
Full Data Lifecycle
Ingestion, transformation, modeling, orchestration, data quality, and the BI layer analysts actually consume.
Built for Safe Reruns
Idempotent load patterns, deterministic object naming, and quarantine checks pipelines designed so failures are recoverable, not catastrophic.
Production Troubleshooting
Root-cause analysis across ADF, Databricks, Snowflake, and Splunk, with runbooks so fixes don't depend on one person's memory.
Works With the Business
Translating reporting requirements into source mappings and validation rules alongside analysts, QA, and support teams.
Interested in discussing data engineering opportunities?
I'm actively looking for Data Engineer roles. Let's talk about how I can contribute to your team.
Contact
Let's connect.
I'm actively open to Data Engineer opportunities. If my background aligns with your team's needs, I'd love to hear from you.