AWS Certified Data Engineer AWS Certified Solutions Architect United States

Naveed Mohiuddin

Data Engineer

Hands-on experience building batch and incremental pipelines across AWS, Azure, and Databricks with an emphasis on reliable reruns, data quality, production troubleshooting, and analytics-ready datasets. Two AWS certifications and a Master's in Computer Science.

$ naveed.skills("AWS", "Azure", "PySpark", "Databricks")
2
AWS Certifications
MS
Computer Science
AWS + Azure
Multi-Cloud

About Me

Engineer, not just a resume.

I'm a Data Engineer with experience building cloud-native data platforms that serve real business needs from relational and API source systems in insurance/financial data domains to serverless lakehouse architectures on AWS. My work sits at the intersection of software engineering and data infrastructure: I design ETL systems, optimize distributed processing, model dimensional data, and automate everything I can.

I hold a Master's in Computer Science from Illinois Institute of Technology and a Bachelor's from Osmania University. I'm AWS certified in both Solutions Architecture and Data Engineering, and I've worked across AWS and Azure stacks in production settings. I care about building systems that are reliable, cost-efficient, and maintainable not just technically interesting.

MS Computer Science
Illinois Institute of Technology
BE Computer Science
Osmania University

Tech Stack

Tools I work with daily.

Production-tested across AWS, Azure, and open-source data ecosystems.

Cloud Platforms
AWS S3GlueGlue Data CatalogLambdaRedshiftAthenaEventBridgeKinesisSNSEC2RDSLake FormationAzure Data FactoryAzure DatabricksSynapse AnalyticsADLS Gen2Unity CatalogAzure DevOps
Data Engineering
ETL/ELTData-Lake & Medallion ArchitectureDimensional ModelingSCD Type 2Batch PipelinesIncremental PipelinesStreaming PipelinesData QualityData Governance & LineageSchema EvolutionReconciliation
Big Data & Lakehouse
Apache SparkPySparkSpark SQLStructured StreamingDelta LakeDelta Live Tables (Lakeflow)Auto LoaderDatabricks WorkflowsKafkaHadoopHDFSHiveMapReduceParquet
Databases & Warehouses
Amazon RedshiftAthenaAzure Synapse AnalyticsSnowflakeSQL ServerOraclePostgreSQLMySQL
Orchestration
Apache AirflowAmazon MWAADAG DesignSLA MonitoringRetries & BackfillAmazon EventBridgeAzure Data FactoryDatabricks Workflows
DevOps & Version Control
GitGitHubBitbucketAzure DevOpsGitHub ActionsJenkinsCI/CDDockerLinux
Programming
PythonSQLPySparkShell Scripting (Bash)JavaPandas
Analytics & BI
Amazon QuickSightPower BIAthena SQLDatabricks SQL

Certifications

AWS Certified, production validated.

Industry-recognized credentials demonstrating cloud architecture and data engineering depth.

SAA

AWS Certified Solutions Architect – Associate

Demonstrates ability to design secure, scalable, and cost-optimized cloud architectures using AWS services.

EC2VPCIAMLambdaS3RDS
DEA

AWS Certified Data Engineer – Associate

Validates expertise in designing and maintaining data pipelines, data stores, data processing, security, and governance on AWS.

S3GlueRedshiftAthenaLake FormationKinesis
AZ

Microsoft Certified: Azure Fundamentals

Covers core Azure services, cloud concepts, security, governance, and pricing fundamentals.

Azure CoreStorageGovernancePricing

Experience

Where I've built things that matter.

Production systems, real data, measurable outcomes.

Data Engineer

Jan 2025 – Present

Benda Infotech — Remote, US

AWS S3LambdaEventBridgeGlueAthenaMWAAAirflowAzure Data FactoryAzure DatabricksDelta LakeADLS Gen2PySparkSpark SQLDatabricks SQLParquetQuickSightPower BIPythonSQLGitAzure DevOps
  • Build cloud data pipelines across AWS, Azure, and Databricks covering batch ingestion, transformation, validation, and reporting use cases — working across both stacks depending on where the source systems and consumers live.
  • On AWS, develop Lambda and EventBridge-based ingestion into S3 with clearly separated raw and curated zones, designed so any load can be traced back to its source file and reprocessed in isolation without touching downstream data.
  • On Azure, build ADF and Databricks pipelines using Delta Lake medallion layers (Bronze/Silver/Gold) for incremental ingestion, transformation, and curated reporting tables that analysts query directly.
  • Write PySpark and Spark SQL logic for data cleaning, validation, joins, reconciliations, and curated table creation — standardizing types and applying business rules before data reaches reporting layers.
  • Use Delta MERGE INTO alongside deterministic S3 object naming so that reruns are safe by design, which cut down the duplicate-load issues that previously required manual cleanup.
  • Tune Athena and Delta workloads using columnar Parquet, partitioning, partition pruning, OPTIMIZE, and Z-ORDER to reduce both query latency and bytes scanned per query.
  • Add record-count, load-completion, and quarantine checks at layer boundaries so failed or incomplete loads are caught in the pipeline rather than discovered by an analyst looking at a broken dashboard.
  • Manage orchestration through Airflow and MWAA alongside ADF and Databricks Workflows, configuring retries, alerting, task dependencies, and parameterized backfill for reprocessing historical windows.
  • Build Athena, Databricks SQL, QuickSight, and Power BI-ready datasets for operational and business reporting, and partner with analysts to translate reporting requirements into source-to-target mappings, validation rules, and reusable data models.

Data Engineer

Feb 2021 – Jul 2023

Applied Information Sciences — Hyderabad, India

Azure Data FactoryADLS Gen2Azure DatabricksPySparkSpark SQLSnowflakeSQL ServerOracleSplunkAzure DevOpsREST APIsPythonSQLGit
  • Built and supported Azure-based data pipelines for policy, claims, and operational data spanning ADF, Databricks, Snowflake, and SQL-based source systems.
  • Developed ADF ingestion pipelines pulling from SQL Server, Oracle, REST APIs, and flat files into ADLS Gen2 and Snowflake, handling the schema and connectivity differences each source type brought with it.
  • Used control-table driven pipeline patterns so onboarding a new source became largely a configuration change rather than net-new ADF development, which cut repeated build work across the team.
  • Implemented full-load and incremental extraction using watermark columns, runtime parameters, and validation checkpoints so pipelines could run reliably on a schedule without manual date handling.
  • Wrote PySpark, Spark SQL, and SQL transformations to standardize schemas, apply validation rules, and prepare curated datasets for the Snowflake warehouse.
  • Built reconciliation queries comparing source and target counts, checking load timestamps, and surfacing duplicate or missing records ahead of reporting cycles — catching issues before business users did.
  • Investigated production failures across ADF, Databricks, Snowflake, and Splunk, tracing root cause through logs and dashboards, then executing controlled reruns and verifying downstream tables were correct.
  • Maintained runbooks covering recurring failure patterns, validation steps, rerun procedures, and escalation paths so on-call handoffs didn't depend on tribal knowledge.
  • Worked with business analysts, QA, and support teams to confirm reported data issues, validate fixes, and keep reporting cycles moving during month-end crunch periods.

Projects

Engineering projects, not just exercises.

Each project solves a real data engineering problem end-to-end.

Chicago Crime Analytics Data Lake

Scheduled serverless pipeline pulling public Chicago crime data into a queryable, partitioned S3 data lake with BI dashboards on top.

Problem: Public API data needed to land reliably on a schedule, stay replayable when transformation logic changed, and be queryable without standing up a warehouse.
  • Lambda and EventBridge pull records from the Socrata API on a schedule and store raw responses in S3
  • Cleaned output written as date-partitioned Parquet for efficient Athena querying
  • Curated tables registered in the Glue Data Catalog for schema management
  • QuickSight views covering category, location, and time-of-day trends
  • Raw layer preserved so the full history can be replayed when transformation logic changes
AWS LambdaEventBridgeS3GlueAthenaQuickSightPythonSQLParquet
View on GitHub

Azure Databricks Lakehouse Pipeline

Governed medallion lakehouse on Databricks combining batch and streaming ingestion with data-quality enforcement and access control.

Problem: Needed a lakehouse that handled both batch and near-real-time ingestion while enforcing data quality and governance inside the pipeline rather than bolting them on afterward.
  • Bronze/Silver/Gold architecture with ADF for ingestion and PySpark for validation and transformation
  • Delta Lake providing versioned, replayable storage with time travel and schema evolution
  • Delta Live Tables expectations enforcing data quality declaratively within the pipeline
  • Kafka producers with Structured Streaming consumers for near-real-time ingestion
  • Unity Catalog access controls covering governance, lineage, and column-level permissions
DatabricksDelta LakeDelta Live TablesUnity CatalogPySparkStructured StreamingKafkaAzure Data Factory
View on GitHub

Big Data Processing with Spark

Multi-format distributed processing on a Hadoop/HDFS cluster with hands-on Spark performance tuning.

Problem: Datasets exceeded single-machine capacity and an early join implementation was shuffle-bound, making runtimes unpredictable.
  • Processed CSV, JSON, and XML datasets on a GCP Dataproc Hadoop/HDFS cluster
  • PySpark and Spark SQL for schema checks, cleansing, joins, and aggregations against a Hive metastore
  • Traced a shuffle-heavy join through the Spark UI to identify the bottleneck
  • Switched the smaller side to a broadcast join and retuned shuffle partitions, cutting data movement and job runtime
PySparkSpark SQLHDFSHiveGCP DataprocMapReducePython
View on GitHub

Why Hire Me

What I bring to your team.

Not just skills on paper, here's why it translates to real value.

Multi-Cloud, Not Just One Stack

Production work across AWS, Azure, and Databricks comfortable picking the right tool rather than forcing everything into one ecosystem.

AWS Certified

Solutions Architect and Data Engineer Associate certifications, applied to real pipelines rather than kept on paper.

Full Data Lifecycle

Ingestion, transformation, modeling, orchestration, data quality, and the BI layer analysts actually consume.

Built for Safe Reruns

Idempotent load patterns, deterministic object naming, and quarantine checks pipelines designed so failures are recoverable, not catastrophic.

Production Troubleshooting

Root-cause analysis across ADF, Databricks, Snowflake, and Splunk, with runbooks so fixes don't depend on one person's memory.

Works With the Business

Translating reporting requirements into source mappings and validation rules alongside analysts, QA, and support teams.

Interested in discussing data engineering opportunities?

I'm actively looking for Data Engineer roles. Let's talk about how I can contribute to your team.

Contact

Let's connect.

I'm actively open to Data Engineer opportunities. If my background aligns with your team's needs, I'd love to hear from you.

Location
United States

Send a message