
Mayank G
Data Engineer for Databricks PySpark and ETL Pipelines
Compétences

Voir mes services


Portfolio
Expérience professionnelle
Software Developer
Reliance Retail • Temps plein
Jul 2025 - Present • 1 yr 2 mos
Led data pipelines for e-commerce inventory, returns and customer analytics, from CDC ingestion to Tableau reporting. • Reduced pipeline latency from 6+ hours to under 10 minutes. • Processed about 22M rows per day using Databricks, PySpark, Spark SQL and Delta Lake. • Built Oracle GoldenGate and Kafka ingestion with data-quality checks across Bronze and Silver layers. • Optimized incremental loads, reducing compute costs by about 18% while sustaining a 99%+ daily job-success rate. • Developed Gold-layer KPI datasets and improved documentation and governance across 400+ catalogued tables. • Enabled self-service analytics through Unity Catalog metadata and natural-language querying via Databricks MCP. Core tools: Databricks, PySpark, Spark SQL, Delta Lake, Kafka, Oracle GoldenGate, Unity Catalog, Tableau and AWS S3.
Data Science
Celebal Technologies • Temps plein
Oct 2024 - Aug 2025 • 10 mos
Built customer analytics and applied AI solutions for business teams. • Defined churn logic and behavioral features for TE Connectivity using Databricks and PySpark. • Conducted exploratory analysis and feature engineering, and trained Random Forest models on AWS SageMaker. • Built a Customer Data Platform with FastAPI REST APIs and Neo4j. • Deployed a knowledge-graph chatbot using Amazon Bedrock to answer business questions. Core tools: Databricks, PySpark, AWS SageMaker, FastAPI, Neo4j and Amazon Bedrock.