AI Data Engineer with 2 years of experience building data pipelines, preparing AI-ready datasets, developing feature engineering workflows, and supporting machine learning applications. Skilled in Python, Apache Spark, Databricks, Azure Data Factory, SQL, Delta Lake, and Azure Machine Learning. Experienced in data ingestion, ETL/ELT, data validation, model data preparation, and scalable cloud data solutions.
Expertise & skills
Selected projects
AI Feature Store Pipeline
Developed a centralized feature engineering pipeline to prepare reusable datasets for machine learning models.
- Built ETL workflows using PySpark.
- Created reusable feature engineering pipelines.
- Validated data quality and consistency.
- Stored features in Delta Lake.
- Automated scheduled data processing.
Customer Behavior Data Platform
Designed a scalable platform to process customer interaction data for AI model training.
- Integrated multiple data sources.
- Transformed raw datasets into analytics-ready tables.
- Optimized Spark jobs.
- Implemented data validation rules.
- Supported downstream ML teams.
Real-Time Sales Analytics Pipeline
Built a streaming-ready pipeline for sales analytics and forecasting.
- Processed transactional datasets.
- Developed aggregation workflows.
- Published cleaned datasets.
- Improved pipeline performance.
- Generated reporting datasets.
Data Quality Monitoring Framework
Created an automated framework to monitor AI training data quality.
- Implemented validation checks.
- Detected missing and duplicate records.
- Generated quality reports.
- Configured automated alerts.
- Documented governance procedures.
Ways of working
Certifications
- Microsoft Certified: Azure Data Fundamentals (DP-900)
- Microsoft Certified: Azure AI Fundamentals (AI-900)
- Databricks Fundamentals
- Python for Data Engineering
Education
- 2024
Master of Technology (M.Tech.) – Artificial Intelligence
Gujarat Technological University (GTU)
- 2022
Bachelor of Engineering (B.E.) – Computer Engineering
Gujarat Technological University (GTU)
Achievements
- Built scalable AI data pipelines for machine learning workloads.
- Improved dataset quality through automated validation.
- Optimized Spark processing for faster ETL execution.
- Delivered AI-ready datasets for predictive modeling.
Similar engineers
Other AI / ML engineers with comparable experience.



