Data Science2024

Large-Scale Data Analysis and Machine Learning Lifecycle

Built distributed processing pipelines to analyze customer reviews, influencer behavior, and sentiment across regions and cuisines.

Methodology

Distributed pipelines in Apache Spark with missing-value treatment, time-series resampling, polynomial feature transforms, and MLflow-tracked hyperparameter optimization.

Findings

Delivered models for star-rating prediction and wind power forecasting, evaluated on accuracy, F1, and regression metrics with full experiment reproducibility.

Tools & Methods

Apache SparkMLflowNLPFeature Engineering
Fine-Tuning BERT for Domain-Specific Named Entity RecognitionModular Data Science Pipeline with DVC & Dagger