MLFlow
Senior DevOps Engineer with a strong background in CICD and Observability and Monitoring and skilled in tools like Elasticsearch, Docker, Kubernetes,Terraform, and Ansible. I focus on automating using DevOps tools or scripting using shell and python.
Experiment Tracking
Process of tracking everything that data scientist do as a part of model training.tracking helps to capture parameters used for each model and compare them to decide the best model.
| Round1 | Model = output of (Randomforest ,CSV) | 67% efficacy |
| Round2 | Model = output of (Logical regression ,CSV) | 79% efficacy |
| Round3 | Model = output of (XGBoot ,CSV) | 87% efficacy |
| ….. | ||
| Round100 | Model = output of (Alg ,CSV) | 81% efficacy |
Tracking Parameters:
| parameters | Learning rate |
| code version | Git version of code,alg,csv |
| dataset version | dvc |
| metrics | accuracy of the model |
| artifacts | model file |
| system information etc | windows,linux |
MLFlow
Widely used platform for experiment tracking
Versioning model
Deploying model
| Excel sheet | MLFlow |
| we might forget to add few experiments | Automates tracking experiments |
| Lost track of most important experiment | Provides UI to compare experiments |
| Does not follow standardization | Standardizes the tracking |
| Stores artifacts,csvs,runs,model information | |
| Python program invoke MLFlow module | |
| SSO can be added to make it secure |
MLOps Engineer install/ setup MLFlow
Data scientists does instrumentation setup python MLFlow module,Connect, record parameters for experiment tracking
They can also do that in the MLflow platform to make it very secure.