Bridging into Python Ecosystem with Cloud-Native Distributed Machine Learning Pipelines

Speaker Card

Abstract:

In recent years, advances in machine learning have made tremendous progress yet large scale machine learning still remains challenging. With the variety of machine learning frameworks such as TensorFlow and PyTorch, it’s not easy to automate the process of training machine learning models on distributed Kubernetes clusters. Machine learning researchers and algorithm engineers with less or zero DevOps experience cannot easily launch, manage, monitor, and optimize distributed machine learning pipelines.

This talk presents various best practices and challenges on building large, efficient, scalable, and reliable distributed machine learning pipelines using cloud-native technologies such as Argo Workflows and Kubeflow as well as how they fit into Python ecosystem with cutting-edge distributed machine learning frameworks such as TensorFlow and PyTorch.

Slides
Schedule
Video
Code: Argo Workflows examples in Python and in YAML.
3,893 registrants.

public-talks

Slides, videos, and supporting files for my public talks

Bridging into Python Ecosystem with Cloud-Native Distributed Machine Learning Pipelines