Deep Gradient Compression for Distributed Training with Song Han

EPISODE 146

May 31, 2018

LISTEN

Banner Image: Song Han - Podcast Interview

Join our list for notifications and early access to events

About this Episode

On today's show I chat with Song Han, assistant professor in MIT's EECS department, about his research on Deep Gradient Compression.

In our conversation, we explore the challenge of distributed training for deep neural networks and the idea of compressing the gradient exchange to allow it to be done more efficiently. Song details the evolution of distributed training systems based on this idea, and provides a few examples of centralized and decentralized distributed training architectures such as Uber's Horovod, as well as the approaches native to Pytorch and Tensorflow. Song also addresses potential issues that arise when considering distributed training, such as loss of accuracy and generalizability, and much more.

Connect with Song

Resources

Deep Gradient Compression
Horovod
Scaling Machine Learning at Uber with Mike Del Balso - Talk #115
Stochastic Gradient Descent
Data Parallelism
Exploding Gradients Explained
AlexNet
TWIML Presents: Series page

Deep Gradient Compression for Distributed Training with Song Han

About this Episode

Connect with Song

Resources

More from TWIML

Leave a Reply Cancel reply