paper-with-me

홈 › Papers

Training Kinetics in 15 Minutes: Large-scale Distributed Training on Videos

2019-10-01 · Ji Lin, Chuang Gan, Song Han

Deep video recognition is more computationally expensive than image recognition, especially on large-scale datasets like Kinetics [1]. Therefore, training scalability is essential to handle a large amount of videos. In this paper, we study the factors that impact the training scalability of video networks. We recognize three bottlenecks, including data loading (data movement from disk to GPU), communication (data movement over networking), and computation FLOPs. We propose three design guidelines to improve the scalability: (1) fewer FLOPs and hardware-friendly operator to increase the computation efficiency; (2) fewer input frames to reduce the data movement and increase the data loading efficiency; (3) smaller model size to reduce the networking traffic and increase the networking efficiency. With these guidelines, we designed a new operator Temporal Shift Module (TSM) that is efficient and scalable for distributed training. TSM model can achieve 1.8x higher throughput compared to previous I3D models. We scale up the training of the TSM model to 1,536 GPUs, with a mini-batch of 12,288 video clips/98,304 images, without losing the accuracy. With such hardware-aware model design, we are able to scale up the training on Summit supercomputer and reduce the training time on Kinetics dataset from 49 hours 55 minutes to 14 minutes 13 seconds, achieving a top-1 accuracy of 74.0%, which is 1.6x and 2.9x faster than previous 3D video models with higher accuracy. The code and more details can be found here: http://tsm-hanlab.mit.edu.

📄 PDF Abstract BibTeX arXiv:1910.00932

Code (1)

MIT-HAN-LAB/temporal-shift-module pytorch

Tasks

GPUVideo Recognition

Similar Papers 제목 키워드 기반

Distributed Deep Reinforcement Learning: Learn how to play Atari games in 21 minutes

2018-01-09 · Igor Adamski, Robert Adamski, Tomasz Grel, Adam Jędrych 외

We present a study in Distributed Deep Reinforcement Learning (DDRL) focused on scalability of a state-of-the-art Deep Reinforcement Learning algorithm known as Batch Asynchronous Advantage ActorCritic (BA3C). We show th…

Atari GamesCPUDeep Reinforcement LearningPlaying the Game of 2048+3

Echo: Simulating Distributed Training At Scale

2024-12-17 · Yicheng Feng, Yuetao Chen, Kaiwen Chen, Jingzong Li 외

Simulation offers unique values for both enumeration and extrapolation purposes, and is becoming increasingly important for managing the massive machine learning (ML) clusters and large-scale distributed training jobs. I…

GPU

Breaking Boundaries: Distributed Domain Decomposition with Scalable Physics-Informed Neural PDE Solvers

2023-08-28 · Arthur Feeney, Zitong Li, Ramin Bostanabad, Aparna Chandramowlishwaran

Mosaic Flow is a novel domain decomposition method designed to scale physics-informed neural PDE solvers to large domains. Its unique approach leverages pre-trained networks on small domains to solve partial differential…

scientific discovery

Asynchronous Training of Word Embeddings for Large Text Corpora

2018-12-07 · Avishek Anand, Megha Khosla, Jaspreet Singh, Jan-Hendrik Zab 외

Word embeddings are a powerful approach for analyzing language and have been widely popular in numerous tasks in information retrieval and text mining. Training embeddings over huge corpora is computationally expensive b…

Information RetrievalRetrievalWord Embeddings

Image Classification at Supercomputer Scale

2018-11-16 · Chris Ying, Sameer Kumar, Dehao Chen, Tao Wang 외

Deep learning is extremely computationally intensive, and hardware vendors have responded by building faster accelerators in large clusters. Training deep learning models at petaFLOPS scale requires overcoming both algor…

ClassificationDeep LearningGeneral Classificationimage-classification+1