Speaker
Piyush Panchal
Description
The single-node, single-GPU implementation of the FDD algorithm surpassed Presto's CPU based implementation by several orders of magnitude. We extend it to a multi-node, multi-GPU implementation using MPI and the overlap-add technique. It enables faster processing of long timeseries data, too big for single compute nodes, by leveraging modern multi-GPU compute nodes and distributed filesystems for high I/O throughput.