Efficient broadcasts and simple algorithms for parallel linear algebra computing in clusters
Tinetti, Fernando Gustavo
Efficient broadcasts and simple algorithms for parallel linear algebra computing in clusters - ^p Datos electrónicos (1 archivo : 246 KB) .
Formato de archivo: PDF. -- Este documento es producción intelectual de la Facultad de Informática-UNLP (Colección BIPA / Biblioteca.) -- Disponible también en línea (Cons. 13/03/2009)
This paper presents a natural and efficient implementation for the classical broadcast message passing routine which optimizes performance of Ethernet based clusters. A simple algorithm for parallel matrix multiplication is specifically designed to take advantage of both, parallel computing facilities (CPUs) provided by clusters, and optimized performance of broadcast messages on Ethernet based clusters. Also, this simple parallel algorithm proposed for matrix multiplication takes into account the possibly heterogeneous computing hardware and maintains a balanced workload of computers according to their relative computing power. Performance tests are presented on a heterogeneous cluster as well as on a homogeneous cluster, where it is compared with the parallel matrix multiplication provided by the ScaLAPACK library. Another simple parallel algorithm is proposed for LU matrix factorization (a general method to solve dense systems of equations) following the same guidelines used for the parallel matrix multiplication algorithm. Some performance tests are presented over a homogeneous cluster.
DIF-M2650
PROCESAMIENTO PARALELO
ALGORITMOS PARALELOS
CLUSTERS
ÁLGEBRA LINEAL
REDES LOCALES
INTERCONEXIÓN DE REDES
COMUNICACIÓN DE DATOS
Efficient broadcasts and simple algorithms for parallel linear algebra computing in clusters - ^p Datos electrónicos (1 archivo : 246 KB) .
Formato de archivo: PDF. -- Este documento es producción intelectual de la Facultad de Informática-UNLP (Colección BIPA / Biblioteca.) -- Disponible también en línea (Cons. 13/03/2009)
This paper presents a natural and efficient implementation for the classical broadcast message passing routine which optimizes performance of Ethernet based clusters. A simple algorithm for parallel matrix multiplication is specifically designed to take advantage of both, parallel computing facilities (CPUs) provided by clusters, and optimized performance of broadcast messages on Ethernet based clusters. Also, this simple parallel algorithm proposed for matrix multiplication takes into account the possibly heterogeneous computing hardware and maintains a balanced workload of computers according to their relative computing power. Performance tests are presented on a heterogeneous cluster as well as on a homogeneous cluster, where it is compared with the parallel matrix multiplication provided by the ScaLAPACK library. Another simple parallel algorithm is proposed for LU matrix factorization (a general method to solve dense systems of equations) following the same guidelines used for the parallel matrix multiplication algorithm. Some performance tests are presented over a homogeneous cluster.
DIF-M2650
PROCESAMIENTO PARALELO
ALGORITMOS PARALELOS
CLUSTERS
ÁLGEBRA LINEAL
REDES LOCALES
INTERCONEXIÓN DE REDES
COMUNICACIÓN DE DATOS