Quickstart
First, ensure that mpiarray is installed.
A distributed array
Section titled “A distributed array”import numpy as np
import mpiarray as mpa
a = mpa.array([1, 2, 3, 4])print(a) # this rank's block; printing never communicates
# a 64 x 48 grid split along x, one halo layer, periodic in xrho = mpa.zeros((64, 48), halo=1, periodic=(True, False))rank = rho.layout.rankrho.local[...] = rank + 1.0 # write this rank's blockrho.update_halos() # halo cells <- neighbours' boundary values
total = rho.sum() # collective: the same value on every rankfull = rho.gather() # collective: the whole array on every ranksines = np.sin(a).gather() # collective, so outside the `if`if rank == 0: print(total, full.shape, sines)Save it as example.py and run it on two ranks:
mpiexec -n 2 python example.pyRank 0 prints its block [1, 2], rank 1 prints [3, 4]. The same script also runs
serially (python example.py), as one rank holding the whole array: without an MPI
launcher mpiarray uses maybempi’s stand-in for mpi4py.MPI and never starts MPI.
What to remember
Section titled “What to remember”- Arrays are split along their first axis by default. Pass
split=(0, 1)to split over two axes, orsplit=Noneto give every rank the whole array. - Elementwise arithmetic and
a.localnever communicate. Reductions,gather(),a[i, j]and halo updates are collective: every rank must call them, also when only rank 0 needs the result. update_halos()copies the neighbours’ boundary values into the halo cells, as needed before a stencil;accumulate_halos()adds the halo cells into the neighbours, as needed after depositing particles.
GPU arrays
Section titled “GPU arrays”Select CuPy with CUNUMPY_BACKEND=cupy (or xp.set_backend("cupy")), Device buffers go through host memory
unless you tell cunumpy that the MPI library can take them:
import cunumpy as xpfrom maybempi import MPI # mpi4py.MPI under mpiexec, a serial stand-in otherwise
xp.mpi.mpi_is_cuda_aware(MPI.COMM_WORLD) # collective probe, or:xp.mpi.set_mpi_cuda_aware(True) # MPI takes device buffers (default: through the host)Next steps
Section titled “Next steps”- Distributed arrays: creating arrays, storage, halo cells and gathering.
- Layouts: how arrays are split and who neighbours whom.
- Halo exchange:
update_halosfor stencils,accumulate_halosfor deposition. - Operations and reductions: arithmetic, indexing, reductions, and which calls are collective.
- GPU arrays: CuPy and CUDA-aware MPI.
- The API reference for every class and function.