Skip to content

Quickstart

First, ensure that mpiarray is installed.

import numpy as np
import mpiarray as mpa
a = mpa.array([1, 2, 3, 4])
print(a) # this rank's block; printing never communicates
# a 64 x 48 grid split along x, one halo layer, periodic in x
rho = mpa.zeros((64, 48), halo=1, periodic=(True, False))
rank = rho.layout.rank
rho.local[...] = rank + 1.0 # write this rank's block
rho.update_halos() # halo cells <- neighbours' boundary values
total = rho.sum() # collective: the same value on every rank
full = rho.gather() # collective: the whole array on every rank
sines = np.sin(a).gather() # collective, so outside the `if`
if rank == 0:
print(total, full.shape, sines)

Save it as example.py and run it on two ranks:

Terminal window
mpiexec -n 2 python example.py

Rank 0 prints its block [1, 2], rank 1 prints [3, 4]. The same script also runs serially (python example.py), as one rank holding the whole array: without an MPI launcher mpiarray uses maybempi’s stand-in for mpi4py.MPI and never starts MPI.

  • Arrays are split along their first axis by default. Pass split=(0, 1) to split over two axes, or split=None to give every rank the whole array.
  • Elementwise arithmetic and a.local never communicate. Reductions, gather(), a[i, j] and halo updates are collective: every rank must call them, also when only rank 0 needs the result.
  • update_halos() copies the neighbours’ boundary values into the halo cells, as needed before a stencil; accumulate_halos() adds the halo cells into the neighbours, as needed after depositing particles.

Select CuPy with CUNUMPY_BACKEND=cupy (or xp.set_backend("cupy")), Device buffers go through host memory unless you tell cunumpy that the MPI library can take them:

import cunumpy as xp
from maybempi import MPI # mpi4py.MPI under mpiexec, a serial stand-in otherwise
xp.mpi.mpi_is_cuda_aware(MPI.COMM_WORLD) # collective probe, or:
xp.mpi.set_mpi_cuda_aware(True) # MPI takes device buffers (default: through the host)