Skip to content

NOTIFIED_RMA_RDMA: libfabric implementation. - #33

Open
joe-explr wants to merge 5 commits into
devreal:notified-rmafrom
joe-explr:notified-rma-lib
Open

joe-explr wants to merge 5 commits into
devreal:notified-rmafrom
joe-explr:notified-rma-lib

Conversation

@joe-explr

@joe-explr joe-explr commented Sep 17, 2026 •

Copy link
Copy Markdown

Title: osc/rdma, btl/ofi: add MPI-5.1 notified RMA support

Summary

This PR implements part of MPI-5.1 notified RMA (§12.6) in osc/rdma:
MPI_Put_notify, MPI_Get_notify, and the counter-management calls
(MPI_Win_set_num_notify, MPI_Win_get_num_notify,
MPI_Win_get_notify_value, MPI_Win_reset_notify_value, and the
MPI_WIN_NOTIFICATION_* window attributes).

It also adds an optional hardware notification counter interface to the
BTL framework, with an implementation in btl/ofi built on libfabric
fi_cntr counters bound to memory regions. When every rank can use it,
osc/rdma sends a notified put as a single network operation, and the NIC
increments the counter.

Design

osc/rdma has two ways to deliver a notification:

  • Portable (atomic) path, always available. Normal put or get, wait
    for it to complete, then send a BTL atomic +1 to a counter in the
    target's window state. It works with any BTL that supports atomics.

  • Hardware-counter path (btl/ofi only). The target registers its window
    memory once for each notification index. Each registration has its own
    fi_cntr, and the NIC increments it when a remote write to that region
    completes. The origin selects the counter by writing through that
    registration's rkey. That is one operation, and the origin does not wait
    between the data and the notification.

Key decisions:

  • Counters live in the existing window state region. That region is
    already registered and exchanged with every peer, so the atomic path needs
    no new registration. The size is fixed at OMPI_OSC_RDMA_NOTIFY_MAX = 16,
    so NUM_SB = NUM_UB = 16. Unlike osc/sm and osc/ucx, the counters do
    not grow on demand.

  • The hardware path is used only when a transfer is exactly one BTL put
    through the NIC.
    The NIC counts operations, but each notified call must
    produce exactly one increment. Local (load/store) peers, non-contiguous
    datatypes, and transfers larger than put_limit fall back to the atomic
    path. MPI_Win_get_notify_value returns the sum of the hardware counter
    and the state counter.

  • Gets always use the atomic path. libfabric (fi_mr(3)) defines only
    FI_REMOTE_WRITE as a counter event for memory regions, so remote reads
    are never counted.

  • Whether to use hardware counters is decided by the whole group. The
    origin chooses the path, and the target reads with the matching method, so
    all ranks must agree (allreduce(MIN)). If any registration fails on any
    rank, the window quietly falls back to atomics.

  • Hardware counters are never zeroed. Zeroing a counter while writes
    are in flight can lose increments. Reset records a base value and
    subtracts it. MPI_Win_set_num_notify re-registers fresh counters, which
    is safe because no epoch is open.

  • Collectives fail together. MPI_Win_set_num_notify checks for local
    errors (bad count, open epoch, allocation failure) before changing any
    state, and all ranks agree on the result with one allreduce. A rank that
    returned early would leave the others stuck in the allgather.

Changes by area

OPAL: BTL interface (opal/mca/btl/btl.h, btl_base_frame.c)

  • New flag MCA_BTL_FLAGS_NOTIFIED_RMA, also added to the flag enum so
    btl_*_flags and ompi_info show it.
  • The opaque type mca_btl_base_notification_t and four function pointers:
    btl_register_notification, btl_deregister_notification,
    btl_notification_read, and btl_notification_wait. Registration returns
    an ordinary mca_btl_base_registration_handle_t, so btl_put works
    unchanged.
  • The new pointers are placed in the existing padding union, so
    mca_btl_base_module_t stays the same size for all other BTLs.
  • Documented contract: only operations that modify memory (puts, atomics)
    are counted, and counters only ever increase.

opal/mca/btl/ofi

  • btl_ofi_notification.c (new): opens an fi_cntr with
    FI_CNTR_EVENTS_COMP, registers an MR with FI_RMA_EVENT, binds it with
    FI_REMOTE_WRITE, handles FI_MR_ENDPOINT and FI_MR_RMA_EVENT (enable
    after binding), and bypasses the rcache so each registration gets a
    distinct rkey. read checks fi_cntr_readerr first, so a failed remote
    write is reported instead of leaving the target polling forever.
  • btl_ofi_component.c: FI_RMA_EVENT is a secondary capability, so adding
    it to the selection hints could change which provider is chosen. Instead,
    the provider is selected exactly as before, then queried again with
    FI_RMA_EVENT added. The result is accepted only if the provider and
    domain are the same. If that fails, the plain variant is used. The
    feature is enabled only on modules that register memory.
  • New MCA parameter btl_ofi_enable_rma_event (default true) to turn the
    feature off.

ompi/mca/osc/rdma

  • osc_rdma_component.c: fills in put_notify, get_notify,
    get/reset value, set/get num, and get_notify_bounds.
  • osc_rdma_comm.c: put_notify (hardware or atomic path), get_notify
    (atomic path only), and notify_increment (CPU atomic for peers where
    CPU and BTL atomics are coherent, otherwise a BTL atomic tracked by the
    sync, so MPI_Win_flush covers the notification).
  • osc_rdma_module.c: set_num_notify (agree on errors, reset, allgather
    counts, rebuild hardware counters), get_notify_value (progresses first,
    then rmb), reset_notify_value (atomic read-and-zero of the state
    counter, using a BTL cswap when CPU atomics are not safe), bounds, and
    cleanup in free.

ompi/mca/osc/sm (separate commit)

  • The same bug where one rank returns early from a collective: a malformed
    mpi_assert_max_num_notify at window creation, and errors in
    set_num_notify. The error now travels through the existing collectives,
    and the reset happens only after all ranks agree. Rank 0 unlinks a grown
    segment if attaching fails. This commit can be split into its own PR.

joe-explr and others added 5 commits September 16, 2026 22:21
Implement the notified communication interface (MPI Standard section
12.6) in the osc/rdma component, which covers the libfabric/ofi BTL
path:

 - MPI_PUT_NOTIFY / MPI_GET_NOTIFY, incrementing the target's
   notification counter with a btl atomic after the data movement has
   completed there
 - MPI_WIN_GET_NOTIFY_VALUE / MPI_WIN_RESET_NOTIFY_VALUE, using a btl
   compare-and-swap when it is not safe to mix CPU and btl atomics
 - MPI_WIN_SET_NUM_NOTIFY / MPI_WIN_GET_NUM_NOTIFY, publishing each
   rank's attached counter count to the group with an allgather

A fixed region of OMPI_OSC_RDMA_NOTIFY_MAX counters is allocated in the
window state so origins can update them with btl atomics; only the first
notify_counts[rank] are considered attached.

Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Add a btl interface for counters that the network adapter increments at
the target once a remote operation completes there, and use it in
osc/rdma so a notified put or get costs a single network operation
instead of a data transfer followed by an atomic to announce it.

opal/mca/btl:

 - MCA_BTL_FLAGS_NOTIFIED_RMA, set by a btl that provides
   btl_register_notification, btl_deregister_notification,
   btl_notification_read and btl_notification_wait
 - registering a region this way yields an ordinary registration handle
   alongside the counter, so an origin targets that handle in a put or
   get and the target's adapter counts the operation as a side effect
 - counters are monotonic and never reset by the btl; a consumer that
   needs a reset records a base value and subtracts it

opal/mca/btl/ofi:

 - a notification is an fi_cntr bound to a memory region with
   FI_REMOTE_WRITE | FI_REMOTE_READ, so it requires the provider to
   have been opened with FI_RMA_EVENT
 - FI_RMA_EVENT is requested as an optional capability: fi_getinfo
   treats requested caps as mandatory, so the query is retried without
   it before any other fallback, since counters are an optimization and
   should be the first thing given up
 - these registrations bypass the rcache, which keys on the address
   range and would hand back one registration where several distinct
   counters over the same buffer are wanted
 - btl_ofi_enable_rma_event MCA parameter to suppress the request

ompi/mca/osc/rdma:

 - each rank registers its window base once per attached notification
   index and publishes the resulting handles to the group; an origin
   picks a counter by choosing which handle to target
 - the accelerated path is taken only when the transfer maps to exactly
   one btl operation, since the counter counts operations and the
   standard requires exactly one increment per notified call; local
   peers, non-contiguous datatypes, size mismatches, transfers over the
   put/get limit and unaligned gets fall back to the atomic path
 - MPI_WIN_RESET_NOTIFY_VALUE records the raw counter value as a base
   rather than resetting it, which avoids racing increments in flight

Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Request FI_RMA_EVENT only from the selected provider, register with
FI_RMA_EVENT, and bind FI_REMOTE_WRITE only (reads are not counted).

Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Agree on errors before set_num_notify changes state, report bounds,
make reset safe against btl atomics, and use atomics for get_notify.

Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Also unlink the grown counter segment when an attach fails.

Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
@joe-explr
joe-explr changed the base branch from main to notified-rma September 17, 2026 17:56
@joe-explr joe-explr changed the title NOTIFIED_RMA_RDMA: Uses lib fabric implementation. NOTIFIED_RMA_RDMA: libfabric implementation. Sep 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant