Conversation
Implement the notified communication interface (MPI Standard section 12.6) in the osc/rdma component, which covers the libfabric/ofi BTL path: - MPI_PUT_NOTIFY / MPI_GET_NOTIFY, incrementing the target's notification counter with a btl atomic after the data movement has completed there - MPI_WIN_GET_NOTIFY_VALUE / MPI_WIN_RESET_NOTIFY_VALUE, using a btl compare-and-swap when it is not safe to mix CPU and btl atomics - MPI_WIN_SET_NUM_NOTIFY / MPI_WIN_GET_NUM_NOTIFY, publishing each rank's attached counter count to the group with an allgather A fixed region of OMPI_OSC_RDMA_NOTIFY_MAX counters is allocated in the window state so origins can update them with btl atomics; only the first notify_counts[rank] are considered attached. Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Add a btl interface for counters that the network adapter increments at the target once a remote operation completes there, and use it in osc/rdma so a notified put or get costs a single network operation instead of a data transfer followed by an atomic to announce it. opal/mca/btl: - MCA_BTL_FLAGS_NOTIFIED_RMA, set by a btl that provides btl_register_notification, btl_deregister_notification, btl_notification_read and btl_notification_wait - registering a region this way yields an ordinary registration handle alongside the counter, so an origin targets that handle in a put or get and the target's adapter counts the operation as a side effect - counters are monotonic and never reset by the btl; a consumer that needs a reset records a base value and subtracts it opal/mca/btl/ofi: - a notification is an fi_cntr bound to a memory region with FI_REMOTE_WRITE | FI_REMOTE_READ, so it requires the provider to have been opened with FI_RMA_EVENT - FI_RMA_EVENT is requested as an optional capability: fi_getinfo treats requested caps as mandatory, so the query is retried without it before any other fallback, since counters are an optimization and should be the first thing given up - these registrations bypass the rcache, which keys on the address range and would hand back one registration where several distinct counters over the same buffer are wanted - btl_ofi_enable_rma_event MCA parameter to suppress the request ompi/mca/osc/rdma: - each rank registers its window base once per attached notification index and publishes the resulting handles to the group; an origin picks a counter by choosing which handle to target - the accelerated path is taken only when the transfer maps to exactly one btl operation, since the counter counts operations and the standard requires exactly one increment per notified call; local peers, non-contiguous datatypes, size mismatches, transfers over the put/get limit and unaligned gets fall back to the atomic path - MPI_WIN_RESET_NOTIFY_VALUE records the raw counter value as a base rather than resetting it, which avoids racing increments in flight Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Request FI_RMA_EVENT only from the selected provider, register with FI_RMA_EVENT, and bind FI_REMOTE_WRITE only (reads are not counted). Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Agree on errors before set_num_notify changes state, report bounds, make reset safe against btl atomics, and use atomics for get_notify. Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Also unlink the grown counter segment when an attach fails. Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Title: osc/rdma, btl/ofi: add MPI-5.1 notified RMA support
Summary
This PR implements part of MPI-5.1 notified RMA (§12.6) in
osc/rdma:MPI_Put_notify,MPI_Get_notify, and the counter-management calls(
MPI_Win_set_num_notify,MPI_Win_get_num_notify,MPI_Win_get_notify_value,MPI_Win_reset_notify_value, and theMPI_WIN_NOTIFICATION_*window attributes).It also adds an optional hardware notification counter interface to the
BTL framework, with an implementation in
btl/ofibuilt on libfabricfi_cntrcounters bound to memory regions. When every rank can use it,osc/rdmasends a notified put as a single network operation, and the NICincrements the counter.
Design
osc/rdmahas two ways to deliver a notification:Portable (atomic) path, always available. Normal put or get, wait
for it to complete, then send a BTL atomic
+1to a counter in thetarget's window state. It works with any BTL that supports atomics.
Hardware-counter path (btl/ofi only). The target registers its window
memory once for each notification index. Each registration has its own
fi_cntr, and the NIC increments it when a remote write to that regioncompletes. The origin selects the counter by writing through that
registration's rkey. That is one operation, and the origin does not wait
between the data and the notification.
Key decisions:
Counters live in the existing window state region. That region is
already registered and exchanged with every peer, so the atomic path needs
no new registration. The size is fixed at
OMPI_OSC_RDMA_NOTIFY_MAX = 16,so
NUM_SB = NUM_UB = 16. Unlikeosc/smandosc/ucx, the counters donot grow on demand.
The hardware path is used only when a transfer is exactly one BTL put
through the NIC. The NIC counts operations, but each notified call must
produce exactly one increment. Local (load/store) peers, non-contiguous
datatypes, and transfers larger than
put_limitfall back to the atomicpath.
MPI_Win_get_notify_valuereturns the sum of the hardware counterand the state counter.
Gets always use the atomic path. libfabric (fi_mr(3)) defines only
FI_REMOTE_WRITEas a counter event for memory regions, so remote readsare never counted.
Whether to use hardware counters is decided by the whole group. The
origin chooses the path, and the target reads with the matching method, so
all ranks must agree (allreduce(MIN)). If any registration fails on any
rank, the window quietly falls back to atomics.
Hardware counters are never zeroed. Zeroing a counter while writes
are in flight can lose increments. Reset records a base value and
subtracts it.
MPI_Win_set_num_notifyre-registers fresh counters, whichis safe because no epoch is open.
Collectives fail together.
MPI_Win_set_num_notifychecks for localerrors (bad count, open epoch, allocation failure) before changing any
state, and all ranks agree on the result with one allreduce. A rank that
returned early would leave the others stuck in the allgather.
Changes by area
OPAL: BTL interface (
opal/mca/btl/btl.h,btl_base_frame.c)MCA_BTL_FLAGS_NOTIFIED_RMA, also added to the flag enum sobtl_*_flagsandompi_infoshow it.mca_btl_base_notification_tand four function pointers:btl_register_notification,btl_deregister_notification,btl_notification_read, andbtl_notification_wait. Registration returnsan ordinary
mca_btl_base_registration_handle_t, sobtl_putworksunchanged.
mca_btl_base_module_tstays the same size for all other BTLs.are counted, and counters only ever increase.
opal/mca/btl/ofibtl_ofi_notification.c(new): opens anfi_cntrwithFI_CNTR_EVENTS_COMP, registers an MR withFI_RMA_EVENT, binds it withFI_REMOTE_WRITE, handlesFI_MR_ENDPOINTandFI_MR_RMA_EVENT(enableafter binding), and bypasses the rcache so each registration gets a
distinct rkey.
readchecksfi_cntr_readerrfirst, so a failed remotewrite is reported instead of leaving the target polling forever.
btl_ofi_component.c:FI_RMA_EVENTis a secondary capability, so addingit to the selection hints could change which provider is chosen. Instead,
the provider is selected exactly as before, then queried again with
FI_RMA_EVENTadded. The result is accepted only if the provider anddomain are the same. If that fails, the plain variant is used. The
feature is enabled only on modules that register memory.
btl_ofi_enable_rma_event(default true) to turn thefeature off.
ompi/mca/osc/rdmaosc_rdma_component.c: fills input_notify,get_notify,get/reset value, set/get num, and
get_notify_bounds.osc_rdma_comm.c:put_notify(hardware or atomic path),get_notify(atomic path only), and
notify_increment(CPU atomic for peers whereCPU and BTL atomics are coherent, otherwise a BTL atomic tracked by the
sync, so
MPI_Win_flushcovers the notification).osc_rdma_module.c:set_num_notify(agree on errors, reset, allgathercounts, rebuild hardware counters),
get_notify_value(progresses first,then
rmb),reset_notify_value(atomic read-and-zero of the statecounter, using a BTL cswap when CPU atomics are not safe), bounds, and
cleanup in
free.ompi/mca/osc/sm(separate commit)mpi_assert_max_num_notifyat window creation, and errors inset_num_notify. The error now travels through the existing collectives,and the reset happens only after all ranks agree. Rank 0 unlinks a grown
segment if attaching fails. This commit can be split into its own PR.