Skip to content

Post small RDMA messages inline - #3569

Open
alxrxs wants to merge 6 commits into
apache:masterfrom
alxrxs:rdma-inline-small-messages
Open

alxrxs wants to merge 6 commits into
apache:masterfrom
alxrxs:rdma-inline-small-messages

Conversation

@alxrxs

@alxrxs alxrxs commented Sep 28, 2026 •

Copy link
Copy Markdown

What problem does this PR solve?

Issue Number: none

Problem Summary:

RdmaEndpoint::CutFromIOBufList() posts every message by reference, so for a small RPC request or response the NIC first has to DMA-read the data from host memory. The QP requests no inline data (MAX_INLINE_DATA was never used and was commented out in #2876).

What is changed and the side effects?

Changed:

  • AllocateQp() requests 236 bytes of inline data, less if the device refuses (see below), and keeps the granted size in RdmaResource.
  • CutFromIOBufList() sets IBV_SEND_INLINE when a message fits that size and all of its blocks come from the RDMA block pool. User-registered memory is never inlined, since it may be GPU memory.

236 bytes is the largest Send whose inlined WQE fits the 256-byte BlueFlame buffer of current mlx5 NICs. Larger messages are posted as before. The 236-byte limit applies only on Mellanox/NVIDIA NICs (vendor ID 0x02c9); other devices inline up to the size they grant. The verbs API cannot report the inline limit, so devices whose driver has a fixed one try it first, by vendor ID (Intel 216, 101 or 48 depending on the irdma generation, 101 on the E810 and 48 on the X722; Alibaba eRDMA 96), and a device that refuses a size is asked for 16 bytes less each time until it accepts, down to none.

Side effects:

  • Performance effects: removes a host memory read before each small message is sent. In a separate ping-pong microbenchmark that sends the same small RDMA write with and without IBV_SEND_INLINE, inlining cut one-way latency by about 0.43 us on Xeon 8480C hosts with ConnectX-7 and 0.89 us on EPYC 7742 hosts with ConnectX-6. On brpc's own rdma_performance example (Xeon 8360Y, ConnectX-6 Dx, RoCEv2), an echo RPC took about 1 us less when both messages were inlined (29 us before); the numbers are in a comment below. With the default rdma_max_sge the send WQE does not grow, but with a small rdma_max_sge it can.

  • Breaking backward compatibility: no.


Check List:

Tested on soft-RoCE (rxe) with the rdma_performance server and an echo client: attachments from 0 to 300 KB compared byte for byte, messages up to the 512 bytes rxe grants went inline (rxe is not a Mellanox device) and larger ones by reference, and the same again with a test hook that made QP creation refuse inline sizes above a device limit (216, 101 or 48 with an Intel vendor ID, 96 with Alibaba's, 101 with rxe's own), where every QP got the expected size (216, 101, 48, 96 and 92 bytes) and messages up to it went inline. Unit tests in brpc_rdma_unittest cover the size selection and the retry loop with a stubbed IbvCreateQp, including a non-EINVAL failure that stops at once, and the whole test binary passes (88 tests, without RDMA hardware). libbrpc.a builds with WITH_RDMA=ON.

Generated with Claude Code

CutFromIOBufList() posts every message by reference, so the NIC has to
read a small RPC from host memory before sending it. Request 236 bytes
of inline data at QP creation, retrying without it if the device
refuses, and inline a message that fits the granted size and comes
entirely from the RDMA block pool. 236 bytes is the largest Send whose
WQE fits the 256-byte BlueFlame buffer of current mlx5 NICs.

Generated-by: Claude Code (Claude Opus 5.5)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@alxrxs
alxrxs force-pushed the rdma-inline-small-messages branch from 91efa9b to 7da4d18 Compare October 2, 2026 12:56
@alxrxs
alxrxs marked this pull request as ready for review October 2, 2026 13:57
@chenBright
chenBright requested a balanced review from Copilot October 3, 2026 03:54

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The new inline-selection and QP-fallback paths need automated regression coverage.

Review effort: Balanced
Findings: 2 Low severity

Open (2)
What changed in this PR

Adds inline posting for small RDMA messages to reduce DMA-read latency.

Changes:

  • Requests up to 236 bytes of inline QP capacity with fallback.
  • Inlines eligible pool-backed messages while excluding user-registered memory.
File Description
src/​brpc/​rdma/​rdma_endpoint.h Stores granted inline capacity.
src/​brpc/​rdma/​rdma_endpoint.cpp Negotiates capacity and selects inline sends.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/brpc/rdma/rdma_endpoint.cpp Outdated
Comment on lines +907 to +908
if (in_pool && this_len <= _resource->max_inline_data) {
wr.send_flags |= IBV_SEND_INLINE;

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added a test for the send path's inline decision in 8e40af7 (should_post_inline): a pool-backed message at the granted size is inlined, one byte more is not, and user-registered memory never is.

Comment thread src/brpc/rdma/rdma_endpoint.cpp Outdated
Comment on lines +1157 to +1160
if (qp == nullptr) {
// The device may not support inline data, try again without it
attr.cap.max_inline_data = 0;
qp = IbvCreateQp(GetRdmaPd(), &attr);

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Covered by the tests added in 9b2a7c8 (create_qp_with_inline_data).

alxrxs and others added 2 commits October 4, 2026 00:30
The 236-byte limit keeps an inlined Send within the 256-byte BlueFlame
buffer of mlx5 NICs. Apply it only on Mellanox/NVIDIA devices (vendor
ID 0x02c9, kept from the device query GlobalRdmaInitializeOrDie already
does); on other devices the inline size the QP was granted is the limit.

If QP creation with 236 bytes of inline data fails, try 64 bytes before
falling back to none. irdma (Intel E810) rejects requests above 101
bytes, so on those NICs nothing was inlined.

Generated-by: Claude Code (Claude Opus 5.5)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After a refused 236-byte inline request, QP creation was retried with 64
bytes, then with none. Devices take less than 236 by different amounts:
Intel's irdma refuses more than 216, 101 or 48 bytes depending on the
generation (101 on an E810, 48 on an X722), Alibaba's erdma more than
96. With 64 as the only middle step, an E810 or erdma QP got 64 bytes, a
newer irdma QP 64 instead of 216, and an X722 QP none.

The verbs API cannot report the limit, so step down to it. A device
whose kernel driver has a fixed limit, found by the vendor ID kept from
the device query GlobalRdmaInitializeOrDie already does, starts there:
216, then 101, then 48 on Intel, 96 on Alibaba, each capped at 236.
Other devices start at 236. Each refused size is followed by the next
one 16 bytes smaller, down to 0, so a failure unrelated to inline data
still ends as before, after a few more attempts at QP creation. Devices
that accept 236 (mlx5, rxe) are unchanged. libfabric's verbs provider
also probes the limit by trial QP creation (vrb_find_max_inline()).

Generated-by: Claude Code (Claude Opus 5.5)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@alxrxs

alxrxs commented Oct 4, 2026

Copy link
Copy Markdown
Author

Replaced the 64-byte fallback: devices whose driver has a fixed inline limit now start there, by vendor ID (Intel 216, 101 or 48 depending on the irdma generation, 101 on the E810 and 48 on the X722; Alibaba eRDMA 96), and anything else that refuses steps down 16 bytes at a time until it is accepted. The verbs API can't report the limit; libfabric's verbs provider probes it too (vrb_find_max_inline). mlx5 and rxe accept 236 on the first try, so nothing changes there.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The hardware-dependent negotiation logic lacks required automated regression coverage.

Review effort: Balanced
Findings: 1 Medium severity · 2 Low severity

Open (3)

Comment thread src/brpc/rdma/rdma_endpoint.cpp Outdated
Comment on lines +1162 to +1167
static uint32_t InlineDataToRequest(uint32_t vendor_id, uint32_t size,
bool refused) {
for (const auto& known : KNOWN_INLINE_LIMITS) {
if (known.vendor_id == vendor_id &&
(!refused || known.max_inline < size)) {
return std::min(known.max_inline, size);

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added unit tests for the size selection and the retry loop, with a stubbed IbvCreateQp.

@wwbmmm wwbmmm left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[replied by brpc-oncall robot] The inline-send path and the QP inline-capability fallback look correct, and the eligibility check conservatively excludes user-registered memory (which may be device memory). One robustness suggestion on the QP retry loop below.

Comment thread src/brpc/rdma/rdma_endpoint.cpp Outdated
attr.qp_type = IBV_QPT_RC;
return IbvCreateQp(GetRdmaPd(), &attr);
ibv_qp* qp = IbvCreateQp(GetRdmaPd(), &attr);
while (qp == nullptr && inline_size > 0) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[replied by brpc-oncall robot] This loop steps down and retries on any IbvCreateQp failure, not only on failures caused by the requested inline size. On a device that rejects every inline size (or that fails for an unrelated reason such as ENOMEM or reaching the device QP limit), AllocateQpCq now performs up to 16 failed ibv_create_qp calls per resource, and the PLOG(WARNING) << "Fail to create QP" in the caller reports the errno of the last (0-inline) attempt rather than the root cause. Consider logging each step-down (e.g. LOG_IF(INFO, FLAGS_rdma_trace_verbose), following the convention used elsewhere in this file) and stepping down only when the failure can plausibly come from the inline request (e.g. errno == EINVAL).

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point. It now steps down only when ibv_create_qp fails with EINVAL, logs each step under FLAGS_rdma_trace_verbose, and any other failure is reported as before.

AllocateQp() retried with a smaller inline size after any ibv_create_qp
failure, so a failure unrelated to inline data, such as ENOMEM or the
device's QP limit, led to up to 16 failed calls, and the caller's PLOG
reported the errno of the last attempt. Step down only when the failed
call left errno == EINVAL, which is how providers refuse an inline size;
any other failure is returned at once with its errno. Each step is
logged under FLAGS_rdma_trace_verbose.

QP creation moves into CreateQpWithInlineData(), which takes the vendor
ID as a parameter. It and InlineDataToRequest() are no longer static, so
that brpc_rdma_unittest can test them with a stubbed IbvCreateQp: the
first request per vendor, the Mellanox cap at 236 bytes, Intel 216, 101
and 48, Alibaba 96, an unknown device stepping down to 92, a non-EINVAL
failure stopping at once, and a device refusing every size ending at 0.

Generated-by: Claude Code (Claude Opus 5.5)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The safety-critical send-path eligibility logic lacks focused automated coverage.

Review effort: Balanced
Findings: 1 Medium severity · 2 Low severity

Open (3)

CutFromIOBufList() inlines a message only when it fits the inline size
the QP was granted and all of its blocks come from the block pool; user
registered memory may be device memory, which the CPU cannot copy from.
Move that decision into ShouldPostInline(), exposed for UT, and test it
in brpc_rdma_unittest: a pool-backed message at the granted size is
inlined, one byte more is not, and a message in user registered memory
never is, however small.

CutFromIOBufList() itself needs an initialized RDMA device: it returns
early in the unit test binary, sizes its SGE list from the device and
looks blocks up in the registered block pool. So the decision is tested
on its own; the send path is otherwise unchanged.

Generated-by: Claude Code (Claude Opus 5.5)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@alxrxs

alxrxs commented Oct 6, 2026

Copy link
Copy Markdown
Author

I've now measured this on hardware with brpc's own rdma_performance example.

Setup: two hosts with Intel Xeon Platinum 8360Y and ConnectX-6 Dx (RoCEv2, 100 Gb/s), Ubuntu 24.04, kernel 6.8, rdma-core 50. Server and client each pinned to one core on the NIC's NUMA node, one thread, depth 1, echo on. Base is this PR's merge base (535d8f8), head is this PR (8e40af7). 10 rounds per size, base and head alternating. I took the time per RPC from the client's printed throughput (attachment bytes / MB/s), because Avg-Latency is printed in whole microseconds.

The sizes below are the lengths of the messages as posted. On these NICs the head inlines a message when it is 236 bytes or less. I checked that with a counting build: every send of 236 bytes or less went inline, and none above.

request / response (bytes) inlined with this PR base (us per RPC) change with this PR (us), 95% interval
64 / 38 both 29.03 -1.12 [-1.31, -0.93]
128 / 102 both 29.21 -0.98 [-1.19, -0.77]
200 / 174 both 29.33 -1.03 [-1.15, -0.91]
232 / 206 both 29.41 -0.99 [-1.13, -0.86]
241 / 215 response only 29.38 -0.53 [-0.68, -0.37]
512 / 486 neither 29.73 -0.04 [-0.19, +0.11]
267 / 241 neither 29.37 +0.08 [-0.00, +0.16]

So when both messages are inlined, an echo RPC takes about 1 us less (about 3.5%). With only the response inlined it's about half that, and when neither message fits there's no change I can resolve.

Comment thread src/brpc/rdma/rdma_endpoint.cpp
All QPs are created on the same device with the same attributes apart
from their queue sizes, so a device that refuses the first inline size
made every one of the rdma_prepared_qp_cnt QPs repeat the same step-down
(216, 101, 48 on an Intel device limited to 48 bytes). Remember the size
the last QP was created with and start there. If that size is refused,
the EINVAL step-down continues from it as before. Each QP still takes its
own granted size from attr.cap.max_inline_data.

Generated-by: Claude Code (Claude Opus 5.5)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Cache the effective granted inline size to avoid repeated oversized QP-creation retries.

Review effort: Lite
Findings: 2 Medium severity · 2 Low severity

Open (4)

Comment on lines +1210 to +1216
if (qp != nullptr) {
g_inline_data_request.store(inline_size, butil::memory_order_relaxed);
// ibv_create_qp writes the granted inline data size back into attr
*max_inline_data = attr->cap.max_inline_data;
if (vendor_id == MELLANOX_VENDOR_ID) {
*max_inline_data = std::min(*max_inline_data, BF_MAX_INLINE_DATA);
}

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is on purpose. ibv_create_qp() either fails or grants at least what was asked: ibv_create_qp(3) says "the values will be greater than or equal to the values requested". So a size that was accepted once is accepted again on the next QP with the same attributes, and the step-down doesn't repeat. Caching the request keeps every QP asking for exactly what the first one asked for. Each QP still reads its own grant from attr.cap.max_inline_data, and the Mellanox cap is applied to that.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants