Repository navigation
Conversation
CutFromIOBufList() posts every message by reference, so the NIC has to read a small RPC from host memory before sending it. Request 236 bytes of inline data at QP creation, retrying without it if the device refuses, and inline a message that fits the granted size and comes entirely from the RDMA block pool. 236 bytes is the largest Send whose WQE fits the 256-byte BlueFlame buffer of current mlx5 NICs. Generated-by: Claude Code (Claude Opus 5.5) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
91efa9b to
7da4d18
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The new inline-selection and QP-fallback paths need automated regression coverage.
Review effort: Balanced
Findings: 2
Open (2)
What changed in this PR
Adds inline posting for small RDMA messages to reduce DMA-read latency.
Changes:
- Requests up to 236 bytes of inline QP capacity with fallback.
- Inlines eligible pool-backed messages while excluding user-registered memory.
| File | Description |
|---|---|
src/brpc/rdma/rdma_endpoint.h |
Stores granted inline capacity. |
src/brpc/rdma/rdma_endpoint.cpp |
Negotiates capacity and selects inline sends. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| if (in_pool && this_len <= _resource->max_inline_data) { | ||
| wr.send_flags |= IBV_SEND_INLINE; |
There was a problem hiding this comment.
Added a test for the send path's inline decision in 8e40af7 (should_post_inline): a pool-backed message at the granted size is inlined, one byte more is not, and user-registered memory never is.
| if (qp == nullptr) { | ||
| // The device may not support inline data, try again without it | ||
| attr.cap.max_inline_data = 0; | ||
| qp = IbvCreateQp(GetRdmaPd(), &attr); |
There was a problem hiding this comment.
Covered by the tests added in 9b2a7c8 (create_qp_with_inline_data).
The 236-byte limit keeps an inlined Send within the 256-byte BlueFlame buffer of mlx5 NICs. Apply it only on Mellanox/NVIDIA devices (vendor ID 0x02c9, kept from the device query GlobalRdmaInitializeOrDie already does); on other devices the inline size the QP was granted is the limit. If QP creation with 236 bytes of inline data fails, try 64 bytes before falling back to none. irdma (Intel E810) rejects requests above 101 bytes, so on those NICs nothing was inlined. Generated-by: Claude Code (Claude Opus 5.5) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After a refused 236-byte inline request, QP creation was retried with 64 bytes, then with none. Devices take less than 236 by different amounts: Intel's irdma refuses more than 216, 101 or 48 bytes depending on the generation (101 on an E810, 48 on an X722), Alibaba's erdma more than 96. With 64 as the only middle step, an E810 or erdma QP got 64 bytes, a newer irdma QP 64 instead of 216, and an X722 QP none. The verbs API cannot report the limit, so step down to it. A device whose kernel driver has a fixed limit, found by the vendor ID kept from the device query GlobalRdmaInitializeOrDie already does, starts there: 216, then 101, then 48 on Intel, 96 on Alibaba, each capped at 236. Other devices start at 236. Each refused size is followed by the next one 16 bytes smaller, down to 0, so a failure unrelated to inline data still ends as before, after a few more attempts at QP creation. Devices that accept 236 (mlx5, rxe) are unchanged. libfabric's verbs provider also probes the limit by trial QP creation (vrb_find_max_inline()). Generated-by: Claude Code (Claude Opus 5.5) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Replaced the 64-byte fallback: devices whose driver has a fixed inline limit now start there, by vendor ID (Intel 216, 101 or 48 depending on the irdma generation, 101 on the E810 and 48 on the X722; Alibaba eRDMA 96), and anything else that refuses steps down 16 bytes at a time until it is accepted. The verbs API can't report the limit; libfabric's verbs provider probes it too (vrb_find_max_inline). mlx5 and rxe accept 236 on the first try, so nothing changes there. |
| static uint32_t InlineDataToRequest(uint32_t vendor_id, uint32_t size, | ||
| bool refused) { | ||
| for (const auto& known : KNOWN_INLINE_LIMITS) { | ||
| if (known.vendor_id == vendor_id && | ||
| (!refused || known.max_inline < size)) { | ||
| return std::min(known.max_inline, size); |
There was a problem hiding this comment.
Added unit tests for the size selection and the retry loop, with a stubbed IbvCreateQp.
wwbmmm
left a comment
There was a problem hiding this comment.
[replied by brpc-oncall robot] The inline-send path and the QP inline-capability fallback look correct, and the eligibility check conservatively excludes user-registered memory (which may be device memory). One robustness suggestion on the QP retry loop below.
| attr.qp_type = IBV_QPT_RC; | ||
| return IbvCreateQp(GetRdmaPd(), &attr); | ||
| ibv_qp* qp = IbvCreateQp(GetRdmaPd(), &attr); | ||
| while (qp == nullptr && inline_size > 0) { |
There was a problem hiding this comment.
[replied by brpc-oncall robot] This loop steps down and retries on any IbvCreateQp failure, not only on failures caused by the requested inline size. On a device that rejects every inline size (or that fails for an unrelated reason such as ENOMEM or reaching the device QP limit), AllocateQpCq now performs up to 16 failed ibv_create_qp calls per resource, and the PLOG(WARNING) << "Fail to create QP" in the caller reports the errno of the last (0-inline) attempt rather than the root cause. Consider logging each step-down (e.g. LOG_IF(INFO, FLAGS_rdma_trace_verbose), following the convention used elsewhere in this file) and stepping down only when the failure can plausibly come from the inline request (e.g. errno == EINVAL).
There was a problem hiding this comment.
Good point. It now steps down only when ibv_create_qp fails with EINVAL, logs each step under FLAGS_rdma_trace_verbose, and any other failure is reported as before.
AllocateQp() retried with a smaller inline size after any ibv_create_qp failure, so a failure unrelated to inline data, such as ENOMEM or the device's QP limit, led to up to 16 failed calls, and the caller's PLOG reported the errno of the last attempt. Step down only when the failed call left errno == EINVAL, which is how providers refuse an inline size; any other failure is returned at once with its errno. Each step is logged under FLAGS_rdma_trace_verbose. QP creation moves into CreateQpWithInlineData(), which takes the vendor ID as a parameter. It and InlineDataToRequest() are no longer static, so that brpc_rdma_unittest can test them with a stubbed IbvCreateQp: the first request per vendor, the Mellanox cap at 236 bytes, Intel 216, 101 and 48, Alibaba 96, an unknown device stepping down to 92, a non-EINVAL failure stopping at once, and a device refusing every size ending at 0. Generated-by: Claude Code (Claude Opus 5.5) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CutFromIOBufList() inlines a message only when it fits the inline size the QP was granted and all of its blocks come from the block pool; user registered memory may be device memory, which the CPU cannot copy from. Move that decision into ShouldPostInline(), exposed for UT, and test it in brpc_rdma_unittest: a pool-backed message at the granted size is inlined, one byte more is not, and a message in user registered memory never is, however small. CutFromIOBufList() itself needs an initialized RDMA device: it returns early in the unit test binary, sizes its SGE list from the device and looks blocks up in the registered block pool. So the decision is tested on its own; the send path is otherwise unchanged. Generated-by: Claude Code (Claude Opus 5.5) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
I've now measured this on hardware with brpc's own rdma_performance example. Setup: two hosts with Intel Xeon Platinum 8360Y and ConnectX-6 Dx (RoCEv2, 100 Gb/s), Ubuntu 24.04, kernel 6.8, rdma-core 50. Server and client each pinned to one core on the NIC's NUMA node, one thread, depth 1, echo on. Base is this PR's merge base (535d8f8), head is this PR (8e40af7). 10 rounds per size, base and head alternating. I took the time per RPC from the client's printed throughput (attachment bytes / MB/s), because Avg-Latency is printed in whole microseconds. The sizes below are the lengths of the messages as posted. On these NICs the head inlines a message when it is 236 bytes or less. I checked that with a counting build: every send of 236 bytes or less went inline, and none above.
So when both messages are inlined, an echo RPC takes about 1 us less (about 3.5%). With only the response inlined it's about half that, and when neither message fits there's no change I can resolve. |
All QPs are created on the same device with the same attributes apart from their queue sizes, so a device that refuses the first inline size made every one of the rdma_prepared_qp_cnt QPs repeat the same step-down (216, 101, 48 on an Intel device limited to 48 bytes). Remember the size the last QP was created with and start there. If that size is refused, the EINVAL step-down continues from it as before. Each QP still takes its own granted size from attr.cap.max_inline_data. Generated-by: Claude Code (Claude Opus 5.5) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
039dfe0 to
8887173
Compare
| if (qp != nullptr) { | ||
| g_inline_data_request.store(inline_size, butil::memory_order_relaxed); | ||
| // ibv_create_qp writes the granted inline data size back into attr | ||
| *max_inline_data = attr->cap.max_inline_data; | ||
| if (vendor_id == MELLANOX_VENDOR_ID) { | ||
| *max_inline_data = std::min(*max_inline_data, BF_MAX_INLINE_DATA); | ||
| } |
There was a problem hiding this comment.
This is on purpose. ibv_create_qp() either fails or grants at least what was asked: ibv_create_qp(3) says "the values will be greater than or equal to the values requested". So a size that was accepted once is accepted again on the next QP with the same attributes, and the step-down doesn't repeat. Caching the request keeps every QP asking for exactly what the first one asked for. Each QP still reads its own grant from attr.cap.max_inline_data, and the Mellanox cap is applied to that.


What problem does this PR solve?
Issue Number: none
Problem Summary:
RdmaEndpoint::CutFromIOBufList() posts every message by reference, so for a small RPC request or response the NIC first has to DMA-read the data from host memory. The QP requests no inline data (MAX_INLINE_DATA was never used and was commented out in #2876).
What is changed and the side effects?
Changed:
236 bytes is the largest Send whose inlined WQE fits the 256-byte BlueFlame buffer of current mlx5 NICs. Larger messages are posted as before. The 236-byte limit applies only on Mellanox/NVIDIA NICs (vendor ID 0x02c9); other devices inline up to the size they grant. The verbs API cannot report the inline limit, so devices whose driver has a fixed one try it first, by vendor ID (Intel 216, 101 or 48 depending on the irdma generation, 101 on the E810 and 48 on the X722; Alibaba eRDMA 96), and a device that refuses a size is asked for 16 bytes less each time until it accepts, down to none.
Side effects:
Performance effects: removes a host memory read before each small message is sent. In a separate ping-pong microbenchmark that sends the same small RDMA write with and without IBV_SEND_INLINE, inlining cut one-way latency by about 0.43 us on Xeon 8480C hosts with ConnectX-7 and 0.89 us on EPYC 7742 hosts with ConnectX-6. On brpc's own rdma_performance example (Xeon 8360Y, ConnectX-6 Dx, RoCEv2), an echo RPC took about 1 us less when both messages were inlined (29 us before); the numbers are in a comment below. With the default rdma_max_sge the send WQE does not grow, but with a small rdma_max_sge it can.
Breaking backward compatibility: no.
Check List:
Tested on soft-RoCE (rxe) with the rdma_performance server and an echo client: attachments from 0 to 300 KB compared byte for byte, messages up to the 512 bytes rxe grants went inline (rxe is not a Mellanox device) and larger ones by reference, and the same again with a test hook that made QP creation refuse inline sizes above a device limit (216, 101 or 48 with an Intel vendor ID, 96 with Alibaba's, 101 with rxe's own), where every QP got the expected size (216, 101, 48, 96 and 92 bytes) and messages up to it went inline. Unit tests in brpc_rdma_unittest cover the size selection and the retry loop with a stubbed IbvCreateQp, including a non-EINVAL failure that stops at once, and the whole test binary passes (88 tests, without RDMA hardware). libbrpc.a builds with WITH_RDMA=ON.
Generated with Claude Code