Skip to content

Give a task cloned while its parent has stashed signals the right mask - #4108

Open
Keno wants to merge 1 commit into
rr-debugger:masterfrom
ChronitonAI:clone-child-sigmask
Open

Keno wants to merge 1 commit into
rr-debugger:masterfrom
ChronitonAI:clone-child-sigmask

Conversation

@Keno

@Keno Keno commented Oct 6, 2026

Copy link
Copy Markdown
Member

[Encountered, debugged and patch by AI]

While a task has stashed signals, will_resume_execution blocks every signal except rr's own in it whenever rr resumes it, and did_wait restores its mask at the next stop. If the task runs a clone, fork or vfork in that window, the kernel copies that temporary mask into the new task, and nothing restores it there. The new task then starts with almost every signal blocked: signals sent to it stay pending, and a signal sent to a thread is lost when the thread exits.

This happens when a signal arrives while the task is stopped at the syscall entry of a clone: the clone fails with ERESTARTNOINTR, prepare_clone reenters the syscall, Task::enter_syscall stashes the signal, and the restarted clone runs with the temporary mask. In chaos mode the short time slices make the time slice signal arrive there often enough to notice. A tracer whose tracee creates threads (PTRACE_O_TRACECLONE) that each send themselves SIGUSR1 sometimes never saw a thread's signal-delivery stop: the thread had SIGUSR1 blocked and exited with the signal pending.

So when we set up the new task after the clone, if the parent has stashed signals, give the new task the parent's real mask, which did_wait restored at the parent's last stop. We do that in RecordTask::post_wait_clone, which already copies the signal handlers from the parent: Task::clone calls it as soon as the new task has stopped for the first time, before anything else touches the new task.

C libraries hide this in some cases, because they block all signals around the clone and set the mask in the new task: glibc does so in pthread_create since 2.32, in fork since 2.41, and in posix_spawn. Raw clone, fork and vfork syscalls, glibc's vfork() (a bare syscall in every version), and threads and forks with older glibc versions see the wrong mask, and the mask survives execve.

The new test ptrace_signal_during_clone triggers this deterministically. The tracee blocks SIGUSR2 and makes a raw clone syscall (SIGCHLD). The tracer stops it at the clone's syscall-entry stop (PTRACE_SYSCALL), sends it SIGUSR1 and resumes it, so the clone is restarted. The clone child checks that it has the same mask and reports the result in its exit status: with the wrong mask SIGABRT is blocked, so test_assert can't abort.

While a task has stashed signals, will_resume_execution blocks every
signal except rr's own in it whenever rr resumes it, and did_wait
restores its mask at the next stop. If the task runs a clone, fork or
vfork in that window, the kernel copies that temporary mask into the
new task, and nothing restores it there. The new task then starts with
almost every signal blocked: signals sent to it stay pending, and a
signal sent to a thread is lost when the thread exits.

This happens when a signal arrives while the task is stopped at the
syscall entry of a clone: the clone fails with ERESTARTNOINTR,
prepare_clone reenters the syscall, Task::enter_syscall stashes the
signal, and the restarted clone runs with the temporary mask. In chaos
mode the short time slices make the time slice signal arrive there
often enough to notice. A tracer whose tracee creates threads
(PTRACE_O_TRACECLONE) that each send themselves SIGUSR1 sometimes never
saw a thread's signal-delivery stop: the thread had SIGUSR1 blocked and
exited with the signal pending.

So when we set up the new task after the clone, if the parent has
stashed signals, give the new task the parent's real mask, which
did_wait restored at the parent's last stop. We do that in
RecordTask::post_wait_clone, which already copies the signal handlers
from the parent: Task::clone calls it as soon as the new task has
stopped for the first time, before anything else touches the new task.

C libraries hide this in some cases, because they block all signals
around the clone and set the mask in the new task: glibc does so in
pthread_create since 2.32, in fork since 2.41, and in posix_spawn. Raw
clone, fork and vfork syscalls, glibc's vfork() (a bare syscall in
every version), and threads and forks with older glibc versions see the
wrong mask, and the mask survives execve.

ptrace_signal_during_clone triggers this deterministically. The tracee
blocks SIGUSR2 and makes a raw clone syscall (SIGCHLD). The tracer
stops it at the clone's syscall-entry stop (PTRACE_SYSCALL), sends it
SIGUSR1 and resumes it, so the clone is restarted. The clone child
checks that it has the same mask and reports the result in its exit
status: with the wrong mask SIGABRT is blocked, so test_assert can't
abort.

Results on x86_64 (Linux 7.0, glibc 2.31), in each of the four test
variants (64- and 32-bit, with and without the syscall buffer):
- Without this change, ptrace_signal_during_clone fails in every run
  (20 of 20 per variant): the clone child has the wrong mask.
- With this change, it passes in 50 of 50 runs per variant.
- Natively, it passes in 500 of 500 runs per bitness.

Chaos-mode recordings, 6 in parallel, of a tracer whose tracee creates
threads one after another (PTRACE_O_TRACECLONE), each of which sends
itself SIGUSR1:
- with 200 threads, 9 of 300 recordings lost a signal without this
  change, and none of 300 with it;
- with 50 threads, 2 of 300 without this change, none of 300 with it;
- ptrace_wnohang_poll from rr-debugger#4105 (50 threads, polling with WNOHANG),
  on top of rr-debugger#4105: 2 of 600 without this change, none of 1800 with it.
In chaos recordings of a program whose threads and fork children check
their own masks (400 threads and 100 forks per run), 47 of 120 saw a
wrong mask without this change, and none of 120 with it.

The full test suite shows no new failures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Keno added a commit to ChronitonAI/rr that referenced this pull request Oct 6, 2026
AutoRemoteSyscalls can be used in a task that is stopped at a seccomp
stop. To get back to that stop afterwards, restore_state_to() moves the
ip back to the syscall instruction and resumes the task, which executes
the syscall instruction again and stops at the seccomp stop again. On
x86-64 the syscall instruction sets rcx to the return address and r11 to
rflags, so the task's registers no longer match the ones it had at the
stop. rr had canonicalized them there (rcx = -1, r11 = 0x246) and
doesn't do it again, so the syscall runs with the clobbered values. rr
records the syscall's exit with those, replay canonicalizes them, and
replay fails:

  Mismatched registers, replay vs rec: rcx 0xffffffffffffffff != <ip>

A restarted syscall makes it worse. rr records the restarted syscall's
entry with the registers it saved at the interrupted syscall's exit,
so the clobbered rcx can also show up at the entry and in the next
asynchronous signal's event. Then replay reaches the signal with
different registers, fast-forwards past it, runs away and hangs.

One way to get there: when a process that shares an address space with
others (CLONE_VM, e.g. posix_spawn()'s child) execs, rr unmaps its
syscall buffer and scratch memory from that address space with
AutoRemoteSyscalls in the next task there that it processes
(Task::unmap_dead_syscallbufs_if_required at the end of record_step).
If that task is stopped at a seccomp stop, e.g. at a syscall that was
restarted after a signal interrupted it, the registers are clobbered. A
program whose threads spawn processes while another thread waits in
pthread_join() hit this in 3 of 69 recordings with --chaos.

So after getting back to the seccomp stop, restore the registers that
the task had there.

Test: in x86/seccomp_stop_remote_syscall, a ptracer stops its tracee at
the entry of a direct getsid() syscall (PTRACE_SYSCALL), which rr
emulates at the syscall's seccomp stop. While the tracee is stopped
there, its CLONE_VM child execs. Then the ptracer resumes the tracee,
and rr's cleanup of the child's buffers runs in the tracee at the
seccomp stop. This is deterministic. On 32-bit x86 the int $0x80
doesn't clobber registers, so the test only fails on x86-64 without
this change.

Results on x86_64 (Linux 7.0, glibc 2.31), in each of the four test
variants (64- and 32-bit, with and without the syscall buffer):
- Without this change, the test fails in every run of the 64-bit
  variants (25 of 25 each) with the register mismatch, and passes in the
  32-bit ones.
- With this change, it passes in 50 of 50 runs per variant.
- Natively, it passes in 500 of 500 runs per bitness.
- With this change and rr-debugger#4108 (without which that program
  often dies during recording), 39 of 40 --chaos recordings of that
  program replay with the same output. The other one hits a separate
  replay bug, unrelated to this change: rr unmaps the wrong part of its
  own copy of a syscall buffer.
- The full test suite shows no new failures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Keno added a commit to ChronitonAI/rr that referenced this pull request Oct 6, 2026
AutoRemoteSyscalls can be used in a task that is stopped at a seccomp
stop. To get back to that stop afterwards, restore_state_to() moves the
ip back to the syscall instruction and resumes the task, which executes
the syscall instruction again and stops at the seccomp stop again. On
x86-64 the syscall instruction sets rcx to the return address and r11 to
rflags, so the task's registers no longer match the ones it had at the
stop. rr had canonicalized them there (rcx = -1, r11 = 0x246) and
doesn't do it again, so the syscall runs with the clobbered values. rr
records the syscall's exit with those, replay canonicalizes them, and
replay fails:

  Mismatched registers, replay vs rec: rcx 0xffffffffffffffff != <ip>

A restarted syscall makes it worse. rr records the restarted syscall's
entry with the registers it saved at the interrupted syscall's exit,
so the clobbered rcx can also show up at the entry and in the next
asynchronous signal's event. Then replay reaches the signal with
different registers, fast-forwards past it, runs away and hangs.

One way to get there: when a process that shares an address space with
others (CLONE_VM, e.g. posix_spawn()'s child) execs, rr unmaps its
syscall buffer and scratch memory from that address space with
AutoRemoteSyscalls in the next task there that it processes
(Task::unmap_dead_syscallbufs_if_required at the end of record_step).
If that task is stopped at a seccomp stop, e.g. at a syscall that was
restarted after a signal interrupted it, the registers are clobbered. A
program whose threads spawn processes while another thread waits in
pthread_join() hit this in 3 of 69 recordings with --chaos.

So after getting back to the seccomp stop, restore the registers that
the task had there.

Test: in x86/seccomp_stop_remote_syscall, a ptracer stops its tracee at
the entry of a direct getsid() syscall (PTRACE_SYSCALL), which rr
emulates at the syscall's seccomp stop. While the tracee is stopped
there, its CLONE_VM child execs. Then the ptracer resumes the tracee,
and rr's cleanup of the child's buffers runs in the tracee at the
seccomp stop. This is deterministic. On 32-bit x86 the int $0x80
doesn't clobber registers, so the test only fails on x86-64 without
this change.

Results on x86_64 (Linux 7.0, glibc 2.31), in each of the four test
variants (64- and 32-bit, with and without the syscall buffer):
- Without this change, the test fails in every run of the 64-bit
  variants (25 of 25 each) with the register mismatch, and passes in the
  32-bit ones.
- With this change, it passes in 50 of 50 runs per variant.
- Natively, it passes in 500 of 500 runs per bitness.
- With this change and rr-debugger#4108 (without which that program
  often dies during recording), 39 of 40 --chaos recordings of that
  program replay with the same output. The other one hits a separate
  replay bug, unrelated to this change: when rr drops a dead syscall
  buffer from replay's model of the address space, it can also unmap its
  own copy of a live syscall buffer mapped just below it.
- The full test suite shows no new failures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant