Skip to content

bug(podman): stale supervisor event can stop a restarted sandbox workload #3776

Description

@purp

User Story

As an OpenShell user on the Podman driver, I want openshell sandbox start right after openshell sandbox stop to reliably bring the sandbox back, so that stop/start cycles don't randomly end in ContainerExited and CI stays green.

Problem Statement

The Podman watcher can stop a freshly restarted workload container because of a stale supervisor container event from the previous run.

start_sandbox records a lifecycle fence so delayed stop/die events from the previous run are ignored (driver.rs#L1526-L1536). It then stops the supervisor, starts the workload, and starts the supervisor last (driver.rs#L1537-L1583).

In the watcher, only workload events check that fence (watcher.rs#L274-L296). Supervisor-role events return early and skip it (watcher.rs#L251-L272). They go straight to inspect_workload, which contains a running workload whose supervisor is not running by calling stop_container on it (watcher.rs#L377-L396).

A delayed supervisor die event can arrive after start_container(workload) and before start_container(supervisor). The watcher then sees a running workload and an exited supervisor, stops the new workload, and the gateway marks the sandbox Error / ContainerExited.

This analysis comes from code reading plus the timeline below. The supervisor-event branch doesn't log, so the log can't show directly which event triggered the containment.

Impact / Why This Matters

Acceptance Criteria

  • Stale supervisor container events from before the latest start_sandbox do not stop or fail the restarted workload.
  • Real supervisor exits after start still trigger workload containment. The containment behavior in inspect_workload is preserved.
  • A unit test covers a delayed supervisor die event delivered between the workload start and the supervisor start.
  • The sandbox_lifecycle conformance scenario passes on fedora-podman-rootful across repeated runs.

Reproduction Steps

Intermittent, timing dependent:

  1. Run the gateway with the Podman driver (rootful reproduced in CI).
  2. openshell sandbox create --name ss --detach -- <long-running command> and wait for Ready.
  3. openshell sandbox stop ss, then immediately openshell sandbox start ss.
  4. Occasionally, start fails with Error: × ContainerExited.

Environment

Logs

Gateway timeline (UTC, 2026-09-28), sandbox ct-k1tdbf91py-ss:

03:38:06.042 StopSandbox
03:38:06.193 StartSandbox
03:38:06.341 podman::watcher: Ignoring container stop event from before the latest sandbox start action=die finished_at="2026-09-28T03:38:06.069Z"
03:38:06.585 podman::watcher: Ignoring unhandled Podman event action=init
03:38:06.757 Sandbox phase changed old_phase=Starting new_phase=Error
03:38:06.757 Sandbox failed to become ready reason=ContainerExited

The workload's stale die event was correctly fenced at 06.341. No workload event explains the failure at 06.757, which points at the unfenced supervisor path.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions