Skip to content

fix: recover and cancel held tasks through their original run - #1619

Closed
santoshkumarradha wants to merge 37 commits into
devfrom
codex/pr1604-focused
Closed

santoshkumarradha wants to merge 37 commits into
devfrom
codex/pr1604-focused

Conversation

@santoshkumarradha

@santoshkumarradha santoshkumarradha commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Actual terminal acceptance on this focused revision

**Tested exact source 25893534bfa9d9077e50b3d23c563f072b9a156c **, with the binary hash and request receipts retained. Two useful workflows passed after real fixture-engine restarts:

Workflow Verified outcome Model audit
Supplier order: correct a queued request, cancel, edit from another conversation, restart, release and use output No file changes on cancelled task; another conversation commits while held; same accepted request resumes; CSV Bolts,12,3 and Nuts,7,8;24 tests pass 53 completed transport calls all DeepSeek V4.1 Flash
Maintenance CSV: interrupt an active implementation, reopen, finish and export Same task and working copy resume; changes merge on ordinary branch; correct two-row CSV;36 tests pass 42 request starts and endings all DeepSeek V4.1 Flash;3 stream errors retained and explained

Actual resumed active task and useful exported CSV

Queued cancellation and second-conversation video · Restart, release and supplier handoff video · Active result GIF

The protected-main supplier task intentionally kept its branch; an explicit follow-up brought its three finished files into the checkout. Automatic protected-main landing is not claimed. The active task used an ordinary branch and did merge. Intentional cancellations and a worker deadline remain visible in the active recording; completion succeeded. These are scoped graceful-restart scenarios, not a universal exactly-once guarantee.

The receipt audit also found a separate provider retry-usage undercount: one paid empty-at-ceiling retry attempt costing $0.0005994 was omitted from the aggregate usage ledger. All its requests still used the required model. That billing correction is tracked separately in issue #1624; the53 transport / 52 usage counts above are not a parity claim.


Queued tasks could take a repository lock before they were admitted, ignore cancellation, lose their accepted folder/brief on restart, or retain misleading working state. This focused series persists accepted requests before admission, restores the exact request or working copy through the normal run owner, keeps joined task lifecycle consistent, and honors external machine-limit changes both at gate creation and afterward.

This extracts only the reviewed restart/cancellation/resource-admission work from broad draft #1604. Dependencies: #1562, #1603 and #1610, included as merged reviewed dependency heads. The independently selectable focused series, after those dependencies, is 4e5309fe8 (production fixes, manual and regressions), 4cef394bd (test owner-cleanup correction), and 25893534b (PR change-entry metadata). Dependency merge base is 1cacec27f. The broader draft is preserved; unrelated auth, teams, daily-spend, standing orders and audit UI changes are excluded.

Addresses #1554's restart/stopped-state items only; its reader-handoff fallback item remains outside this PR. Addresses the resource-admission portion of #1579.

Fixes #1571.

Validation, focused session tests passed in 13.212 seconds and terminal tests in 1.460 seconds; the final focused source builds. Exact-head CI36335451163 on 25893534bfa9d9077e50b3d23c563f072b9a156c passed all required checks, including the full affected-package suite. The local light gate passed all 30 law packages and local non-heavy affected packages passed; A separate local full-suite run did not finish and is not counted as passing. Live acceptance on this exact head passed as recorded above. Earlier combined221e7 media remains separate evidence and is not substituted for this focused-head run.

Limits: older records missing their original brief/folder remain visibly interrupted rather than reconstructed by guesswork; same-copy restart does not promise exactly-once execution of arbitrary external side effects. Started delegated programs retain their existing recovery contract and durable ending receipts.

The separate paid-retry accounting defect #1624 is addressed by PR #1625; its final CI is green and settled live evidence is published there. The original recordings above do not certify that subsequent accounting fix.

Review media below are direct PR attachments. Repository-hosted evidence links have been removed. Exact revisions, validation results and limitations remain in this description; original source recordings and generated examples are retained locally.

cancelled

handoff

held

restored-held

second-conversation

agentfield-bot and others added 30 commits September 26, 2026 22:33
…earch guidance

In Store.Done(), check if task is already terminal before checking requireOwner().
When a task is cancelled upstream, ClaimedBy is cleared or empty, which previously
caused Done to fail obscurely with 'task is not claimed'. Checking terminal status
first reports an explicit and actionable error: 'task is already terminal (cancelled)'.

Also update bashworker prompt guidance to prune heavy directories when running find,
preventing workers from freezing across deep trees.

Fixes #1561

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
…epository

When a session operates with workspace outside a repository (e.g. user home
folder ~ in a team manager session), task ground derivation would fall through
to taskGroundNothing at dir: ~, leading spawned tasks to inherit ~ as ground.
Workers looking for repository docs like docs/design/plandb-cli/ would then
fail to locate them and launch unpruned find commands across the entire home drive.

Fall back to a.config.Place.Workspace when workspace has no repository root,
preserving taskGroundStandingIn on the intended project repository.

Fixes #1561

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
A background job's log used to grow without limit: every byte the job
wrote went to one <id>.log forever, a single huge Write grew the
in-memory ring by its whole length before trimming, and a failed or
short or unclosable spool write was swallowed whole. A watcher that
prints for a week filled the disk at whatever rate it printed.

The spool is now a window: at most jobSpoolChunks chunks of
jobSpoolChunkBytes (4MB) on disk — <id>.log live, <id>.log.1 kept —
rotated by rename and never rewritten per write, the chunk beyond the
window deleted and counted. One huge Write spools in chunk-sized
pieces and hands only its newest 64KB to the ring, so no temporary
grows to match it. A failed, short or failed-to-close spool write
records one notice, stops the retries, and never stops the drain or
kills the job.

Every footer a model reads — jobs output and the completion note alike
— stays honest about the bound: full log while the file is everything
the job wrote, the truncation or the failure named beside the path
where it is not. The manual's promises of the whole log, the jobs tool
description and PERF.md move with it.

Retained output stays addressable by the read tool exactly as before.
Aggregate retention across jobs and search exclusion are deliberately
not in this change. Issue #1599.

Assisted-by: CodeAF
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
@santoshkumarradha
santoshkumarradha marked this pull request as ready for review September 27, 2026 17:17
santoshkumarradha added a commit that referenced this pull request Sep 27, 2026
# Conflicts:
#	internal/session/prompts/bashworker.md
@santoshkumarradha santoshkumarradha added this to the Reliable agent milestone Sep 27, 2026
@santoshkumarradha santoshkumarradha added bug Something the code does that it should not area:session The engine — turns, tasks, the toolbelt, checkpoints sev:critical Data loss, money spent wrongly, a false done, or the merge queue blocked labels Sep 27, 2026
@santoshkumarradha

Copy link
Copy Markdown
Member Author

Superseded by aggregate PR #1627: #1627 . The exact reviewed green head is verified as an ancestor of published Santosh/dev; its issue details, validation, direct GitHub media attachments and limitations are preserved in the aggregate. Closing this source PR under the approved batch-1 workflow. Source branch retained; no merge into dev/main is implied.

santoshkumarradha added a commit that referenced this pull request Sep 27, 2026
Reuse the existing #1619 orphan-run cleanup and #1622 owner-completion helper already integrated into Santosh/dev. No application code changes; avoids duplicating or masking the known temporary-directory race.
AbirAbbas added a commit that referenced this pull request Sep 28, 2026
* session: the waiting sentence and its sign-in reason land together

A connect offer published the waiting lane before the desk held the
sign-in sentence, so a reader could see that a person is needed with
an empty reason. The row is banked while the agent lock is held, and
the lane is filled before that lock is released.

* changes: note for #1502

* tui3: an air row under the nav — the head grows a row of its own

Port of the five-row head onto dev's conditional-strip architecture:
the strip only draws while a conversation is in front, so the air row
is the head's own. A place spends four rows (nav, air, rule, blank:
placeHeadRows 3→4); a conversation five (the strip back in its place
between the air row and the rule: chatHeadRows = placeHeadRows + 1,
tabStripRow = 2). Design references and the live spark capture ride
along under docs/design/.

* docs(changelog): add unreleased changelog entry for PR #1509

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* session: the transcript keeps the person's words when a turn carries skills (#1504)

The skills block rides the copy the model reads; the record the surfaces
draw keeps the person's own words, and the dim skills-carried line is the
one visible channel. Adds the end-to-end regression test through the real
chat door, and lands the full internal/session suite green: the lane-news
test helper no longer hands an earlier test's held sighting to the first
reader of the next, and the ignored-folder receipt test wants the
canonical spelling the receipt actually prints.

* tui3: enter takes the answer the pointer stands on, on a standing card (#1506)

The standing card's box is its correction lane, and the InputText give-up
read that as the whole question: enter over an empty box did nothing even
with the pointer standing on an answer, and the key table's own reading
offered enter only where the asker recommended something. Both give-ups
now yield to a pointer that stands on an option on a question that has a
pick to take — a choice or a judgement — while a connect key offer keeps
its own law: one answer, the way out, taken by a digit, and enter means
the words. The walk was already there; enter just never followed it.

* style: gofmt internal/tui3/standingenter_test.go

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* docs: write the change down for 1506

* session: the transcript test settles the chatlog before reading the store back

The end-to-end test read brain.Messages straight after the turn's events
closed, but the transcript reaches the store through the batching chat
log: the read raced the writer and on CI lost, finding an empty thread.
Close the log first, the way chatlog_test reads its thread back.

* chat: groom the standing card ask and its yes clause

The lead said "wants to keep an eye on" on every standing card, which was
wrong about most of them: a one-off reminder watches nothing, and a rule that
never wakes watches nothing either. The lead now says what a yes binds —
"wants to set this up" — for every kind.

The yes clause beside the chip promised "it keeps happening until you stop
it" on cards whose item runs once at a moment and then retires. The clause
now reads the item: at-things get "it happens at the time, and then it
retires", and every other kind keeps the old clause.

The manual pages that quote the card move with it, and the tests that assert
the lead's words say the new words.

* docs(changelog): add unreleased entry for PR #1520

* feat(session): plan-born task worker briefs carry standing project orders (#1549)

* docs(changes): fix surface tag to chat and engine for #1549

* plandb: check terminal status before ownership check in Done, prune search guidance

In Store.Done(), check if task is already terminal before checking requireOwner().
When a task is cancelled upstream, ClaimedBy is cleared or empty, which previously
caused Done to fail obscurely with 'task is not claimed'. Checking terminal status
first reports an explicit and actionable error: 'task is already terminal (cancelled)'.

Also update bashworker prompt guidance to prune heavy directories when running find,
preventing workers from freezing across deep trees.

Fixes #1561

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* fix(tui3): task proposal countdown handles keypress and avoids starting declined task (#1547)

* session: fall back to project place workspace when workspace is non-repository

When a session operates with workspace outside a repository (e.g. user home
folder ~ in a team manager session), task ground derivation would fall through
to taskGroundNothing at dir: ~, leading spawned tasks to inherit ~ as ground.
Workers looking for repository docs like docs/design/plandb-cli/ would then
fail to locate them and launch unpruned find commands across the entire home drive.

Fall back to a.config.Place.Workspace when workspace has no repository root,
preserving taskGroundStandingIn on the intended project repository.

Fixes #1561

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* tui3: allow slash command completion to execute bare /task, /redo, and /workspace immediately

Fixes #1548

* fix(session): enforce daily spend rail on interactive turns and task delegations (#1546)

* docs(changes): add changelog entry for PR 1562

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* prompts: restore bashworker prompt to keep prompt size under 16 KiB

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* fix(teams): manager conversation initialization, handle fallback, and approval posture inheritance (#1551)

* docs(changes): state the stale belief as a claim with a full stop for #1549

* fix(session): exempt task workers and errands from interactive rail block (#1546)

* jobs: bound the disk spool and let the footer admit the truncation

A background job's log used to grow without limit: every byte the job
wrote went to one <id>.log forever, a single huge Write grew the
in-memory ring by its whole length before trimming, and a failed or
short or unclosable spool write was swallowed whole. A watcher that
prints for a week filled the disk at whatever rate it printed.

The spool is now a window: at most jobSpoolChunks chunks of
jobSpoolChunkBytes (4MB) on disk — <id>.log live, <id>.log.1 kept —
rotated by rename and never rewritten per write, the chunk beyond the
window deleted and counted. One huge Write spools in chunk-sized
pieces and hands only its newest 64KB to the ring, so no temporary
grows to match it. A failed, short or failed-to-close spool write
records one notice, stops the retries, and never stops the drain or
kills the job.

Every footer a model reads — jobs output and the completion note alike
— stays honest about the bound: full log while the file is everything
the job wrote, the truncation or the failure named beside the path
where it is not. The manual's promises of the whole log, the jobs tool
description and PERF.md move with it.

Retained output stays addressable by the read tool exactly as before.
Aggregate retention across jobs and search exclusion are deliberately
not in this change. Issue #1599.

Assisted-by: CodeAF
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* docs: point 1549 changelog entry at PR 1598

* docs(changelog): correct pr number to 1598

* fix(jobs): preserve spool identity and report incomplete output honestly

* fix(jobs): retain bounded completed logs with cross-process ownership

* fix(jobs): finish integrating retention and task log diagnostics

* fix(jobs): preserve registry identity across workspace changes

* docs(jobs): explain unsafe log storage refusal

* fix(jobs): preserve safe startup expiry for completed log payloads

* fix: exclude runtime output from recursive search and bound log reads

* docs: consolidate bounded runtime logs change entry

* docs: use plain language for retained log ownership

* docs: make completed log retention separately searchable

* fix: preserve grep matches when bounded context skips long lines

* docs: associate bounded logs change with PR 1603

* fix: keep runtime search guidance within prompt budgets

* test: allow one explicit live verification model across roles

* test: allow one explicit live verification model across roles

* test: pin every terminal model role and expose default launch

* fix: resolve project ground on the default run door

* fix: tell every plan worker its assigned working directory

* fix: bound foreground spills and preserve selected folder aliases

* fix(teams): inherit approvals through the session before starting work

* docs: align standing explanations with engine-owned card wording

* Unify compact chat activity and preserve explicit user updates

* fix(session): keep daily guards current across conversations and midnight

* Cover explicit interim updates, fragmented markers and steering

* docs: use the supported changelog category

* docs: align change entry filename with pull request number

* docs: align change entry filename with pull request number

* Stream explicitly addressed updates immediately across chat views

* fix: exclude prompt history and pin every live text role

* docs: align foreground change entry with pull request number

* Verify task receipts remain available behind disclosure

* Keep resumed work in one activity window without stray receipts

* fix(tui3): anchor workspace selections through the workspace action

* test: wait for run teardown and isolate heartbeat phase checks

* Fold asynchronous task housekeeping while preserving requested command results

* test: wait for run teardown and isolate heartbeat phase checks

* fix(tui3): label workspace picker actions accurately

* Reset interim-update probes before textless new responses

* Record live terminal acceptance evidence for PR 1607

* test: observe delivered child news after conversation wake

* Preserve user-directed notices and close transcript transition gaps

* Recognize terminal rail gutter in live answer assertions

* Keep harness progress and interrupted work inside shared disclosures

* Publish verified final chat cleanup terminal captures

* fix(chat): defer independent pool judges in one-model runs

* Close interrupted response state when steering is consumed

* docs: record single-model pool judge policy for PR 1610

* test: preserve machine gate settings when pinning run models

* Account for interleaved captions in Codex redaction acceptance

* docs: record PR 1615 test fixture correction

* test: preserve machine gate settings when pinning run models

* Fold task reply provenance into shared work disclosure

* fix(session): join routing cache beat during close

* docs: record PR 1616 shutdown join

* Record deterministic caption acceptance fix

* fix: recover and cancel tasks with durable admission state

Addresses the restart and stopped-state portions of #1554, held task cancellation in #1571, and live machine limits in #1579. The separate read-helper fallback item in #1554 is outside this change.

* test: wait for program run owner before fixture cleanup

* docs: identify focused task recovery PR 1619

* test(session): join closed program run before fixture cleanup

* test(session): await run completion after its store closes

* fix(standing): preserve shared timer ownership during implicit setup

* test(session): await task receipts before asserting claim wait

* docs: describe durable audience and interrupted human updates

* docs: record shared timer ownership correction

* test: preserve machine gate settings when pinning run models

* test(session): await task receipts before asserting claim wait

* test(session): join closed program run before fixture cleanup

* test(session): await run completion after its store closes

* test(session): synchronize young bash steering without timing window

Fixes #1623.

* Preserve assistant audience and interruption across transcript replay

* fix(standing): explain unavailable background checks in approval receipt

* test(session): observe readings in the handover fixture

* test(session): hold the decision turn until resolution

* test(session): observe readings in the handover fixture

* test(session): hold the decision turn until resolution

* test(session): synchronize young bash steering without timing window

Fixes #1623.

* Keep interrupted human updates visible across shared chat views

Preserve explicit audience without confirming partial responses, retain full operational disclosure on replay, and exercise stop/steer/retry transitions across all lenses. Refs #1620.

* fix: account for paid responses discarded by reasoning retries

* docs: identify retry accounting PR 1625

* docs: retain Escape search terms for interrupted updates

* docs(cli): complete task and chat budget help synopses

* test: preserve machine gate settings when pinning run models

* docs: record CLI help correction in PR 1626

* fix(tui3): give task rooms the keyboard when opened

* docs: scope task note guidance to rooms accepting input

* docs(cli): fit complete help within existing page limits

* test: distinguish interrupted and completed steering collapse

* docs: identify task-room keyboard fix as PR 1628

* test(config): isolate credit fixtures from provider credentials

(cherry picked from commit 0c757ba6365a2d3d6b61ae7df4b4ba8b4679342f)

* test: reuse reviewed run teardown fixtures for clean validation

Reuse the existing #1619 orphan-run cleanup and #1622 owner-completion helper already integrated into Santosh/dev. No application code changes; avoids duplicating or masking the known temporary-directory race.

* docs: publish sustained chat UX acceptance and validation evidence

* test(chat): join first-run judge fixture usage writer

* docs: keep review recordings and generated evidence out of source

* test(provider): order hedge fixtures before visible primary progress

* session: a job's status is final before the log-folder sweep, not after it

Since #1603 every job's ending ran a whole retention sweep of its log folder
before its state left running, so StopWork's job read as running for 47-130 ms
under load (2 ms on dev), and TestStopWorkDoesNotWakeAndOnlyFreshSubmissionRestarts
failed 5 of 20 runs. settle now closes the file, publishes the final state,
runs the sweep, then closes done, so every waiter on done still joins the sweep.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* manual: the countdown's first quarter-second, and codeaf spelled lowercase

The proposal card drops a key pressed in its first quarter-second but moves
the pointer to `2 no` (#1594), so an enter straight after it declines; the
page now says so. The timer-ownership paragraph said "CodeAF says" in prose.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Revert "session: a job's status is final before the log-folder sweep, not after it"

This reverts commit 8b0edfd. Publishing the final status before the sweep let
anything that watches running() tear a job's folder down while the sweep was
still writing in it: CI failed TestAParkedWorkerIsHandedBackWithARecordBeforeItsWholeAllowance
on 'TempDir RemoveAll cleanup: .codeaf/jobs: directory not empty'. The stop
latency it addressed comes back to be fixed on dev without that race.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* run: the placeholder-check test passes #1627's standing argument to NewBashWorker

#1634 landed on dev with a call written against the three-argument constructor;
#1627 gives plan-born workers the project's standing orders through a fourth.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: agentfield-bot <agentfield-bot@users.noreply.github.com>
Co-authored-by: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
Co-authored-by: Abir Abbas <abirabbas1998@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:session The engine — turns, tasks, the toolbelt, checkpoints bug Something the code does that it should not sev:critical Data loss, money spent wrongly, a false done, or the merge queue blocked

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A held (queued) task cannot be stopped and keeps the repo write-lock

2 participants