fix: recover and cancel held tasks through their original run - #1619
Closed
santoshkumarradha wants to merge 37 commits into
Closed
santoshkumarradha wants to merge 37 commits into
santoshkumarradha wants to merge 37 commits into
Conversation
…earch guidance In Store.Done(), check if task is already terminal before checking requireOwner(). When a task is cancelled upstream, ClaimedBy is cleared or empty, which previously caused Done to fail obscurely with 'task is not claimed'. Checking terminal status first reports an explicit and actionable error: 'task is already terminal (cancelled)'. Also update bashworker prompt guidance to prune heavy directories when running find, preventing workers from freezing across deep trees. Fixes #1561 Assisted-by: CodeAF (gemini-3.8-flash-high) Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
…epository When a session operates with workspace outside a repository (e.g. user home folder ~ in a team manager session), task ground derivation would fall through to taskGroundNothing at dir: ~, leading spawned tasks to inherit ~ as ground. Workers looking for repository docs like docs/design/plandb-cli/ would then fail to locate them and launch unpruned find commands across the entire home drive. Fall back to a.config.Place.Workspace when workspace has no repository root, preserving taskGroundStandingIn on the intended project repository. Fixes #1561 Assisted-by: CodeAF (gemini-3.8-flash-high) Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
Assisted-by: CodeAF (gemini-3.8-flash-high) Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
Assisted-by: CodeAF (gemini-3.8-flash-high) Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
A background job's log used to grow without limit: every byte the job wrote went to one <id>.log forever, a single huge Write grew the in-memory ring by its whole length before trimming, and a failed or short or unclosable spool write was swallowed whole. A watcher that prints for a week filled the disk at whatever rate it printed. The spool is now a window: at most jobSpoolChunks chunks of jobSpoolChunkBytes (4MB) on disk — <id>.log live, <id>.log.1 kept — rotated by rename and never rewritten per write, the chunk beyond the window deleted and counted. One huge Write spools in chunk-sized pieces and hands only its newest 64KB to the ring, so no temporary grows to match it. A failed, short or failed-to-close spool write records one notice, stops the retries, and never stops the drain or kills the job. Every footer a model reads — jobs output and the completion note alike — stays honest about the bound: full log while the file is everything the job wrote, the truncation or the failure named beside the path where it is not. The manual's promises of the whole log, the jobs tool description and PERF.md move with it. Retained output stays addressable by the read tool exactly as before. Aggregate retention across jobs and search exclusion are deliberately not in this change. Issue #1599. Assisted-by: CodeAF Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
…1604-focused # Conflicts: # internal/session/prompts/bashworker.md
santoshkumarradha
marked this pull request as ready for review
September 27, 2026 17:17
santoshkumarradha
added a commit
that referenced
this pull request
Sep 27, 2026
# Conflicts: # internal/session/prompts/bashworker.md
Member
Author
|
Superseded by aggregate PR #1627: #1627 . The exact reviewed green head is verified as an ancestor of published Santosh/dev; its issue details, validation, direct GitHub media attachments and limitations are preserved in the aggregate. Closing this source PR under the approved batch-1 workflow. Source branch retained; no merge into dev/main is implied. |
AbirAbbas
added a commit
that referenced
this pull request
Sep 28, 2026
* session: the waiting sentence and its sign-in reason land together A connect offer published the waiting lane before the desk held the sign-in sentence, so a reader could see that a person is needed with an empty reason. The row is banked while the agent lock is held, and the lane is filled before that lock is released. * changes: note for #1502 * tui3: an air row under the nav — the head grows a row of its own Port of the five-row head onto dev's conditional-strip architecture: the strip only draws while a conversation is in front, so the air row is the head's own. A place spends four rows (nav, air, rule, blank: placeHeadRows 3→4); a conversation five (the strip back in its place between the air row and the rule: chatHeadRows = placeHeadRows + 1, tabStripRow = 2). Design references and the live spark capture ride along under docs/design/. * docs(changelog): add unreleased changelog entry for PR #1509 Assisted-by: CodeAF (gemini-3.8-flash-high) Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com> * session: the transcript keeps the person's words when a turn carries skills (#1504) The skills block rides the copy the model reads; the record the surfaces draw keeps the person's own words, and the dim skills-carried line is the one visible channel. Adds the end-to-end regression test through the real chat door, and lands the full internal/session suite green: the lane-news test helper no longer hands an earlier test's held sighting to the first reader of the next, and the ignored-folder receipt test wants the canonical spelling the receipt actually prints. * tui3: enter takes the answer the pointer stands on, on a standing card (#1506) The standing card's box is its correction lane, and the InputText give-up read that as the whole question: enter over an empty box did nothing even with the pointer standing on an answer, and the key table's own reading offered enter only where the asker recommended something. Both give-ups now yield to a pointer that stands on an option on a question that has a pick to take — a choice or a judgement — while a connect key offer keeps its own law: one answer, the way out, taken by a digit, and enter means the words. The walk was already there; enter just never followed it. * style: gofmt internal/tui3/standingenter_test.go Assisted-by: CodeAF (gemini-3.8-flash-high) Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com> * docs: write the change down for 1506 * session: the transcript test settles the chatlog before reading the store back The end-to-end test read brain.Messages straight after the turn's events closed, but the transcript reaches the store through the batching chat log: the read raced the writer and on CI lost, finding an empty thread. Close the log first, the way chatlog_test reads its thread back. * chat: groom the standing card ask and its yes clause The lead said "wants to keep an eye on" on every standing card, which was wrong about most of them: a one-off reminder watches nothing, and a rule that never wakes watches nothing either. The lead now says what a yes binds — "wants to set this up" — for every kind. The yes clause beside the chip promised "it keeps happening until you stop it" on cards whose item runs once at a moment and then retires. The clause now reads the item: at-things get "it happens at the time, and then it retires", and every other kind keeps the old clause. The manual pages that quote the card move with it, and the tests that assert the lead's words say the new words. * docs(changelog): add unreleased entry for PR #1520 * feat(session): plan-born task worker briefs carry standing project orders (#1549) * docs(changes): fix surface tag to chat and engine for #1549 * plandb: check terminal status before ownership check in Done, prune search guidance In Store.Done(), check if task is already terminal before checking requireOwner(). When a task is cancelled upstream, ClaimedBy is cleared or empty, which previously caused Done to fail obscurely with 'task is not claimed'. Checking terminal status first reports an explicit and actionable error: 'task is already terminal (cancelled)'. Also update bashworker prompt guidance to prune heavy directories when running find, preventing workers from freezing across deep trees. Fixes #1561 Assisted-by: CodeAF (gemini-3.8-flash-high) Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com> * fix(tui3): task proposal countdown handles keypress and avoids starting declined task (#1547) * session: fall back to project place workspace when workspace is non-repository When a session operates with workspace outside a repository (e.g. user home folder ~ in a team manager session), task ground derivation would fall through to taskGroundNothing at dir: ~, leading spawned tasks to inherit ~ as ground. Workers looking for repository docs like docs/design/plandb-cli/ would then fail to locate them and launch unpruned find commands across the entire home drive. Fall back to a.config.Place.Workspace when workspace has no repository root, preserving taskGroundStandingIn on the intended project repository. Fixes #1561 Assisted-by: CodeAF (gemini-3.8-flash-high) Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com> * tui3: allow slash command completion to execute bare /task, /redo, and /workspace immediately Fixes #1548 * fix(session): enforce daily spend rail on interactive turns and task delegations (#1546) * docs(changes): add changelog entry for PR 1562 Assisted-by: CodeAF (gemini-3.8-flash-high) Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com> * prompts: restore bashworker prompt to keep prompt size under 16 KiB Assisted-by: CodeAF (gemini-3.8-flash-high) Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com> * fix(teams): manager conversation initialization, handle fallback, and approval posture inheritance (#1551) * docs(changes): state the stale belief as a claim with a full stop for #1549 * fix(session): exempt task workers and errands from interactive rail block (#1546) * jobs: bound the disk spool and let the footer admit the truncation A background job's log used to grow without limit: every byte the job wrote went to one <id>.log forever, a single huge Write grew the in-memory ring by its whole length before trimming, and a failed or short or unclosable spool write was swallowed whole. A watcher that prints for a week filled the disk at whatever rate it printed. The spool is now a window: at most jobSpoolChunks chunks of jobSpoolChunkBytes (4MB) on disk — <id>.log live, <id>.log.1 kept — rotated by rename and never rewritten per write, the chunk beyond the window deleted and counted. One huge Write spools in chunk-sized pieces and hands only its newest 64KB to the ring, so no temporary grows to match it. A failed, short or failed-to-close spool write records one notice, stops the retries, and never stops the drain or kills the job. Every footer a model reads — jobs output and the completion note alike — stays honest about the bound: full log while the file is everything the job wrote, the truncation or the failure named beside the path where it is not. The manual's promises of the whole log, the jobs tool description and PERF.md move with it. Retained output stays addressable by the read tool exactly as before. Aggregate retention across jobs and search exclusion are deliberately not in this change. Issue #1599. Assisted-by: CodeAF Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com> * docs: point 1549 changelog entry at PR 1598 * docs(changelog): correct pr number to 1598 * fix(jobs): preserve spool identity and report incomplete output honestly * fix(jobs): retain bounded completed logs with cross-process ownership * fix(jobs): finish integrating retention and task log diagnostics * fix(jobs): preserve registry identity across workspace changes * docs(jobs): explain unsafe log storage refusal * fix(jobs): preserve safe startup expiry for completed log payloads * fix: exclude runtime output from recursive search and bound log reads * docs: consolidate bounded runtime logs change entry * docs: use plain language for retained log ownership * docs: make completed log retention separately searchable * fix: preserve grep matches when bounded context skips long lines * docs: associate bounded logs change with PR 1603 * fix: keep runtime search guidance within prompt budgets * test: allow one explicit live verification model across roles * test: allow one explicit live verification model across roles * test: pin every terminal model role and expose default launch * fix: resolve project ground on the default run door * fix: tell every plan worker its assigned working directory * fix: bound foreground spills and preserve selected folder aliases * fix(teams): inherit approvals through the session before starting work * docs: align standing explanations with engine-owned card wording * Unify compact chat activity and preserve explicit user updates * fix(session): keep daily guards current across conversations and midnight * Cover explicit interim updates, fragmented markers and steering * docs: use the supported changelog category * docs: align change entry filename with pull request number * docs: align change entry filename with pull request number * Stream explicitly addressed updates immediately across chat views * fix: exclude prompt history and pin every live text role * docs: align foreground change entry with pull request number * Verify task receipts remain available behind disclosure * Keep resumed work in one activity window without stray receipts * fix(tui3): anchor workspace selections through the workspace action * test: wait for run teardown and isolate heartbeat phase checks * Fold asynchronous task housekeeping while preserving requested command results * test: wait for run teardown and isolate heartbeat phase checks * fix(tui3): label workspace picker actions accurately * Reset interim-update probes before textless new responses * Record live terminal acceptance evidence for PR 1607 * test: observe delivered child news after conversation wake * Preserve user-directed notices and close transcript transition gaps * Recognize terminal rail gutter in live answer assertions * Keep harness progress and interrupted work inside shared disclosures * Publish verified final chat cleanup terminal captures * fix(chat): defer independent pool judges in one-model runs * Close interrupted response state when steering is consumed * docs: record single-model pool judge policy for PR 1610 * test: preserve machine gate settings when pinning run models * Account for interleaved captions in Codex redaction acceptance * docs: record PR 1615 test fixture correction * test: preserve machine gate settings when pinning run models * Fold task reply provenance into shared work disclosure * fix(session): join routing cache beat during close * docs: record PR 1616 shutdown join * Record deterministic caption acceptance fix * fix: recover and cancel tasks with durable admission state Addresses the restart and stopped-state portions of #1554, held task cancellation in #1571, and live machine limits in #1579. The separate read-helper fallback item in #1554 is outside this change. * test: wait for program run owner before fixture cleanup * docs: identify focused task recovery PR 1619 * test(session): join closed program run before fixture cleanup * test(session): await run completion after its store closes * fix(standing): preserve shared timer ownership during implicit setup * test(session): await task receipts before asserting claim wait * docs: describe durable audience and interrupted human updates * docs: record shared timer ownership correction * test: preserve machine gate settings when pinning run models * test(session): await task receipts before asserting claim wait * test(session): join closed program run before fixture cleanup * test(session): await run completion after its store closes * test(session): synchronize young bash steering without timing window Fixes #1623. * Preserve assistant audience and interruption across transcript replay * fix(standing): explain unavailable background checks in approval receipt * test(session): observe readings in the handover fixture * test(session): hold the decision turn until resolution * test(session): observe readings in the handover fixture * test(session): hold the decision turn until resolution * test(session): synchronize young bash steering without timing window Fixes #1623. * Keep interrupted human updates visible across shared chat views Preserve explicit audience without confirming partial responses, retain full operational disclosure on replay, and exercise stop/steer/retry transitions across all lenses. Refs #1620. * fix: account for paid responses discarded by reasoning retries * docs: identify retry accounting PR 1625 * docs: retain Escape search terms for interrupted updates * docs(cli): complete task and chat budget help synopses * test: preserve machine gate settings when pinning run models * docs: record CLI help correction in PR 1626 * fix(tui3): give task rooms the keyboard when opened * docs: scope task note guidance to rooms accepting input * docs(cli): fit complete help within existing page limits * test: distinguish interrupted and completed steering collapse * docs: identify task-room keyboard fix as PR 1628 * test(config): isolate credit fixtures from provider credentials (cherry picked from commit 0c757ba6365a2d3d6b61ae7df4b4ba8b4679342f) * test: reuse reviewed run teardown fixtures for clean validation Reuse the existing #1619 orphan-run cleanup and #1622 owner-completion helper already integrated into Santosh/dev. No application code changes; avoids duplicating or masking the known temporary-directory race. * docs: publish sustained chat UX acceptance and validation evidence * test(chat): join first-run judge fixture usage writer * docs: keep review recordings and generated evidence out of source * test(provider): order hedge fixtures before visible primary progress * session: a job's status is final before the log-folder sweep, not after it Since #1603 every job's ending ran a whole retention sweep of its log folder before its state left running, so StopWork's job read as running for 47-130 ms under load (2 ms on dev), and TestStopWorkDoesNotWakeAndOnlyFreshSubmissionRestarts failed 5 of 20 runs. settle now closes the file, publishes the final state, runs the sweep, then closes done, so every waiter on done still joins the sweep. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * manual: the countdown's first quarter-second, and codeaf spelled lowercase The proposal card drops a key pressed in its first quarter-second but moves the pointer to `2 no` (#1594), so an enter straight after it declines; the page now says so. The timer-ownership paragraph said "CodeAF says" in prose. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Revert "session: a job's status is final before the log-folder sweep, not after it" This reverts commit 8b0edfd. Publishing the final status before the sweep let anything that watches running() tear a job's folder down while the sweep was still writing in it: CI failed TestAParkedWorkerIsHandedBackWithARecordBeforeItsWholeAllowance on 'TempDir RemoveAll cleanup: .codeaf/jobs: directory not empty'. The stop latency it addressed comes back to be fixed on dev without that race. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * run: the placeholder-check test passes #1627's standing argument to NewBashWorker #1634 landed on dev with a call written against the three-argument constructor; #1627 gives plan-born workers the project's standing orders through a fourth. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: agentfield-bot <agentfield-bot@users.noreply.github.com> Co-authored-by: CodeAF <267109073+agentfield-bot@users.noreply.github.com> Co-authored-by: Abir Abbas <abirabbas1998@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Actual terminal acceptance on this focused revision
**Tested exact source
25893534bfa9d9077e50b3d23c563f072b9a156c**, with the binary hash and request receipts retained. Two useful workflows passed after real fixture-engine restarts:Bolts,12,3andNuts,7,8;24 tests passQueued cancellation and second-conversation video · Restart, release and supplier handoff video · Active result GIF
The protected-main supplier task intentionally kept its branch; an explicit follow-up brought its three finished files into the checkout. Automatic protected-main landing is not claimed. The active task used an ordinary branch and did merge. Intentional cancellations and a worker deadline remain visible in the active recording; completion succeeded. These are scoped graceful-restart scenarios, not a universal exactly-once guarantee.
The receipt audit also found a separate provider retry-usage undercount: one paid empty-at-ceiling retry attempt costing $0.0005994 was omitted from the aggregate usage ledger. All its requests still used the required model. That billing correction is tracked separately in issue #1624; the53 transport / 52 usage counts above are not a parity claim.
Queued tasks could take a repository lock before they were admitted, ignore cancellation, lose their accepted folder/brief on restart, or retain misleading working state. This focused series persists accepted requests before admission, restores the exact request or working copy through the normal run owner, keeps joined task lifecycle consistent, and honors external machine-limit changes both at gate creation and afterward.
This extracts only the reviewed restart/cancellation/resource-admission work from broad draft #1604. Dependencies: #1562, #1603 and #1610, included as merged reviewed dependency heads. The independently selectable focused series, after those dependencies, is
4e5309fe8(production fixes, manual and regressions),4cef394bd(test owner-cleanup correction), and25893534b(PR change-entry metadata). Dependency merge base is1cacec27f. The broader draft is preserved; unrelated auth, teams, daily-spend, standing orders and audit UI changes are excluded.Addresses #1554's restart/stopped-state items only; its reader-handoff fallback item remains outside this PR. Addresses the resource-admission portion of #1579.
Fixes #1571.
Validation, focused session tests passed in 13.212 seconds and terminal tests in 1.460 seconds; the final focused source builds. Exact-head CI36335451163 on
25893534bfa9d9077e50b3d23c563f072b9a156cpassed all required checks, including the full affected-package suite. The local light gate passed all 30 law packages and local non-heavy affected packages passed; A separate local full-suite run did not finish and is not counted as passing. Live acceptance on this exact head passed as recorded above. Earlier combined221e7 media remains separate evidence and is not substituted for this focused-head run.Limits: older records missing their original brief/folder remain visibly interrupted rather than reconstructed by guesswork; same-copy restart does not promise exactly-once execution of arbitrary external side effects. Started delegated programs retain their existing recovery contract and durable ending receipts.
The separate paid-retry accounting defect #1624 is addressed by PR #1625; its final CI is green and settled live evidence is published there. The original recordings above do not certify that subsequent accounting fix.
Review media below are direct PR attachments. Repository-hosted evidence links have been removed. Exact revisions, validation results and limitations remain in this description; original source recordings and generated examples are retained locally.