Skip to content

fix(standing): preserve shared timer ownership on first approval - #1622

Closed
santoshkumarradha wants to merge 10 commits into
devfrom
codex/standing-timer-ownership
Closed

santoshkumarradha wants to merge 10 commits into
devfrom
codex/standing-timer-ownership

Conversation

@santoshkumarradha

@santoshkumarradha santoshkumarradha commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Real product workflow — recorded on Spark

Goal: Use the pantry CLI to create a shopping report, approve a weekly refresh from a new profile while another profile owns the shared timer, then choose manual use and stop the weekly item.

Observed outcome: The real CLI produced beans buy 3 and rice buy 1. Approval saved the weekly item while leaving both existing shared timer files byte-identical. The model explained that this profile needs a codeaf window open and cannot run after it closes. A natural follow-up retired the item, preserved the report and inventory, and ran all 19 existing tests successfully.

Real CodeAF workflow result

Watch the workflow (MP4) · Animated GIF · Original terminal recording (.cast)

Real CodeAF workflow — speed-adjusted preview

Actual tests, shopping output and stopped weekly item

Evidence: tested revision dddee8c4ef8f91be704ab158647d707bb00b646a; Spark session critical-extended-1618verified-v3. Execution/binary identity, model receipts, checksums and playback details. All 22 recorded calls used OpenRouter deepseek/deepseek-v4.1-flash, including auxiliary calls. Independent tests, timer hashes, inventory and retired-item proof, working report, and coverage and limitations.

Playback and limits: Original .cast is preserved; GIF/MP4 play at 3× speed with idle intervals capped at 3 seconds. Recorded runtime dddee8c; final checkout 7eafed3 differs only in test fixtures. CI is green. Spark light, full cmd/codeaf/standing/manual, and full session checks passed (2 session shards, 303 seconds). This verifies Linux refusal to replace an existing owner, not successful background execution for this profile. Darwin, concurrent claims, canceled locks and explicit takeover use deterministic fake-home regressions. Model Pool was disabled for separately fixed #1608/#1610. An earlier run exposed missing model-facing availability; another accidentally launched an older binary during rebuild. Both are retained locally and excluded from this final acceptance, whose embedded version and binary hash were verified. This does not resolve the coordinating task’s earlier shared-timer cleanup checkpoint.

Approving the first standing item in a new profile could replace another profile’s shared OS timer. Implicit setup now preserves the existing owner, while saving the approved item and explaining that its schedule requires a codeaf window open for this home. Every waking approval carries current availability, including after the first notice; permission rules do not incorrectly require a timer.

One interprocess lock covers the ownership read and all timer mutations. Launch repair rechecks ownership under that lock, avoiding a check/install race. Explicit background-check settings retain deliberate takeover/removal behavior.

Validation on Spark: focused Linux/systemd and Darwin/launchd fake-home tests cover foreign/live owners, explicit takeover, concurrent claims and canceled mutations. Session tests verify successful item saving and truthful first/subsequent receipts. The light gate and every affected package passed: cmd/codeaf, standing, manual, and full session (2 shards, 303 seconds). Final-head CI is green.

Real tmux acceptance used embedded runtime dddee8c, with all 22 model receipts on OpenRouter deepseek/deepseek-v4.1-flash. It built a useful shopping report, approved a weekly refresh without changing either existing shared timer file, explicitly explained the open-window limitation, then stopped the item and ran 19 pantry tests. Independent tests and byte comparisons passed. Final 7eafed3 differs only in test fixtures. Model Pool was off as the disclosed workaround for #1608, separately fixed in #1610. No real shared timer was changed.

Fixes #1618. Related to #360.

@santoshkumarradha santoshkumarradha added bug Something the code does that it should not area:session The engine — turns, tasks, the toolbelt, checkpoints sev:serious Wrong or missing behaviour a person meets in ordinary use labels Sep 27, 2026
@santoshkumarradha santoshkumarradha added this to the Reliable agent milestone Sep 27, 2026
@santoshkumarradha
santoshkumarradha marked this pull request as ready for review September 27, 2026 17:42
santoshkumarradha added a commit that referenced this pull request Sep 27, 2026
# Conflicts:
#	internal/session/task_run_belt_test.go
#	internal/session/task_run_orphan_test.go
@santoshkumarradha

Copy link
Copy Markdown
Member Author

Superseded by the aggregate draft PR #1627. Exact verified source 7eafed3 is integrated into Santosh/dev and verified as its ancestor. The aggregate preserves the live workflow, original recording, 22 exact DeepSeek v4.1 Flash model receipts, timer ownership proof, passing local gates and source CI. This PR is closed without deleting its branch or evidence; no dev/main merge occurred. The earlier shared-timer owner remains unknown, and this task did not restore or change it.

santoshkumarradha added a commit that referenced this pull request Sep 27, 2026
Reuse the existing #1619 orphan-run cleanup and #1622 owner-completion helper already integrated into Santosh/dev. No application code changes; avoids duplicating or masking the known temporary-directory race.
AbirAbbas added a commit that referenced this pull request Sep 28, 2026
* session: the waiting sentence and its sign-in reason land together

A connect offer published the waiting lane before the desk held the
sign-in sentence, so a reader could see that a person is needed with
an empty reason. The row is banked while the agent lock is held, and
the lane is filled before that lock is released.

* changes: note for #1502

* tui3: an air row under the nav — the head grows a row of its own

Port of the five-row head onto dev's conditional-strip architecture:
the strip only draws while a conversation is in front, so the air row
is the head's own. A place spends four rows (nav, air, rule, blank:
placeHeadRows 3→4); a conversation five (the strip back in its place
between the air row and the rule: chatHeadRows = placeHeadRows + 1,
tabStripRow = 2). Design references and the live spark capture ride
along under docs/design/.

* docs(changelog): add unreleased changelog entry for PR #1509

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* session: the transcript keeps the person's words when a turn carries skills (#1504)

The skills block rides the copy the model reads; the record the surfaces
draw keeps the person's own words, and the dim skills-carried line is the
one visible channel. Adds the end-to-end regression test through the real
chat door, and lands the full internal/session suite green: the lane-news
test helper no longer hands an earlier test's held sighting to the first
reader of the next, and the ignored-folder receipt test wants the
canonical spelling the receipt actually prints.

* tui3: enter takes the answer the pointer stands on, on a standing card (#1506)

The standing card's box is its correction lane, and the InputText give-up
read that as the whole question: enter over an empty box did nothing even
with the pointer standing on an answer, and the key table's own reading
offered enter only where the asker recommended something. Both give-ups
now yield to a pointer that stands on an option on a question that has a
pick to take — a choice or a judgement — while a connect key offer keeps
its own law: one answer, the way out, taken by a digit, and enter means
the words. The walk was already there; enter just never followed it.

* style: gofmt internal/tui3/standingenter_test.go

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* docs: write the change down for 1506

* session: the transcript test settles the chatlog before reading the store back

The end-to-end test read brain.Messages straight after the turn's events
closed, but the transcript reaches the store through the batching chat
log: the read raced the writer and on CI lost, finding an empty thread.
Close the log first, the way chatlog_test reads its thread back.

* chat: groom the standing card ask and its yes clause

The lead said "wants to keep an eye on" on every standing card, which was
wrong about most of them: a one-off reminder watches nothing, and a rule that
never wakes watches nothing either. The lead now says what a yes binds —
"wants to set this up" — for every kind.

The yes clause beside the chip promised "it keeps happening until you stop
it" on cards whose item runs once at a moment and then retires. The clause
now reads the item: at-things get "it happens at the time, and then it
retires", and every other kind keeps the old clause.

The manual pages that quote the card move with it, and the tests that assert
the lead's words say the new words.

* docs(changelog): add unreleased entry for PR #1520

* feat(session): plan-born task worker briefs carry standing project orders (#1549)

* docs(changes): fix surface tag to chat and engine for #1549

* plandb: check terminal status before ownership check in Done, prune search guidance

In Store.Done(), check if task is already terminal before checking requireOwner().
When a task is cancelled upstream, ClaimedBy is cleared or empty, which previously
caused Done to fail obscurely with 'task is not claimed'. Checking terminal status
first reports an explicit and actionable error: 'task is already terminal (cancelled)'.

Also update bashworker prompt guidance to prune heavy directories when running find,
preventing workers from freezing across deep trees.

Fixes #1561

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* fix(tui3): task proposal countdown handles keypress and avoids starting declined task (#1547)

* session: fall back to project place workspace when workspace is non-repository

When a session operates with workspace outside a repository (e.g. user home
folder ~ in a team manager session), task ground derivation would fall through
to taskGroundNothing at dir: ~, leading spawned tasks to inherit ~ as ground.
Workers looking for repository docs like docs/design/plandb-cli/ would then
fail to locate them and launch unpruned find commands across the entire home drive.

Fall back to a.config.Place.Workspace when workspace has no repository root,
preserving taskGroundStandingIn on the intended project repository.

Fixes #1561

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* tui3: allow slash command completion to execute bare /task, /redo, and /workspace immediately

Fixes #1548

* fix(session): enforce daily spend rail on interactive turns and task delegations (#1546)

* docs(changes): add changelog entry for PR 1562

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* prompts: restore bashworker prompt to keep prompt size under 16 KiB

Assisted-by: CodeAF (gemini-3.8-flash-high)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* fix(teams): manager conversation initialization, handle fallback, and approval posture inheritance (#1551)

* docs(changes): state the stale belief as a claim with a full stop for #1549

* fix(session): exempt task workers and errands from interactive rail block (#1546)

* jobs: bound the disk spool and let the footer admit the truncation

A background job's log used to grow without limit: every byte the job
wrote went to one <id>.log forever, a single huge Write grew the
in-memory ring by its whole length before trimming, and a failed or
short or unclosable spool write was swallowed whole. A watcher that
prints for a week filled the disk at whatever rate it printed.

The spool is now a window: at most jobSpoolChunks chunks of
jobSpoolChunkBytes (4MB) on disk — <id>.log live, <id>.log.1 kept —
rotated by rename and never rewritten per write, the chunk beyond the
window deleted and counted. One huge Write spools in chunk-sized
pieces and hands only its newest 64KB to the ring, so no temporary
grows to match it. A failed, short or failed-to-close spool write
records one notice, stops the retries, and never stops the drain or
kills the job.

Every footer a model reads — jobs output and the completion note alike
— stays honest about the bound: full log while the file is everything
the job wrote, the truncation or the failure named beside the path
where it is not. The manual's promises of the whole log, the jobs tool
description and PERF.md move with it.

Retained output stays addressable by the read tool exactly as before.
Aggregate retention across jobs and search exclusion are deliberately
not in this change. Issue #1599.

Assisted-by: CodeAF
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>

* docs: point 1549 changelog entry at PR 1598

* docs(changelog): correct pr number to 1598

* fix(jobs): preserve spool identity and report incomplete output honestly

* fix(jobs): retain bounded completed logs with cross-process ownership

* fix(jobs): finish integrating retention and task log diagnostics

* fix(jobs): preserve registry identity across workspace changes

* docs(jobs): explain unsafe log storage refusal

* fix(jobs): preserve safe startup expiry for completed log payloads

* fix: exclude runtime output from recursive search and bound log reads

* docs: consolidate bounded runtime logs change entry

* docs: use plain language for retained log ownership

* docs: make completed log retention separately searchable

* fix: preserve grep matches when bounded context skips long lines

* docs: associate bounded logs change with PR 1603

* fix: keep runtime search guidance within prompt budgets

* test: allow one explicit live verification model across roles

* test: allow one explicit live verification model across roles

* test: pin every terminal model role and expose default launch

* fix: resolve project ground on the default run door

* fix: tell every plan worker its assigned working directory

* fix: bound foreground spills and preserve selected folder aliases

* fix(teams): inherit approvals through the session before starting work

* docs: align standing explanations with engine-owned card wording

* Unify compact chat activity and preserve explicit user updates

* fix(session): keep daily guards current across conversations and midnight

* Cover explicit interim updates, fragmented markers and steering

* docs: use the supported changelog category

* docs: align change entry filename with pull request number

* docs: align change entry filename with pull request number

* Stream explicitly addressed updates immediately across chat views

* fix: exclude prompt history and pin every live text role

* docs: align foreground change entry with pull request number

* Verify task receipts remain available behind disclosure

* Keep resumed work in one activity window without stray receipts

* fix(tui3): anchor workspace selections through the workspace action

* test: wait for run teardown and isolate heartbeat phase checks

* Fold asynchronous task housekeeping while preserving requested command results

* test: wait for run teardown and isolate heartbeat phase checks

* fix(tui3): label workspace picker actions accurately

* Reset interim-update probes before textless new responses

* Record live terminal acceptance evidence for PR 1607

* test: observe delivered child news after conversation wake

* Preserve user-directed notices and close transcript transition gaps

* Recognize terminal rail gutter in live answer assertions

* Keep harness progress and interrupted work inside shared disclosures

* Publish verified final chat cleanup terminal captures

* fix(chat): defer independent pool judges in one-model runs

* Close interrupted response state when steering is consumed

* docs: record single-model pool judge policy for PR 1610

* test: preserve machine gate settings when pinning run models

* Account for interleaved captions in Codex redaction acceptance

* docs: record PR 1615 test fixture correction

* test: preserve machine gate settings when pinning run models

* Fold task reply provenance into shared work disclosure

* fix(session): join routing cache beat during close

* docs: record PR 1616 shutdown join

* Record deterministic caption acceptance fix

* fix: recover and cancel tasks with durable admission state

Addresses the restart and stopped-state portions of #1554, held task cancellation in #1571, and live machine limits in #1579. The separate read-helper fallback item in #1554 is outside this change.

* test: wait for program run owner before fixture cleanup

* docs: identify focused task recovery PR 1619

* test(session): join closed program run before fixture cleanup

* test(session): await run completion after its store closes

* fix(standing): preserve shared timer ownership during implicit setup

* test(session): await task receipts before asserting claim wait

* docs: describe durable audience and interrupted human updates

* docs: record shared timer ownership correction

* test: preserve machine gate settings when pinning run models

* test(session): await task receipts before asserting claim wait

* test(session): join closed program run before fixture cleanup

* test(session): await run completion after its store closes

* test(session): synchronize young bash steering without timing window

Fixes #1623.

* Preserve assistant audience and interruption across transcript replay

* fix(standing): explain unavailable background checks in approval receipt

* test(session): observe readings in the handover fixture

* test(session): hold the decision turn until resolution

* test(session): observe readings in the handover fixture

* test(session): hold the decision turn until resolution

* test(session): synchronize young bash steering without timing window

Fixes #1623.

* Keep interrupted human updates visible across shared chat views

Preserve explicit audience without confirming partial responses, retain full operational disclosure on replay, and exercise stop/steer/retry transitions across all lenses. Refs #1620.

* fix: account for paid responses discarded by reasoning retries

* docs: identify retry accounting PR 1625

* docs: retain Escape search terms for interrupted updates

* docs(cli): complete task and chat budget help synopses

* test: preserve machine gate settings when pinning run models

* docs: record CLI help correction in PR 1626

* fix(tui3): give task rooms the keyboard when opened

* docs: scope task note guidance to rooms accepting input

* docs(cli): fit complete help within existing page limits

* test: distinguish interrupted and completed steering collapse

* docs: identify task-room keyboard fix as PR 1628

* test(config): isolate credit fixtures from provider credentials

(cherry picked from commit 0c757ba6365a2d3d6b61ae7df4b4ba8b4679342f)

* test: reuse reviewed run teardown fixtures for clean validation

Reuse the existing #1619 orphan-run cleanup and #1622 owner-completion helper already integrated into Santosh/dev. No application code changes; avoids duplicating or masking the known temporary-directory race.

* docs: publish sustained chat UX acceptance and validation evidence

* test(chat): join first-run judge fixture usage writer

* docs: keep review recordings and generated evidence out of source

* test(provider): order hedge fixtures before visible primary progress

* session: a job's status is final before the log-folder sweep, not after it

Since #1603 every job's ending ran a whole retention sweep of its log folder
before its state left running, so StopWork's job read as running for 47-130 ms
under load (2 ms on dev), and TestStopWorkDoesNotWakeAndOnlyFreshSubmissionRestarts
failed 5 of 20 runs. settle now closes the file, publishes the final state,
runs the sweep, then closes done, so every waiter on done still joins the sweep.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* manual: the countdown's first quarter-second, and codeaf spelled lowercase

The proposal card drops a key pressed in its first quarter-second but moves
the pointer to `2 no` (#1594), so an enter straight after it declines; the
page now says so. The timer-ownership paragraph said "CodeAF says" in prose.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Revert "session: a job's status is final before the log-folder sweep, not after it"

This reverts commit 8b0edfd. Publishing the final status before the sweep let
anything that watches running() tear a job's folder down while the sweep was
still writing in it: CI failed TestAParkedWorkerIsHandedBackWithARecordBeforeItsWholeAllowance
on 'TempDir RemoveAll cleanup: .codeaf/jobs: directory not empty'. The stop
latency it addressed comes back to be fixed on dev without that race.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* run: the placeholder-check test passes #1627's standing argument to NewBashWorker

#1634 landed on dev with a call written against the three-argument constructor;
#1627 gives plan-born workers the project's standing orders through a fourth.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: agentfield-bot <agentfield-bot@users.noreply.github.com>
Co-authored-by: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
Co-authored-by: Abir Abbas <abirabbas1998@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:session The engine — turns, tasks, the toolbelt, checkpoints bug Something the code does that it should not sev:serious Wrong or missing behaviour a person meets in ordinary use

Projects

None yet

Development

Successfully merging this pull request may close these issues.

First standing approval can take another profile’s shared background timer chat: --one-model still calls an independent Model Pool judge

1 participant