[CL-352] Upgrade Paperclip до v2026.720.0 #353
Open
andrei
wants to merge 1 commit from
upgrade/paperclip-v2026.720.0 into master
pull from: upgrade/paperclip-v2026.720.0
merge into: europa-tech-srl:master
europa-tech-srl:master
europa-tech-srl:ceo/pix-11114-gitnexus-metadata
europa-tech-srl:ceo/pix-15506-api-write-500-fix
europa-tech-srl:ceo/pix-15506-api-500-fix
europa-tech-srl:ceo/pix-15497-dependency-security-updates
europa-tech-srl:fix/pix-15500-lockfile-policy
europa-tech-srl:chore/sync-master-ahead-pix-15377
europa-tech-srl:fix/heartbeat-review-wait-retry-scheduling
europa-tech-srl:agent/devops/pix-15345
europa-tech-srl:chore/sync-master-ahead-pix-15193
europa-tech-srl:fix/pix-15110-sentry-production-projects
europa-tech-srl:fix/quota-window-health-json-output
europa-tech-srl:agent/fullstack/pix-14770
europa-tech-srl:agent/fullstack/pix-14729
europa-tech-srl:agent/devops/pix-14664
europa-tech-srl:fix/pm2-waiting-restart-watchdog
europa-tech-srl:agent/devops/pix-14323
europa-tech-srl:fix/sentry-autonomous-incident-routing
europa-tech-srl:fix/sentry-canonical-stack-gate
europa-tech-srl:fix/pix-11114-suppress-loop-success
europa-tech-srl:chore/pix-11114-gitnexus-metadata
europa-tech-srl:agent/devops/pix-14274
europa-tech-srl:pix-14255-auto-prompt-check-speed
europa-tech-srl:agent/cto/pix-14252-codex-config
europa-tech-srl:agent/fullstack/pix-13817
europa-tech-srl:agent/devops/pix-14192
europa-tech-srl:agent/devops/pix-14184-quota-windows-health
europa-tech-srl:fix/watchdog-ceo-triggerdetail-system
europa-tech-srl:fix/watchdog-wake-ceo-autofix-20260717
europa-tech-srl:agent/fullstack/pix-13215
europa-tech-srl:fix/pix-14170-timer-coalesces
europa-tech-srl:fix/heartbeat-watchdog-interval-aware-20260717
europa-tech-srl:agent/devops/pix-13719
europa-tech-srl:hotfix/paperclip-runtime-exports-20260708
europa-tech-srl:fullstack/pix-13367-rebase
europa-tech-srl:fix/shutdown-retry-chains
europa-tech-srl:fix/server-vitest-config-root
europa-tech-srl:fix/pix-14129-master-drift-delivery
europa-tech-srl:fix/wake-prompt-payload-cap
europa-tech-srl:fix/codex-worker-prompt-compaction
europa-tech-srl:fix/sweeper-smoke-alert-guard-20260716
europa-tech-srl:agent/cto/pix-14128-clean
europa-tech-srl:agent/cto/pix-14128
europa-tech-srl:agent/fullstack/pix-13959
europa-tech-srl:fix/suppress-ceo-loop-no-action-20260716
europa-tech-srl:fix/master-drift-terminal-run-cancel
europa-tech-srl:agent/devops/pix-14069
europa-tech-srl:agent/fullstack/pix-14084
europa-tech-srl:agent/devops/pix-14027
europa-tech-srl:agent/fullstack/pix-13996
europa-tech-srl:agent/cto/pix-13983
europa-tech-srl:fix/git-delivery-gate-costs-telemetry
europa-tech-srl:agent/fullstack/pix-13907-pr304
europa-tech-srl:agent/devops/pix-13910
europa-tech-srl:fix/watchdog-ownership-timeout-delivery
europa-tech-srl:fix/watchdog-ownership-timeout
europa-tech-srl:agent/fullstack/pix-13894
europa-tech-srl:agent/devops/pix-13864
europa-tech-srl:agent/fullstack/pix-13777
europa-tech-srl:agent/fullstack/pix-13762
europa-tech-srl:agent/fullstack/pix-13661
europa-tech-srl:fix/codex-retry-attempt-promotion
europa-tech-srl:fix/suppress-git-delivery-handoff-20260715
europa-tech-srl:agent/cto/pix-13838
europa-tech-srl:agent/devops/pix-13841
europa-tech-srl:fix/preserve-valid-queued-locks-20260715
europa-tech-srl:fix/codex-auth-lock-wait-heartbeat-20260715
europa-tech-srl:fix/control-plane-stale-execution-locks
europa-tech-srl:agent/cto/pix-13774-forward-fix
europa-tech-srl:fix/paperclip-pre-spawn-token-burn-20260715
europa-tech-srl:agent/cto/pix-13791
europa-tech-srl:fix/control-plane-watchdog-branch-drift
europa-tech-srl:agent/cto/pix-13774
europa-tech-srl:agent/cto/pix-13773-delivery
europa-tech-srl:fix/scheduled-retry-watchdog-active-runs
europa-tech-srl:ops/paperclip-auth-holds-active-adapters-20260714
europa-tech-srl:agent/fullstack/pix-13767
europa-tech-srl:fix/paperclip-forever-control-plane-guardrails-20260715
europa-tech-srl:agent/devops/pix-13736
europa-tech-srl:fix/paperclip-agent-handoff-report-hygiene-20260714
europa-tech-srl:fix/codex-runtime-guard-master
europa-tech-srl:fix/codex-output-monitor-stderr
europa-tech-srl:agent/devops/pix-13720
europa-tech-srl:ops/paperclip-autonomy-refill-20260714
europa-tech-srl:agent/cto/pix-13695
europa-tech-srl:ops/paperclip-runtime-stabilize-20260714
europa-tech-srl:fix/codex-cli-global-flags
europa-tech-srl:operator/pix-13685-inherited-claude-env
europa-tech-srl:operator/pix-13671-control-plane-drift
europa-tech-srl:operator/pix-13662-branch-drift
europa-tech-srl:agent/fullstack/pix-13641
europa-tech-srl:agent/devops/pix-13632
europa-tech-srl:agent/fullstack/pix-13620
europa-tech-srl:agent/devops/pix-13612
europa-tech-srl:agent/cto/pix-13608
europa-tech-srl:agent/fullstack/pix-13596
europa-tech-srl:agent/devops/pix-13336
europa-tech-srl:agent/ceo/pix-11114-forgejo-pr-files-fallback
europa-tech-srl:agent/fullstack/pix-13497
europa-tech-srl:agent/cto/pix-13525
europa-tech-srl:agent/ceo/pix-11114-lightrag-policy-skip
europa-tech-srl:agent/fullstack/pix-13521
europa-tech-srl:agent/cto/pix-13349
europa-tech-srl:fix/watchdog-pm2-hygiene-20260710
europa-tech-srl:agent/fullstack/pix-13437
europa-tech-srl:agent/cto/pix-13457
europa-tech-srl:operator/control-plane-api-autorecover
europa-tech-srl:agent/cto/pix-13391
europa-tech-srl:agent/cto/pix-13423
europa-tech-srl:agent/cto/pix-13419
europa-tech-srl:agent/cto/pix-13388
europa-tech-srl:operator/model-profile-schema-fix
europa-tech-srl:agent/devops/pix-13422
europa-tech-srl:agent/devops/pix-13361
europa-tech-srl:agent/devops/pix-13396-sync
europa-tech-srl:agent/cto/pix-13383
europa-tech-srl:agent/fullstack/pix-13355
europa-tech-srl:agent/devops/pix-13357-gitignore-agent-manifest
europa-tech-srl:agent/fullstack/pix-13338
europa-tech-srl:agent/fullstack/pix-13328
europa-tech-srl:agent/devops/pix-13334
europa-tech-srl:agent/cto/pix-13326
europa-tech-srl:agent/cto/pix-13273
europa-tech-srl:agent/fullstack/pix-13312
europa-tech-srl:agent/devops/pix-13299
europa-tech-srl:agent/cto/pix-13272
europa-tech-srl:agent/devops/pix-13286
europa-tech-srl:agent/fullstack/pix-13268
europa-tech-srl:agent/cto/pix-13257
europa-tech-srl:agent/fullstack/pix-13264
europa-tech-srl:agent/fullstack/pix-13260
europa-tech-srl:agent/fullstack/pix-13247
europa-tech-srl:agent/devops/pix-13236
europa-tech-srl:recovery/pix-13213-heartbeat-fixes
europa-tech-srl:agent/devops/pix-13213-delivery
europa-tech-srl:delivery/pix-12637
europa-tech-srl:fix/approval-answer-label-fix
europa-tech-srl:delivery/pix-13034
europa-tech-srl:delivery/pix-13142
europa-tech-srl:agent/devops/pix-13081-guard-delivery
europa-tech-srl:agent/fullstack/pix-13186
europa-tech-srl:agent/fullstack/pix-13174
europa-tech-srl:agent/fullstack/pix-13144
europa-tech-srl:agent/fullstack/pix-13145
europa-tech-srl:agent/devops/pix-13136
europa-tech-srl:agent/fullstack/pix-13130
europa-tech-srl:feature/claude-issue-comment-wakeup-unhandled-rejection
europa-tech-srl:agent/devops/pix-13125
europa-tech-srl:feature/claude-fix-flaky-issue-list-limit
europa-tech-srl:agent/devops/pix-13092
europa-tech-srl:agent/devops/pix-13081
europa-tech-srl:agent/cto/pix-13064
europa-tech-srl:agent/fullstack/pix-13045
europa-tech-srl:agent/cto/pix-13590-fix-master-ci-issue-list-mocks
europa-tech-srl:agent/fullstack/pix-13034
europa-tech-srl:agent/fullstack/pix-13019
europa-tech-srl:agent/cto/pix-13030
europa-tech-srl:agent/fullstack/pix-13011
europa-tech-srl:agent/cto/pix-12986-terminal-disposition-fix
europa-tech-srl:merge-queue/pix-13034-ci-api-job
europa-tech-srl:merge/pix-12937-batch2
europa-tech-srl:agent/devops/pix-12872
europa-tech-srl:agent/cto/pix-12884-cherry-pick-forgejo-sweeper
europa-tech-srl:fix/pix-12885-pr-gates-issue-scope
europa-tech-srl:agent/cto/pix-12887
europa-tech-srl:agent/europatech-hermes-agent/pix-11114
europa-tech-srl:agent/devops/pix-12840
europa-tech-srl:agent/devops/approvalbot-compound-issue
europa-tech-srl:agent/cto/pix-12778
europa-tech-srl:agent/devops/pix-12774
europa-tech-srl:fix/issues-execution-state-read-response
europa-tech-srl:fix/pix-12752-timer-heartbeat-checked
europa-tech-srl:agent/cto/pix-12749
europa-tech-srl:agent/devops/pix-13081-guard
europa-tech-srl:agent/cto/pix-12726
europa-tech-srl:agent/cto/pix-12674-heartbeat-shutdown-grace
europa-tech-srl:agent/fullstack/pix-12684
europa-tech-srl:agent/devops/pix-12641
europa-tech-srl:agent/devops/pix-13883
europa-tech-srl:agent/devops/pix-13781
europa-tech-srl:fix/cl-watchdog-flock-pathfix
europa-tech-srl:agent/devops/pix-13218
europa-tech-srl:fix/claude-watchdog-argparse
europa-tech-srl:agent/devops/pix-13138
europa-tech-srl:fix/claude-daemon-audit
europa-tech-srl:agent/fullstack/pix-13102
europa-tech-srl:fix/telegram-interaction-replies
europa-tech-srl:feature/claude-agents-lockfile-sync
europa-tech-srl:fix/ci-lockfile-mismatch
europa-tech-srl:agent/devops/pix-13039
europa-tech-srl:agent/fullstack/pix-13055
europa-tech-srl:agent/fullstack/pix-12998
europa-tech-srl:agent/fullstack/pix-12993
europa-tech-srl:agent/fullstack/pix-12985
europa-tech-srl:fix/pagination-response-format
europa-tech-srl:fix/pix-12761-cancel-done-runs
europa-tech-srl:review-pix12826
europa-tech-srl:agent/fullstack/pix-12826
europa-tech-srl:fix/pix-12815-bounded-workspace-api
europa-tech-srl:fix/heartbeat-workspace-policy
europa-tech-srl:agent/devops/pix-12759
europa-tech-srl:agent/cto/pix-12761
europa-tech-srl:agent/fullstack/pix-12765
europa-tech-srl:fix/pix-12751-done-target-live-run
europa-tech-srl:agent/devops/pix-12740
europa-tech-srl:agent/fullstack/pix-12708
europa-tech-srl:fix/pix-12732-terminal-live-runs
europa-tech-srl:agent/fullstack/pix-12682
europa-tech-srl:agent/devops/pix-12667
europa-tech-srl:agent/devops/pix-10781
europa-tech-srl:pix-10781-fix-merged
europa-tech-srl:agent/devops/pix-11070
europa-tech-srl:agent/cto/pix-12560
europa-tech-srl:agent/fullstack/pix-12103
europa-tech-srl:agent/devops/pix-11100
europa-tech-srl:pix-11070-remote
europa-tech-srl:codex/paperclip-memory-system-loopback-proxy-recovered
europa-tech-srl:agent/fullstack/pix-12212
europa-tech-srl:agent/fullstack/pix-12181
europa-tech-srl:agent/fullstack/pix-12095
europa-tech-srl:agent/cto/pix-12065
europa-tech-srl:fix/agents-lockfile-overrides-sync
europa-tech-srl:agent/cto/pix-11803
europa-tech-srl:agent/fullstack/pix-11638
europa-tech-srl:agent/fullstack/pix-11596
europa-tech-srl:agent/sentry/pix-11612
europa-tech-srl:agent/sentry/pix-11604
europa-tech-srl:agent/sentry/pix-11581
europa-tech-srl:agent/devops/pix-11564
europa-tech-srl:agent/fullstack/pix-11528
europa-tech-srl:agent/fullstack/pix-11433
europa-tech-srl:agent/devops/pix-11396
europa-tech-srl:agent/devops/pix-11376
europa-tech-srl:agent/fullstack/pix-11294
europa-tech-srl:agent/devops/pix-10908
europa-tech-srl:fix/PIX-7666-codex-flags-v2
europa-tech-srl:agent/devops/pix-10782
europa-tech-srl:agent/devops/pix-10772
europa-tech-srl:fix/pix-10731-workspace-closed-consistency
europa-tech-srl:agent/devops/pix-10734
europa-tech-srl:agent/devops/pix-10721
europa-tech-srl:agent/devops/pix-11143
europa-tech-srl:agent/cto/pix-12251
europa-tech-srl:agent/cto/pix-12286
europa-tech-srl:agent/cto/pix-12433
europa-tech-srl:agent/cto/pix-12587
europa-tech-srl:agent/cto/pix-12897
europa-tech-srl:agent/cto/pix-12909
europa-tech-srl:agent/cto/pix-13478
europa-tech-srl:agent/devops/pix-12910
europa-tech-srl:agent/devops/pix-12914
europa-tech-srl:agent/devops/pix-12941
europa-tech-srl:agent/devops/pix-12942
europa-tech-srl:agent/devops/pix-12954
europa-tech-srl:agent/devops/pix-12957
europa-tech-srl:agent/devops/pix-12961
europa-tech-srl:agent/devops/pix-12962
europa-tech-srl:agent/devops/pix-12980
europa-tech-srl:agent/devops/pix-12984
europa-tech-srl:agent/devops/pix-12988
europa-tech-srl:agent/devops/pix-12991
europa-tech-srl:agent/devops/pix-13010
europa-tech-srl:agent/devops/pix-13108
europa-tech-srl:agent/devops/pix-13123
europa-tech-srl:agent/devops/pix-13164
europa-tech-srl:agent/devops/pix-13167
europa-tech-srl:agent/devops/pix-13175
europa-tech-srl:agent/devops/pix-13177
europa-tech-srl:agent/devops/pix-13182
europa-tech-srl:agent/devops/pix-13188
europa-tech-srl:agent/devops/pix-13193
europa-tech-srl:agent/devops/pix-13196
europa-tech-srl:agent/devops/pix-13252
europa-tech-srl:agent/devops/pix-13267
europa-tech-srl:agent/devops/pix-13274
europa-tech-srl:agent/devops/pix-13280
europa-tech-srl:agent/devops/pix-13288
europa-tech-srl:agent/devops/pix-13295
europa-tech-srl:agent/devops/pix-13300
europa-tech-srl:agent/devops/pix-13306
europa-tech-srl:agent/devops/pix-13308
europa-tech-srl:agent/devops/pix-13312
europa-tech-srl:agent/devops/pix-13362
europa-tech-srl:agent/devops/pix-13369
europa-tech-srl:agent/devops/pix-13385
europa-tech-srl:agent/devops/pix-13417
europa-tech-srl:agent/devops/pix-13436
europa-tech-srl:agent/devops/pix-13439
europa-tech-srl:agent/devops/pix-13450
europa-tech-srl:agent/devops/pix-13452
europa-tech-srl:agent/devops/pix-13471
europa-tech-srl:agent/fullstack/pix-12313
europa-tech-srl:agent/fullstack/pix-12651
europa-tech-srl:deploy/paperclip-v2026.618.0-20260619
europa-tech-srl:fix/PIX-7823-v2-local
europa-tech-srl:fix/PIX-7823-v2
europa-tech-srl:fix/pix-7555-benign-cancel-full
europa-tech-srl:feature/claude-recovery-spread
europa-tech-srl:reconcile/prod-sync-20260615
europa-tech-srl:feature/claude-stranded-recovery-autoclose
europa-tech-srl:fix/PIX-8194-productivity-delivery-signal-015229
europa-tech-srl:fix/run-liveness-terminal-cancelled
europa-tech-srl:fix/PIX-7723-v2-git-status-gate
europa-tech-srl:fix/PIX-8262-transient-retry-escalation
europa-tech-srl:fix/PIX-8157-transient-productivity-review
europa-tech-srl:fix/PIX-8111-scheduled-retry-blocker
europa-tech-srl:fix/PIX-7723-done-gate-non-code-sla
europa-tech-srl:fix/PIX-7823-budget-guardrail-normalization
europa-tech-srl:fix/PIX-7814-regulatory-template-claims
europa-tech-srl:fix/PIX-7831-queued-retry-drain
europa-tech-srl:fix/PIX-7733-cost-attribution
europa-tech-srl:fix/PIX-7595-costs-by-agent-usage
europa-tech-srl:fix/PIX-7666-codex-recovery-flags-clean
europa-tech-srl:fix/PIX-7714-lightrag-maintenance
europa-tech-srl:fix/pix-7662-metered-codex-cost
europa-tech-srl:fix/pix-7684-superseded-done-gate-clean
europa-tech-srl:fix/pix-7684-superseded-done-gate
europa-tech-srl:fix/PIX-7666-codex-recovery-flags
europa-tech-srl:fix/pix-7473-costs-current-month
europa-tech-srl:fix/pix-7560-support-sla-done-gate
europa-tech-srl:fix/pix-7632-codex-search-fallback
europa-tech-srl:fix/pix-7555-benign-cancelled-snapshots
europa-tech-srl:fix/pix-7573-cfo-routine-gate
europa-tech-srl:fix/pix-7496-forgejo-ci-devdeps
europa-tech-srl:fix/pix-7497-code-task-gate
europa-tech-srl:fix/agent-deliverable-doc-tg
europa-tech-srl:fix/pix-7437-recovery-blocked-path
europa-tech-srl:fix/pix-7245-codex-search-fallback
europa-tech-srl:agent-cto/pix-7481-git-delivery-gate
europa-tech-srl:fix/pix-7435-research-closeout-gate
europa-tech-srl:fix/pix-7433-budget-gate
europa-tech-srl:agent-cto/pix-7374-agents
europa-tech-srl:fix/pix-7277-git-delivery-fail-closed
europa-tech-srl:feature/claude-pix-7086-impl
europa-tech-srl:feature/claude-pix-7086-done-gate
europa-tech-srl:chore/reconcile-source-drift
europa-tech-srl:archive/master-pre-reconcile-52b4555
europa-tech-srl:europa-v2
europa-tech-srl:cleanup/codex-git-junk-20260424-035731
europa-tech-srl:refactor/ns-modules
europa-tech-srl:europa
europa-tech-srl:feature/observability-dashboard
No reviewers
Labels
Clear labels
No items
No labels
Milestone
Clear milestone
No items
No milestone
Projects
Clear projects
No items
No project
Assignees
Clear assignees
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".
No due date set.
Dependencies
No dependencies set.
Reference
europa-tech-srl/europa-tech-agents!353
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "upgrade/paperclip-v2026.720.0"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Thinking Path
Upstream
v2026.720.0не имеет общей Git-истории с текущим fork после прошлых tree-based upgrades. Для корректной трёхсторонней интеграции создан временный synthetic merge-base от последнего интегрированного upstreamv2026.626.0, затем выполнен merge нового upstream поверх актуального fork.Локальные миграции
0125–0139сохранены. Новые upstream-миграции перенумерованы в0140–0194, журнал Drizzle пересобран без коллизий. Локальные Hermes/daemon/watchdog/Telegram policies сохранены; конфликтующие core-файлы взяты из нового upstream и адаптированы под локальные регрессии.Fixes #352
What Changed
v2026.720.0;sourceTrustpersistence;Verification
pnpm install --no-frozen-lockfilepnpm run buildpnpm run typecheck662 passed24 passed58 passedpnpm test:run— выполняется локально и Forgejo CI195уникальных записей, последняя0194_decision_training_retention_policygit diff --checkRisks / Rollback
rollback/paperclip-v2026.720.0-*+ pre-upgrade gzip PostgreSQL dump.Model Used
gpt-5.6-solChecklist
**Issue (described inline; no existing tracking issue):** Pausing an agent is not durable. Pausing cancels the in-flight run, but a queued or recovery-dispatched run can clobber the agent back to `running` because the execution-start status update is unconditional — so a "paused" agent silently resumes work while `paused_at` is still set. ## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies; agent work executes as "runs" tracked in `heartbeat_runs`, with a recovery/automation layer that re-dispatches work when a run disappears. > - The agent lifecycle has a pause control (status `paused`, `paused_at` set) meant to stop an agent from taking or continuing work. > - The problem: pause is not durable. Pausing cancels the in-flight run, but the execution-start path then sets `agents.status = 'running'` with an unconditional `UPDATE ... WHERE id = ?`, so any queued or recovery-dispatched run can clobber the paused agent back to `running` and execute. > - Why it matters: a "paused" agent silently resuming undermines the core operational control operators rely on to halt runaway, cost-sensitive, or unsafe work. > - This pull request guards the execution-start status flip with an atomic conditional UPDATE, and tags pause-cancellations for observability without changing resume behaviour. > - The benefit is that a paused agent can no longer transition back to `running`; queued/recovery-dispatched runs are cancelled cleanly instead of clobbering status, while un-pausing still resumes in-flight work. ## What Changed - Execution-start guard: replaced the unconditional `UPDATE agents SET status='running' WHERE id = ?` with an atomic conditional `UPDATE ... WHERE id = ? AND status NOT IN ('paused','terminated','pending_approval')`. On a zero-row match the run is cancelled (`errorCode: "agent_not_invokable"`), the issue execution lock is released, and the path returns — instead of clobbering status. - Exported `DIRECT_NON_INVOKABLE_STATUSES` from `agent-invokability.ts` and reused it in `heartbeat.ts` as the single source of truth for the guard. - Pause observability: `cancelActiveForAgentInternal` now accepts an `errorCode` (default `"cancelled"`); the pause-route wrapper `cancelActiveForAgent` passes `"agent_paused"`. This is classification-neutral — `agent_paused` is NOT added to `NON_RETRYABLE_CONTINUATION_ERROR_CODES`, so on un-pause the issue's continuation re-enqueues and work resumes. - Exported `classifyContinuationFailure` from `recovery/service.ts` for unit testing (no logic change). - Added `server/src/services/recovery/service.pause-durability.test.ts` covering continuation classification. ## Verification - `pnpm --filter @paperclipai/server typecheck` — clean. - `pnpm exec vitest run server/src/services/recovery/service.pause-durability.test.ts` — 5 passed. - `pnpm exec vitest run server` — full server suite passes locally. The only failures are pre-existing and environment-specific, unrelated to this change (a git default-branch test fixture, and a known checkout-lock race) — both reproduce identically on clean `master` with this change stashed out. - Behavioural: a paused agent's execution-start now aborts cleanly with no status clobber; non-pause cancellations keep `errorCode "cancelled"` and existing behaviour; un-pausing resumes the in-flight issue. ## Risks - Low risk. No schema change, no migration, no new dependency; four files. The change narrows a single UPDATE to be conditional and adds a rarely-taken abort branch on the run-start path; behaviour for invokable agents is unchanged. The abort's `agent_not_invokable` code is already in `NON_RETRYABLE_CONTINUATION_ERROR_CODES`. The only caller of `cancelActiveForAgent` is the pause route. ## Related upstream work — not duplicates This area has prior and in-flight PRs; #8317 was checked against them and is intentionally distinct: - **#4503** (`fix(heartbeat): make agent pause status guard atomic with status update`) targets a different TOCTOU race on the **post-run / finalize** path (`finalizeAgentStatus`). #8317 targets the **execution-start** race, where a recovery-dispatched run flips a paused agent back to `running` *before the run begins*. #4503 does not cover the proven failure path here: `pause → active run cancelled → recovery dispatches a new run → execution-start overwrites the paused state`. #4503 also does not add the resume semantics below. - **#4356** (`honor system/manual/auto pause at all heartbeat-run enqueue sites`) and **#1067** (`pause guard on queue drain`) protect the **enqueue / queue-drain** layer. They are complementary to — not substitutes for — the execution-start guard, which is the last gate before a run actually starts. - **#6944** (`guard executeRun against paused agent`), **#7140**, and **#7141** attempted similar execution-start ideas but were closed for implementation hygiene / build issues, not because the guard concept was wrong. #8317 implements that concept cleanly: a single atomic conditional UPDATE, a clean abort with `errorCode: "agent_not_invokable"`, a shared `DIRECT_NON_INVOKABLE_STATUSES` source of truth, and passing tests + typecheck. Intentional, minor difference (not a criticism of #4503): #8317's execution-start deny-list is `paused`, `terminated`, and `pending_approval` — the full non-invokable set for run-start invokability — whereas #4503 appears focused on `paused`/`terminated`. The broader set is deliberate for the execution-start guard. Resume semantics: #8317 keeps `agent_paused` as observability-only and classification-neutral (retryable), so a paused agent's in-flight work resumes on un-pause rather than escalating to blocked. ## Model Used - Provider: Anthropic. Model: Claude Opus 4 (`claude-opus-4-8`), via the Claude desktop "Cowork" agent. Mode: agentic/extended reasoning with tool use (shell, file editing, running `tsc`/`vitest`, git). Used to investigate the root cause in source, design the fix, implement it, and validate locally. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots (N/A — no UI change) - [ ] I have updated relevant documentation to reflect my changes (N/A — no doc-facing change) - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge - [x] I have searched the open and closed PR list for similar/duplicate PRs and found none## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Creating an agent starts at **New Agent → "manually" → pick an adapter**, which routes to `/{company}/agents/new?adapterType=claude_local` and renders the `NewAgent` page with the `AgentConfigForm` > - `AgentConfigForm` hands the parent a `triggerTestEnvironment` callback via an `onTestActionChange` effect so the page can wire up its "Test"/"Save + Test" button > - That trigger was rebuilt on every render: it depended on `runEnvironmentTest`, which is derived from a react-query `useMutation` result, and `useMutation` returns a **brand-new result object identity on every render** > - So the `onTestActionChange` effect re-fired every render and pushed a new function into the parent's state, producing an infinite `setState` loop ("Maximum update depth exceeded") that threw during render > - The app's custom router updates location **without remounting**, and there was no error boundary around the routed outlet, so the throw left a dead render tree — a fully **blank page** that stayed blank on back-navigation until a hard refresh > - This pull request stabilizes the trigger with a latest-ref pattern so the effect no longer re-fires, and adds a route-keyed error boundary so any future render throw degrades to a recoverable error card instead of a blank screen > - The benefit is that creating an agent works again, and render-time failures anywhere in the routed UI are contained and recoverable rather than silently blanking the app ## Linked Issues or Issue Description No public GitHub issue exists, so the underlying bug is described inline following the bug report template (`.github/ISSUE_TEMPLATE/bug_report.yml`): ### What happened? In the UI, choosing **New Agent → "manually" → (any adapter, e.g. Claude)** navigates to the agent-config page and renders a **completely blank page**. The browser console shows React's `Maximum update depth exceeded`. Using the back button changes the URL but the page stays blank until a full hard refresh. ### Expected behavior Selecting an adapter shows the agent configuration form so the agent can be created. ### Steps to reproduce 1. Open the app and click **New Agent**. 2. Choose **manually**. 3. Pick an adapter (e.g. Claude / `claude_local`). 4. Observe the blank page (URL becomes `/{company}/agents/new?adapterType=claude_local`). ### Paperclip version or commit Reproduces on `master` (base of this PR). ### Deployment mode Reproduces regardless of deployment mode — it is a client-side render loop. ### Root cause `useMutation` returns a new result object identity each render, so the `runEnvironmentTest`-derived `triggerTestEnvironment` callback was unstable, which made the `onTestActionChange` effect push a new function into parent state every render → infinite update loop → render throw → no boundary → blank tree. ## What Changed - **`ui/src/components/AgentConfigForm.tsx`** — Stabilize the environment-test trigger handed to the parent using a latest-ref pattern: the churny behavior (`runEnvironmentTest`, `testEnvironmentDisabled`) lives in a `useRef` updated by an effect, and the exposed `triggerTestEnvironment` is a `useCallback(() => triggerRef.current(), [])` with an empty dep array, so its identity is stable across renders and the `onTestActionChange` effect no longer re-fires every render. - **`ui/src/components/RouteErrorBoundary.tsx`** (new) — A route-keyed React error boundary that catches render throws and renders a recoverable error card (showing the error message, with "Go back" and "Reload page" actions). It resets automatically when the route (`pathname + search`) changes. - **`ui/src/components/Layout.tsx`** — Wrap the routed `<Outlet />` in `<RouteErrorBoundary>` so a render throw degrades to the error card instead of a blank page. - **`ui/src/components/RouteErrorBoundary.test.tsx`** (new) — Regression test: a throwing child is contained as a recoverable error card (showing the message), "Go back" calls `navigate(-1)`, and the boundary resets to render children again after the route changes. ## Verification - `npx vitest run ui/src/components/RouteErrorBoundary.test.tsx` → 3 passed (regression test for this fix). - `npx vitest run ui/src/components/AgentConfigForm.test.ts` → 9 passed. - `npx tsc -b` in `ui` → 0 errors. - Manual: ran the dev server, clicked **New Agent → manually → Claude**. - **Before:** blank page; console logs `Maximum update depth exceeded`; back button leaves the page blank until hard refresh. - **After:** the agent configuration form renders normally and the agent can be created; navigating away and back works without a hard refresh. - Boundary check: with the loop still in place (pre-fix), the new boundary catches the throw and shows a recoverable error card instead of a blank screen; "Go back" / route change resets it. _Screenshots: the before-state is the React `Maximum update depth exceeded` error and a blank `/agents/new` page; the after-state is the rendered agent-config form. Both were observed locally; rendered images can be attached on request._ ## Risks - **Low risk.** Changes are confined to three UI files with no API, schema, or behavioral change to agent creation beyond fixing the loop. - The latest-ref pattern preserves identical runtime behavior of the test trigger (same guard, same `runEnvironmentTest()` call) — it only stabilizes the callback identity. - The error boundary is additive; on the happy path it renders its children unchanged. Its only behavior is to catch render throws that previously blanked the app. ## Model Used Claude (Anthropic), model `claude-opus-4-8` — extended-thinking-capable, tool-use (file edit, shell, tests). Used to diagnose the infinite render loop, implement the latest-ref fix and route error boundary, and verify via local typecheck/tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots (described textually in Verification — see note) - [ ] I have updated relevant documentation to reflect my changes (N/A — bug fix, no docs affected) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Heartbeats are the control-plane path that turns scheduled, comment-driven, or on-demand wakeups into adapter executions. > - Budgeting and recurring work need enforcement before an adapter starts, not only after model usage is recorded. > - Empty timer wakes also need an opt-in fast-exit path so operators can keep routine schedules without paying for no-op model turns. > - This pull request adds heartbeat preflight gates for daily run and daily cost caps, plus an explicit timer no-work skip policy. > - The benefit is safer autonomous operation: capped agents stop before new execution, queued work is cancelled cleanly at claim time, and proactive agents still run by default unless the operator opts into no-work skipping. ## Linked Issues or Issue Description No public issue exists for this change. Inline bug report: ### What happened? Heartbeat execution can start without enforcing per-agent daily invocation and spend limits at the heartbeat boundary. A run that was queued before a cap was reached can also be claimed later and invoke the adapter unless the cap is checked again immediately before execution. Operators also do not have an explicit opt-in fast-exit policy for generic timer wakes with no actionable assigned work. ### Expected behavior Configured daily run and daily cost caps should stop new heartbeat runs before adapter execution. Already queued runs should be rechecked at claim time and cancelled cleanly when a cap is now reached. Queued issue runs cancelled by daily caps should release their issue execution locks and promote deferred wakeups without entering immediate recovery loops. Generic timer no-work skipping should be opt-in so proactive agents continue to run by default. ### Steps to reproduce 1. Configure an agent heartbeat policy with a one-run daily cap or a daily cost cap. 2. Create or queue heartbeat wakeups for that agent after the cap has already been consumed. 3. Observe that without preflight and claim-time checks, the heartbeat path can still enqueue or claim work that should be blocked before adapter execution. ### Paperclip version or commit Reproduced against `master` before this branch. ### Deployment mode Local dev (`pnpm dev`) / built from source. ### Installation method Built from source (`pnpm dev` / `pnpm build`). ### Agent adapter(s) involved Not adapter-specific (core heartbeat scheduling and claim logic). ### Database mode External Postgres in tests via embedded test harness. ### Access context Not applicable. ### Node.js version Node 20 in CI-compatible local development. ### Operating system macOS local development, Linux CI-compatible tests. ### Relevant logs or output The regression suite added in this PR covers the failing paths: ```shell pnpm exec vitest run server/src/__tests__/heartbeat-stale-queue-invalidation.test.ts ``` ### Relevant config (if applicable) ```json { "heartbeat": { "maxDailyRuns": 1, "maxDailyCostCents": 1, "skipTimerWhenNoActionableWork": true } } ``` ### Additional context This affects recurring/autonomous operation because the safest place to stop excess work is before adapter execution starts. ### Privacy checklist Reviewed for sensitive data; no private logs, credentials, or local instance URLs are included. ## What Changed - Added heartbeat policy parsing for per-agent daily run caps, daily cost caps, and opt-in no-actionable-work timer skipping. - Added pre-queue daily cap checks while preserving same-issue wake coalescing. - Added claim-time cap checks so already queued runs are cancelled before adapter execution when a cap is reached. - Added skipped wakeup metadata for cap and timer fast-exit decisions. - Released issue execution locks for queued issue runs cancelled by daily caps, with deferred wake promotion and without immediate recovery loops while caps are active. - Added regression coverage for timer skipping, proactive default behavior, run caps, cost caps, queued-run cancellation, started cancelled runs, and deferred issue wake promotion. ## Verification - `git diff --check` - `node -c server/src/services/heartbeat.ts` - `pnpm exec vitest run server/src/__tests__/heartbeat-stale-queue-invalidation.test.ts` - `pnpm --filter @paperclipai/server typecheck` - Local autoreview: `skills/autoreview/scripts/autoreview --mode branch --base origin/master --engine codex --model gpt-5.5 --thinking high` - Result: clean, no accepted/actionable findings ## Risks - Medium operational risk because this changes heartbeat scheduling and claim-time behavior. - The no-actionable-work timer fast-exit is explicitly opt-in to avoid suppressing proactive agents unexpectedly. - Daily run caps count runs by `startedAt` so old queued rows do not consume today’s cap, while started runs still count even if they later end as cancelled. - Queued issue-run cap cancellation uses the existing release/promotion path with immediate recovery suppressed to avoid retry loops while caps are active. ## Model Used Codex with GPT-5.5 high reasoning assisted with implementation, local testing, and autoreview. The final review gate used local autoreview with `gpt-5.5` high reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergeplugin target(#8575) 1951c80237## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runs stream their transcripts through per-adapter stdout parsers into the chat/run transcript UI (`ui/src/adapters/transcript.ts`) > - The Cursor CLI (local) streams assistant text as many small `text` events (often a token or word each), and the parser emitted one assistant entry per event and trimmed each > - As a result the chat rendered one bubble per token ("every line a new token") and dropped inter-token whitespace, making Cursor runs hard to read > - The render layer already coalesces consecutive `delta` entries (`appendTranscriptEntry`), but the Cursor parser never tagged streamed text as a delta > - This pull request tags streamed `text` as a delta (without trimming) so the existing render-time coalescer merges them into one assistant block, while a `tool_call`/`tool_result` between deltas still breaks the run > - The benefit is readable Cursor transcripts with correct spacing and preserved tool boundaries, with no change to the canonical event stream (raw view unaffected) ## Linked Issues or Issue Description No existing public issue — describing the bug inline (per `.github/ISSUE_TEMPLATE/bug_report.yml`): **What happened** In the chat/run transcript, Cursor (local) assistant messages render as one bubble per token/word, and inter-token spaces are dropped, making the transcript unreadable. Root cause: `packages/adapters/cursor-local/src/ui/parse-stdout.ts` (`type: "text"` branch) emitted `{ kind: "assistant" }` per streamed `text` event without `delta: true` and trimmed each, so the render-time coalescer (`ui/src/adapters/transcript.ts`) never merged them and whitespace was lost. **Expected behavior** Streamed assistant text should render as a single contiguous prose block, with tool calls preserved as boundaries between blocks. **Steps to reproduce** 1. Run a Cursor (local) agent that streams a multi-word assistant message. 2. Open the run transcript in the chat UI. 3. Observe each streamed token/word rendered as its own bubble, with inter-token spaces missing. **Paperclip version** Reproduced on current `master` (cutover base `e68188c43`). **Deployment mode** Self-hosted, `cursor_local` adapter. ## What Changed - `packages/adapters/cursor-local/src/ui/parse-stdout.ts`: tag streamed `text` events as `{ kind: "assistant", delta: true }` and stop trimming, so the existing `appendTranscriptEntry` coalescer merges consecutive deltas into one block. - `ui/src/adapters/cursor-coalescing.test.ts` (new): dual-shape golden fixtures (Cursor local + cloud) exercising the full render-time projection via `buildTranscript`. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/adapters/cursor-coalescing.test.ts src/adapters/transcript.test.ts` → **10/10 pass**. - `pnpm --filter @paperclipai/ui --filter @paperclipai/adapter-cursor-local typecheck` → **green**. - The golden fixtures assert the run `text → tool_call → tool_result → text → consolidated final` renders as exactly **two prose blocks with the tool between them**, **no duplication** of the consolidated final, and **inter-token whitespace preserved** across coalesced deltas. ## Risks - **Low risk.** Pure classification at parse time; the canonical event stream and the raw view are unchanged — only the "nice" render-time projection changes. The coalescing logic (`appendTranscriptEntry`) is pre-existing and already covered by tests. No schema, migration, or behavioral change outside transcript rendering. ## Model Used - **Claude Opus 4.8** (Anthropic), extended/high reasoning mode, driven via the Cursor agent with tool use + code execution. Diagnosis and fixtures grounded in the repo's actual parser/render code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (none found) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change (`fix/cursor-transcript-coalescing`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes (N/A — no documented behavior changes) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending CI run) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending review) - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Sebastian Heyneman <sebastian@joinnova.com> Co-authored-by: Cursor <cursoragent@cursor.com>## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The dev/server UI exposes a `/api/health` endpoint and a lower-left account drawer, but nothing surfaces *which* build the running instance is on or when it last restarted > - When iterating on a local dev instance it is hard to tell whether the server you're looking at has actually restarted onto your latest commit, or how stale the running process is > - Developers need a lightweight, opt-in way to confirm the running instance's identity without digging through logs or shelling into the host > - This pull request adds an experimental "Server Info Debug View" setting that surfaces the running instance's last-restart time and current commit as read-only rows in the account drawer > - The benefit is a quick, in-UI sanity check of what the live server is actually running, behind an experimental flag so it ships zero cost to users who don't opt in ## Linked Issues or Issue Description No public GitHub issue exists. Describing the underlying request inline following the feature request template: **Problem or motivation:** When working against a local Paperclip dev instance there is no in-UI way to confirm what the running server is — its current commit or when it last restarted. You have to check logs or the host shell to know whether the process picked up your latest build. **Proposed solution:** An opt-in experimental setting ("Server Info Debug View") that, once enabled, renders a small read-only "Server" section at the bottom of the lower-left account drawer showing **Last restarted** (the server process start time) and **Running commit** (the current git HEAD short SHA + subject). **Alternatives considered:** A separate top-right pill/overlay (like the work-life-balance plugin). The account drawer was chosen to reuse existing menu-row styling and avoid adding new always-present chrome. **Roadmap alignment:** Small, self-contained developer-experience aid gated behind an experimental flag; does not overlap planned core roadmap work. ## What Changed - Added `server/src/server-info.ts`: captures a `serverInfo` snapshot once at boot — process start time and current git commit (SHA + subject). Git is read via `execFileSync` with SHA validation and a timeout. - `/api/health` exposes the `serverInfo` snapshot, but only on full-details health responses (board/agent in authenticated mode, or local-trusted dev). - Gated the UI surface behind a new `enableServerInfoDebugView` experimental setting, wired through the shared instance type, validator, settings normalizer, and OpenAPI schema. - UI: added `SidebarServerInfo` rendering the read-only rows in the account drawer (`BreadcrumbBar` / `SidebarAccountMenu`), plus the experimental settings toggle and a typed `health` API client. - Moved `ServerGitInfo` / `ServerInfoSnapshot` into `@paperclipai/shared` so the server and UI share one definition instead of duplicating it. - Added unit tests for the server-info snapshot, health route exposure, validator/normalizer, settings routes, the experimental settings page, and the sidebar component. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/health.test.ts src/__tests__/server-info.test.ts src/__tests__/instance-settings-service.test.ts src/__tests__/instance-settings-routes.test.ts` — 32 passed - `pnpm --filter @paperclipai/ui exec vitest run src/components/SidebarServerInfo.test.tsx src/pages/InstanceExperimentalSettings.test.tsx` — 9 passed - `tsc --noEmit` on both `@paperclipai/server` and `@paperclipai/ui` — clean - Manual: enable **Settings → Experimental → Server Info Debug View**, refresh the UI, open the lower-left account drawer — a "Server" section shows Last restarted and Running commit. ## Risks - Low risk. The UI surface is fully opt-in via an experimental flag and defaults off. - The `serverInfo` field on `/api/health` is access-controlled to full-details responses only (board/agent in authenticated mode, or local-trusted dev) — never anonymous authenticated callers — so the git SHA is not broadly exposed. - The only new server work is a one-time git read at boot, guarded with SHA validation and a timeout; failures degrade gracefully (the git block reports `available: false` rather than throwing). ## Model Used Claude — `claude-opus-4` (Anthropic), extended thinking with tool use, via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge## Thinking Path > - Paperclip orchestrates ai-agents for zero-human companies > - But humans want to watch the agents and oversee their work > - Human users interact with the dashboard UI to monitor costs, schedules, and agent behavior > - The UI has toggle button groups where you click one option and other become unselected (like date range presets, day-of-week selectors) > - Visually the active button look different (filled vs outlined), but screen reader users hear no difference - all buttons sound same > - The `aria-pressed` attribute is what screen readers need to announce which toggle is active and which is not > - I searched entire UI codebase for `aria-pressed` and found zero usage anywhere > - This pull request adds `aria-pressed={isActive}` to two toggle button groups: date range presets in Costs page and day-of-week selector in ScheduleEditor > - Now screen reader users can tell which button is currently selected without relying on visual styling only ## Problem The app has button groups that work like toggles - you click one button to select it and the others become unselected. Visually this work fine because the active button change to a different variant (filled vs outlined). But for screen reader users, ALL buttons sound exactly the same - just "button, 7 days", "button, 30 days", etc with no way to tell which one is currently active. I searched the entire UI codebase for `aria-pressed` and found zero results. This attribute is what screen readers need to announce "7 days, pressed" vs "30 days, not pressed" for toggle button groups. ## What I changed Added `aria-pressed={isActive}` to two toggle button groups: 1. **Costs.tsx** - Date range preset buttons (7d, 30d, 90d, Custom). Screen reader now announce which date range is selected. 2. **ScheduleEditor.tsx** - Day of week selector (Mon, Tue, Wed...). Screen reader now announce which day is selected for weekly schedule. ## How to test 1. Go to Costs page, use VoiceOver (Cmd+F5 on Mac) 2. Tab through the date preset buttons 3. Active button should announce "pressed", others "not pressed" 4. Same for ScheduleEditor - create/edit trigger with weekly preset, tab through day buttons 2 files, 2 lines added.## Thinking Path Two interactive UI controls were missing WAI-ARIA attributes, so screen-reader users couldn't perceive their state. The IssuesList view-mode toggle already had `title` tooltips but no `aria-label`/`aria-pressed` and its container had no `role="group"`; the GoalTree expand/collapse chevron announced only "button". Approach: attributes only, no logic/render changes, reusing the pattern already shipped on the Agents page toggle. Per review feedback, GoalTree uses a **stable** `aria-label` (`` `${goal.title} subtree` ``) with `aria-expanded` for state, rather than a dynamic label that double-announces state. ## Issue _No existing tracking issue — described inline per CONTRIBUTING.md → "Link Issues or Describe Them In-PR"._ **What happened** Two components expose buttons with no accessible name or state, making them unusable via screen reader: (1) the `IssuesList` view-mode toggle doesn't convey which view is active; (2) the `GoalTree` expand/collapse chevrons have no name and no expanded/collapsed state. **Expected behavior** Both controls announce their purpose and current state to assistive technology. **Steps to reproduce** Enable VoiceOver, open the Issues page and Tab to the view-mode toggle, then open the Goals page with nested goals and Tab to a tree chevron — each control announces only "button", with no name and no pressed/expanded state. ## What Changed **`ui/src/components/IssuesList.tsx`** — `role="group"` + `aria-label="View mode"` on the container; `aria-label` ("List view"/"Board view") and `aria-pressed` on each button. **`ui/src/components/GoalTree.tsx`** — stable `aria-label` (`` `${goal.title} subtree` ``) and `aria-expanded` on the chevron button. 2 files, ARIA attributes only, no behavioral change. ## Verification 1. Issues page → toggle announces "List view, pressed" / "Board view, not pressed", grouped as "View mode". 2. Goals page with nested goals → each chevron announces "<goal title> subtree" with expanded/collapsed state. 3. Manual VoiceOver pass; no visual/behavioral change for sighted users. ## Risks Minimal — additive HTML attributes with no impact on logic, rendering, or state. Worst case is a suboptimal announcement string, trivially adjusted. ## Model Used Original change human-authored by @bluzername. Two follow-up commits (stable `aria-label` refinement; removal of a stray tooling file) applied via maintainer edit; the refinement was drafted with Claude Opus 4.8. ## Checklist - [x] I searched the GitHub PR list (open + recently closed) for similar/duplicate PRs before opening — none found. --------- Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip stores heartbeat run ownership context so operators can audit which user was responsible for agent work > - A migration backfills missing heartbeat run `responsible_user_id` values from issue references in each run's context snapshot > - Some context snapshots can store issue identifiers as public ticket strings rather than UUIDs > - The migration needs to resolve both UUID issue ids and issue identifiers without trying to cast identifier strings to UUID > - This pull request tightens the migration query so UUID matching only casts validated UUID-shaped values and identifier matching remains a separate fallback > - A companion repair migration is needed so installations that already recorded `0130` still get the corrected heartbeat-run backfill > - The benefit is that existing installations can apply the responsible-user invariant migrations without failing on non-UUID issue references and without leaving already-migrated databases unrepaired ## Linked Issues or Issue Description No public GitHub issue was found in a quick search for this migration failure. Bug report context following `.github/ISSUE_TEMPLATE/bug_report.yml`: **Pre-submission checklist** - I searched existing open and closed issues and did not find a duplicate. - I can reproduce against current `master` plus the responsible-user invariant migration path. - The error originates in Paperclip's database migration, not in an adapter, API provider, or local configuration. **What happened?** Applying the heartbeat run responsible-user backfill migration could fail when `heartbeat_runs.context_snapshot->>'issueId'` or `taskId` contained an issue identifier such as `PAP-123` instead of a UUID. The migration attempted to use issue refs for UUID matching and identifier fallback, but the UUID path needed to avoid casting non-UUID identifier strings. Because `0130` may already have been applied in some installations, a follow-up repair migration is needed as well. **Expected behavior** The migration should backfill from UUID issue ids when present, from issue identifiers when present, and fall back to the company default responsible user without unsafe UUID casts. Already-migrated installations should receive the repaired heartbeat-run context-ref backfill through a new migration. **Steps to reproduce** 1. Use a migrated database with a company, issue, agent, and heartbeat run. 2. Store a null `heartbeat_runs.responsible_user_id` and a `context_snapshot` like `{"issueId":"PAP-123"}`. 3. Replay/apply the run responsible-user repair migration. 4. Observe that the migration must not cast `PAP-123` to UUID and should backfill from the matching issue identifier. **Paperclip version or commit** Current `master` plus this migration fix branch. **Deployment mode** Database migration during server startup or explicit migration command. **Installation method** Built from source / self-hosted migration path. **Agent adapter(s) involved** Not adapter-specific; core database migration bug. **Database mode** Postgres migration path, including embedded Postgres in development. **Access context** Not applicable; migration-time data backfill. **Relevant logs or output** Unsafe UUID casts can surface as Postgres invalid input syntax errors when a context snapshot issue ref is an identifier rather than a UUID. **Privacy checklist** No private logs, paths, API keys, tokens, company names, or internal Paperclip issue links are included. ## What Changed - Split heartbeat run context issue reference extraction into reusable CTEs. - Only cast `issueId` / `taskId` values to UUID after a UUID-shape regex check. - Preserve fallback matching by issue identifier within the same company. - Keep deterministic candidate priority with `issueId` before `taskId` and UUID matches before identifier matches. - Added `0131_repair_run_responsible_user_context_refs.sql` so installations that already applied `0130` still receive the corrected heartbeat-run backfill. - Added a DB migration regression test that replays the repair migration with an identifier-style heartbeat run issue ref. ## Verification - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/db exec vitest run src/client.test.ts` - Isolated embedded Postgres migration run with a temporary `PAPERCLIP_CONFIG`: `pnpm --filter @paperclipai/db migrate` applied all pending migrations successfully. - GitHub Actions PR workflow is green on commit `1c2655bc457cef3716a43e66463e1e5bc2fdcfab`. - Greptile is 5/5 with no unresolved threads on commit `1c2655bc457cef3716a43e66463e1e5bc2fdcfab`. - Searched for duplicate public issues/PRs with GitHub search; no direct duplicate found. - Checked `ROADMAP.md` for overlap; no related roadmap item found. ## Risks - Low-to-medium migration risk because this modifies an existing data backfill migration and adds a companion repair migration. - The query still relies on `context_snapshot` containing either issue UUIDs or identifiers that match issues in the same company. - Installations with unusual malformed context snapshots now skip unsafe UUID casts and fall through to identifier/default backfill behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex coding agent with tool use and local command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes, or confirmed no docs update is needed for this migration-only fix - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>## Thinking Path > - Paperclip is the open-source app people use to manage AI agents for work > - The task/issue lifecycle subsystem tracks when agents wake up, are suppressed, or are deferred, recording each wake request in `agent_wakeup_requests` and each defer/suppression event in `activity_log` > - When an agent appears stuck or doesn't resume after a dependency resolves, there is currently no read-only API surface to inspect its wake history — operators must query the database directly > - Making wake history queryable via a first-class endpoint lets operators, support, and monitoring tools diagnose "why didn't this agent wake up?" without database access > - This pull request adds `GET /api/issues/:id/diagnostics/wakes`, returning a bounded 14-day/50-row projection of wake requests and defer/suppression activity events, with a deterministic `diagnosis` field and a `likelyReason` inference — including a Case-B inference ("no wake enqueued because a visible blocker is not done") that reuses the blocker readiness data from the companion blocker diagnostics endpoint (see Refs #9114) > - The benefit is that platform operators can answer "why is this agent not waking up?" from a safe, read-only HTTP endpoint rather than needing direct database access, and CI/monitoring can assert expected wake behavior ## Linked Issues or Issue Description Refs #9114 (companion blocker diagnostics endpoint, already merged — this PR extends the same diagnostic surface to wake/activity history) ## What Changed - **New route** `GET /api/issues/:id/diagnostics/wakes` in `server/src/routes/issues.ts`: returns a bounded (14-day window, 50-row cap) projection of `agent_wakeup_requests` rows and wake-relevant `activity_log` rows (defer/suppression events) - **Sanitized projection**: raw `payload`, `details`, `error`, and `triggerDetail` fields are stripped; unknown free-form `source`, `reason`, and `status` values are projected to `"other"` to prevent schema bleed - **Deterministic `diagnosis` and `likelyReason` fields**: includes Case-B inference ("no wake enqueued — visible blocker not done") that calls the existing blocker-readiness helper from Slice 1 (#9114) so the wake surface can explain missing wakes caused by outstanding blockers - **Auth**: `assertCompanyAccess` + `assertIssueReadAllowed`; cross-company requests are denied; Case-B blocker inference filters by caller trust level so hidden (low-trust) blockers are mentioned but not identified - **Types in `@paperclipai/shared`**: `IssueWakeDiagnosticsResponse`, `WakeEvent`, `ActivityEvent` exported from the shared package - **OpenAPI tag registration** for the new route - **Skill reference docs** in `skills/paperclip/references/api-reference.md` documenting the endpoint contract - **Test coverage** (`server/src/__tests__/issue-wake-diagnostics-routes.test.ts`, embedded Postgres): happy path, empty/null diagnosis, Case-B inference, low-trust hidden blocker, cross-company denial, raw blob minimization, cap behaviour, combined blocker+wake test run ## Verification ```bash # Wake diagnostics tests only pnpm exec vitest run server/src/__tests__/issue-wake-diagnostics-routes.test.ts # Wake + blocker diagnostics together (integration) pnpm exec vitest run server/src/__tests__/issue-blocker-diagnostics-routes.test.ts server/src/__tests__/issue-wake-diagnostics-routes.test.ts # Type-check shared and server packages pnpm --filter @paperclipai/shared typecheck pnpm --filter @paperclipai/server typecheck # Whitespace / diff check git diff --check ``` All commands passed locally. ## Risks - **No schema or migration changes** — this is a read-only projection over existing tables; no DDL risk. - **Bounded queries** — 14-day window + 50-row cap limit per call; no unbounded scans. - **Auth boundary** — cross-company access is denied at `assertCompanyAccess`; Case-B inference uses the same per-node trust filtering as the blocker endpoint so low-trust blockers are acknowledged but not identified. - Low overall risk; the endpoint is additive and read-only. ## Model Used - **Provider:** Anthropic - **Model:** Claude Sonnet 4.6 (`claude-sonnet-4-6`) - **Tool use:** yes (file reads, edits, bash execution, Paperclip API calls) - **Reasoning mode:** standard (no extended thinking) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>## Thinking Path > - Paperclip is an open-source platform for orchestrating AI agents; agents run inside execution workspaces that range from a shared container to full git worktrees cloned from a project repository. > - Isolated git-worktree workspaces require a project to determine which repository to clone — without a project the worktree base path cannot be computed. > - A task pinned to `isolated_workspace` + `git_worktree` with no project was previously accepted at creation time but failed late at dispatch with the opaque `workspace_validation_failed` / `git_worktree_base_agent_home` error — only after the heartbeat attempted to provision the workspace. > - Fail-closed validation should happen in two places: (1) explicit create/update pins that contradict the requirement are rejected at the HTTP layer with a structured 422; (2) rows that reach the heartbeat dispatcher with this invalid combination (e.g. through inheritance or a retroactively-removed project) are blocked before any heartbeat run or adapter spawn. > - This PR adds the shared detection policy, the create/update guard in the issues service, and the heartbeat pre-dispatch guard, together with focused unit tests for all three layers. > - The benefit is deterministic early failure with a clear remediation message instead of a late, cryptic runtime error. ## Linked Issues or Issue Description No upstream public GitHub issue — describing the problem inline (bug-report format). **What happened?** Creating an issue with `executionWorkspaceSettings: { mode: "isolated_workspace", type: "git_worktree" }` and no `projectId` was accepted without error. The issue then became blocked at dispatch time with the opaque message `git_worktree_base_agent_home` / `workspace_validation_failed` — surfaced only after the heartbeat attempted to provision the workspace. **Expected behavior** The platform should reject the invalid combination at create/update time with a structured 422 that includes a clear remediation message, before any heartbeat resource is consumed. **Steps to reproduce** 1. Call `POST /api/issues` (or `PATCH /api/issues/:id`) with `executionWorkspaceSettings: { mode: "isolated_workspace", type: "git_worktree" }` and omit `projectId` (or set it to `null`). 2. Observe: request succeeds (200/201). 3. Assign the issue to an agent and watch it enter `blocked` with a cryptic `workspace_validation_failed` error at dispatch. **Related prior fix** — Refs #4844 (`fix(validator): reject static cwd combined with git_worktree strategy`) — same validation area, different dimension (static cwd vs. missing project). **Paperclip version / commit** Latest `master` (pre-this-PR). **Deployment mode** Standard (app-global server). ## What Changed - **`execution-workspace-policy.ts`** — new shared `detectWorkspaceWorktreeRequiresProject` function returning a stable `workspace_worktree_requires_project` policy violation when an isolated git-worktree task has no project, project workspace, or reusable execution workspace; exports canonical remediation text used by both the HTTP guard and the heartbeat guard. - **`issues.ts`** — create and update paths check the new policy before persisting; explicit pins to `isolated_workspace` / `operator_branch` + `git_worktree` with no project are rejected with a 422 including the policy code and remediation text. - **`heartbeat.ts`** — pre-dispatch preflight checks the same policy for rows that reach the heartbeat with the invalid combination (e.g. through inheritance); such rows are marked `blocked` with a skipped wakeup request, durable issue comment, and activity log before any heartbeat run or adapter spawn. - **`execution-workspace-policy.test.ts`** — focused policy-layer unit tests for detection logic and remediation text. - **`issues-service.test.ts`** — create/update 422 guard tests for the new policy. - **`heartbeat-workspace-branch-containment.test.ts`** — pre-dispatch blocking test for inherited/ambiguous invalid rows; also fixes a cleanup race in the existing test suite. ## Verification ```sh pnpm exec vitest run \ server/src/__tests__/execution-workspace-policy.test.ts \ server/src/__tests__/issues-service.test.ts \ server/src/__tests__/heartbeat-workspace-branch-containment.test.ts pnpm --filter @paperclipai/server typecheck # Targeted regression pnpm exec vitest run \ server/src/__tests__/heartbeat-workspace-branch-containment.test.ts \ -t "blocks projectless isolated git-worktree issues before dispatch" ``` All three test files and typecheck passed locally before this PR was opened. ## Risks **Low risk.** The policy detection function is pure with no side effects. The create/update guard only triggers on explicit `isolated_workspace` or `operator_branch` + `git_worktree` pins combined with a missing project — it does not fire on inherited settings (handled by the heartbeat preflight), so there is no false-positive rejection risk for valid tasks. The heartbeat guard fires before any resource is provisioned; the only behavioral change for already-invalid rows is that they receive a clear `blocked` status and durable comment instead of a late cryptic error. ## Model Used Claude Sonnet 4.6 (`claude-sonnet-4-6`), 200k context window, tool use enabled (agentic coding). Used to implement all server-side changes and tests in this PR. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents tackle complex tasks via *plan* flows: a planner decomposes work into child issues, which are accepted by the board and then executed > - When a plan is accepted, `createChild` in `issues.ts` inserts child issues pre-bound to the parent's already-realized execution workspace — carrying over its concrete branch ref > - If the repository's base ref advances between plan acceptance and a child's first heartbeat, the child inherits a stale branch that no longer matches the current base > - At first heartbeat the workspace validator detects the mismatch and freezes the child ("branch freeze"), blocking it from starting any work > - The real fix is to strip the concrete workspace binding when creating accepted-plan children: they should receive only the unresolved *intent* (mode, baseRef, branchTemplate) and realize a fresh workspace from the current base on their own first heartbeat > - This PR implements that strip, adds a regression test that proves a post-base-advance child realizes cleanly, and also fixes `parseIssueExecutionWorkspaceSettings` so `environmentId` is not silently dropped on update round-trips (a latent bug that was masking the original fix) ## Linked Issues or Issue Description No public GitHub issue exists for this bug. Bug description follows the bug-report template: **What happened:** Accepted-plan decomposition pre-binds child issues to the parent's realized execution workspace branch (`executionWorkspaceId` + `executionWorkspaceBranch`). When `origin/master` advances between plan acceptance and the child's first heartbeat, the workspace branch interlock fires and the child is permanently frozen before it can start. **Expected behavior:** Accepted-plan children should receive only unresolved workspace intent (mode, git strategy fields) and realize a fresh isolated worktree from the current base on first heartbeat. A base-ref advance between acceptance and first-run should be transparent. **Steps to reproduce:** 1. Accept a plan that decomposes into one or more child issues (isolated_workspace + git_worktree mode). 2. Allow `origin/master` to advance (new merge). 3. Observe the first child heartbeat: workspace validation fails with a branch-freeze error. **Paperclip version:** current `master` (pre-fix). **Deployment mode:** any (affects all modes that use isolated workspace + git worktree strategy). Supersedes #9227 (earlier attempt, now closed — the fix was incomplete because `environmentId` was silently dropped during `parseIssueExecutionWorkspaceSettings` update round-trips, causing the child workspace to lose its environment binding; this PR includes that fix). ## What Changed - **`server/src/issues.ts` — `createChild` / accepted-plan decomposition path:** strip resolved workspace fields (`executionWorkspaceId`, concrete branch) when creating accepted-plan children; preserve only unresolved intent fields (`mode`, `baseRef`, `branchTemplate`, `environmentId`, runtime/provisioning settings). - **`server/src/execution-workspace-policy.ts` — `parseIssueExecutionWorkspaceSettings`:** preserve `environmentId` through update round-trips (was silently dropped, causing environment to detach on any workspace settings update). - **`server/src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts`:** new regression test — accepted-plan child created after `origin/master` moves realizes a fresh isolated worktree from the moved base and passes workspace execution. - **`server/src/__tests__/issues-service.test.ts`:** extended workspace-linkage and `createChild` tests covering the accepted-plan strip and the unchanged direct-child path. ## Verification ```sh pnpm exec vitest run server/src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts pnpm exec vitest run server/src/__tests__/issues-service.test.ts -t "workspace linkage|accepted plan decomposition|createChild applies" pnpm --filter @paperclipai/server typecheck ``` All three pass on this branch. ## Risks **Low.** The change is scoped to the accepted-plan `createChild` code path. The direct child / follow-up issue creation path (normal non-plan decomposition) is unchanged and covered by existing tests. The `parseIssueExecutionWorkspaceSettings` fix is additive — it now preserves a field that was previously silently dropped, so no consumer loses data. ## Model Used Claude Sonnet 4.6 (`claude-sonnet-4-6`), Anthropic, 200K context window, extended tool use + code generation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (supersedes #9227) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>## Summary When a task runs in a branch-pinned execution workspace, agents sometimes switch or rename the workspace branch, which breaks the worktree contract. This adds a short, one-time prompt hint telling the agent to stay on the pinned branch. - **heartbeat.ts**: after the execution workspace is resolved, attach `executionWorkspace: { branchName }` to the wake payload (only when a branch pin exists — agent-home runs without a branch are untouched). - **server-utils.ts (adapter-utils)**: normalize the new payload field and render one bullet in `renderPaperclipWakePrompt`: > `- execution workspace branch: you are running in an execution workspace on branch \`<name>\`. Do not switch, rename, or re-point this branch; keep all commits on it.` - The hint renders **only on non-resumed sessions** — resume-delta prompts skip it, so it appears the first time an issue's session starts, not on every turn, and it never pollutes the issue thread. One renderer change covers every adapter (claude, codex, cursor, gemini, grok, opencode, pi, hermes, acpx engine) with zero per-adapter edits. ## Tests - `server-utils.test.ts`: branch guard renders on first prompt, absent on resumed-session prompts, absent when no branch is pinned; payload round-trips through `stringifyPaperclipWakePayload`. - `heartbeat-workspace-branch-containment.test.ts`: the finalize-path adapter mock now asserts the wake payload the adapter receives carries the branch pin matching `context.paperclipWorkspace.branchName` (end-to-end heartbeat wiring, embedded postgres). All 6 pass. - Full `server-utils` (56) and acpx-engine execute (34) suites green; adapter-utils typechecks clean; no new server tsc errors. PAP-13326 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing>## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Those agents run as child processes spawned by `@paperclipai/adapter-utils`'s `runChildProcess`, which arms a parent-side wall-clock timer at `timeoutSec` to bound a hung run > - At the deadline `runChildProcess` sends SIGTERM, then after a grace window escalates to SIGKILL — the SIGKILL backstop is what turns a wedged child into a dead PID so the scheduler can reclaim and retry it > - On the **direct-child fallback** path (`signalRunningProcess`, used on win32 and whenever process-group signaling is unavailable or throws) the escalation was gated on `!child.killed` > - But Node sets `ChildProcess.killed` to `true` the instant a signal is *successfully sent*, not when the process exits — so once the earlier SIGTERM has been sent, `child.killed` is already `true`, the `!child.killed` guard is `false`, and the SIGKILL escalation never runs > - A child that ignores SIGTERM (e.g. a graceful-shutdown handler wedged on a socket) is therefore never force-killed, outlives its deadline, and for an unattended scheduler sits running forever with no terminal state > - This PR gates the fallback escalation on real liveness (`exitCode === null && signalCode === null`), so SIGKILL fires precisely while the child is still alive > - The benefit is the hard timeout actually guarantees termination (except true uninterruptible D-state) on every platform/configuration, not just where the process-group path is available ## Linked Issues or Issue Description No existing public issue — describing the bug inline, following `.github/ISSUE_TEMPLATE/bug_report.yml`: ### What happened? When `@paperclipai/adapter-utils`'s `runChildProcess` reaches `timeoutSec` and the spawned child ignores SIGTERM, the SIGKILL escalation on the **direct-child fallback** path (`signalRunningProcess`, taken on win32 or whenever `process.kill(-pgid, …)` is unavailable or throws) never fires, so the child outlives its deadline indefinitely. Root cause: the escalation is gated on `!running.child.killed`, and `ChildProcess.killed` reflects only that a signal was *successfully sent* (per the Node docs it "does not indicate that the child process has been terminated"). After the deadline SIGTERM, `child.killed` is already `true`, so `!child.killed` is `false` and the follow-up SIGKILL is suppressed. ### Expected behavior After the grace window, a child that is still alive is force-killed with SIGKILL regardless of whether SIGTERM was already sent — the hard timeout should guarantee termination (except true uninterruptible D-state) on every platform/configuration. ### Steps to reproduce 1. Spawn a child that installs a no-op `SIGTERM` handler and never exits (e.g. `process.on('SIGTERM', () => {}); setInterval(() => {}, 1000)`). 2. Drive it through the direct-child fallback, i.e. `signalRunningProcess({ child, processGroupId: null }, …)` (the path used on win32 / when group signaling is unavailable). 3. Send SIGTERM (the child swallows it; `child.killed` becomes `true`), then send SIGKILL. 4. On the pre-fix `!child.killed` guard the SIGKILL call is a no-op and the PID survives past its deadline. Covered by the new regression test in this PR. ### Paperclip version or commit Reproduces on `master` (the `signalRunningProcess` fallback). Also present in published `@paperclipai/adapter-utils` (e.g. `2026.325.0`), where the same `!child.killed` guard sits on the single direct-child escalation path. _Searched the open PR list for duplicates/related work on `runChildProcess` / `signalRunningProcess` / SIGKILL escalation; found none._ ## What Changed - `packages/adapter-utils/src/server-utils.ts`: in `signalRunningProcess`, replace the direct-child fallback guard `!running.child.killed` with `running.child.exitCode === null && running.child.signalCode === null` (real liveness). The process-group path is unchanged. - `packages/adapter-utils/src/server-utils.ts`: `export` `signalRunningProcess` so the fallback branch can be unit-tested directly. - `packages/adapter-utils/src/server-utils.test.ts`: add a companion regression test (POSIX-only, like the sibling timeout tests) that forces the fallback (`processGroupId: null`) — sends SIGTERM (child swallows it, `child.killed` becomes `true`), asserts the child is still alive, then sends SIGKILL and asserts the PID dies. Also keeps the end-to-end `runChildProcess` SIGTERM-ignoring test. ## Verification ``` npx vitest run packages/adapter-utils/src/server-utils.test.ts # 52 passed npx tsc --noEmit # clean ``` - **Regression proof:** reverting the guard to `!running.child.killed` makes the new fallback test fail (`waitForPidExit` → false; the child survives); the liveness guard makes it pass. This addresses the prior review note that the existing test only exercised the process-group path (which already escalated correctly on POSIX) and never reached the changed branch. ## Risks Low. A one-line guard change scoped to the direct-child fallback; the process-group path is untouched. SIGKILL is only sent when `exitCode`/`signalCode` are both still `null`, i.e. the process is provably alive, so the change cannot signal an already-reaped/recycled PID. New tests are POSIX-only and `skipIf(win32)`, consistent with the sibling timeout tests in this file. ## Model Used Anthropic **Claude Opus 4.8**, driven via the Cursor agent (extended reasoning + tool use, large context). Diff, tests, and the regression proof above were produced and run by the agent; reviewed by a human before pushing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent configuration includes adapter-specific model settings and a built-in adapter test action so operators can verify runtime configuration before saving changes. > - Cody/Codex-style local adapters can use an adapter default model when the user clears the explicit model field. > - The adapter test path still passed an object containing `model: undefined` in some create/edit flows, which is different from omitting the model and can break default-model behavior. > - The previous fix was reverted because it also included an unrelated skill documentation edit. > - This pull request reapplies only the UI default-model test-config fix, with no doc or skill changes. > - The benefit is that testing Cody/Codex adapter settings with the default model follows the same contract as saving default model settings: no explicit model key is sent. ## Linked Issues or Issue Description Bug report: - Summary: Testing a Cody/Codex local agent after selecting the default model could send an adapter config with an undefined model value instead of omitting the model key. - Expected behavior: Clearing the model to use the adapter default should test with `adapterConfig: {}` unless another model is explicitly selected. - Actual behavior: The UI test-config path could preserve `model: undefined`, causing the adapter test to fail instead of exercising the default model. - Related PRs: Reapplies the UI-only portion of #9361 after #9363 reverted the original PR. ## What Changed - Exported and reused `omitUndefinedEntries` so adapter test config payloads drop undefined adapter config entries before calling the test endpoint. - Hardened the current model display value so create-mode values that are nullish or non-string do not crash the model selector/test flow. - Added render coverage for editing a Codex agent back to the default model and for testing a create form with the default model. ## Verification - `pnpm exec vitest run ui/src/components/AgentConfigForm.render.test.tsx` - `pnpm check:token-gates` - Confirmed `git diff origin/master --name-only` contains only: - `ui/src/components/AgentConfigForm.render.test.tsx` - `ui/src/components/AgentConfigForm.tsx` - `ui/src/lib/agent-config-patch.ts` ## Risks Low risk. The change only removes `undefined` adapter config entries from the UI adapter-test payload and adds focused render coverage. Explicit model values and other adapter config fields are preserved. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent, tool-use enabled. Context window size not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>## Thinking Path > - Paperclip is an open-source agentic AI management platform that ships anonymous usage telemetry to understand product health and adoption. > - The telemetry system has a generated contract (`packages/shared/src/telemetry/generated/paperclip-telemetry.ts`) that types every first-party event the product emits. > - When a product change needs a new first-party event that is not yet in the generated contract, contributors had no public workflow explaining how to propose an event or later promote it into the contract once accepted. > - The gap leads to confusion at call sites: contributors either skip tracking entirely or emit untyped events that bypass the privacy and governance safeguards built into the contract. > - This PR fills that gap by adding `doc/TELEMETRY_WORKFLOW.md` — a public contributor guide that covers the full propose → promote lifecycle: the `@ts-expect-error -- proposed-telemetry(...)` marker, the canonical multi-line `track()` shape, the TS2578 expiry signal, and the look-up-by-event-name promotion step. > - `packages/shared/src/telemetry/README.md` gains a cross-reference so readers of the data-contract doc can find the workflow guide without searching. > - `README.md` gains a one-line pointer in the telemetry section so the workflow is discoverable from the project entry point. ## Linked Issues or Issue Description No pre-existing public GitHub issue covers this doc gap. Inline description: **Problem:** Contributors adding product telemetry for events not yet in the generated contract have no documented workflow. The `@ts-expect-error -- proposed-telemetry(...)` pattern exists in the codebase but is undocumented, leading to inconsistent usage and missing adoption signals. **Solution:** A new public guide (`doc/TELEMETRY_WORKFLOW.md`) documents the marker format, the recommended multi-line `track()` shape that preserves the TS2578 expiry signal, the dimension rules, and the promotion checklist. Cross-references are added to `packages/shared/src/telemetry/README.md` and the top-level `README.md`. No related open PRs found. ## What Changed - **New file `doc/TELEMETRY_WORKFLOW.md`** — public contributor guide for the propose/promote lifecycle: marker format, canonical multi-line example, TS2578 single-line trap, dimension rules, and step-by-step promotion checklist. - **`packages/shared/src/telemetry/README.md`** — added one-line cross-reference pointing at `doc/TELEMETRY_WORKFLOW.md` for proposed events not yet in the generated contract. - **`README.md`** — added one-line pointer in the telemetry section so the new workflow guide is reachable from the top-level project entry. ## Verification This is a docs-only change. Verification steps: 1. Open `doc/TELEMETRY_WORKFLOW.md` and confirm it contains: - The `^[a-z0-9][a-z0-9._:-]{1,63}$` event-name grammar. - The exact `// @ts-expect-error -- proposed-telemetry(<issue>): <rationale>` marker. - The multi-line `client.track()` copy-paste example (directive on line before event-name string). - The explanation of the TS2578 single-line trap and the look-up-by-event-name step. - The promotion checklist in the "Promote An Event" section. 2. Confirm `packages/shared/src/telemetry/README.md` cross-references `doc/TELEMETRY_WORKFLOW.md`. 3. Confirm the `README.md` telemetry section links to `doc/TELEMETRY_WORKFLOW.md`. ## Risks Low risk — docs-only change. No runtime behavior, schema, or existing telemetry emission is affected. The risk is that the guidance could diverge from the actual enforcement in the codebase over time; mitigated by linking to the generated contract and keeping the doc in the same repo. ## Model Used Claude — `claude-sonnet-4-6` (Anthropic Claude Sonnet 4.6). Tool use enabled. Extended context. Produced via the Paperclip agentic workflow with `Co-authored-by: Paperclip <noreply@paperclip.ing>`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents stream their run output live into the web UI, viewed per-run in the `AgentDetail` transcript viewer > - Browser tabs holding these views were sitting at 8–16 GB of memory footprint while their JS heap stayed at ~256 MB — a 60×+ gap, meaning the cost is in retained DOM / render objects, not JS objects > - The `LogViewer` in `AgentDetail.tsx` kept every streamed stdout/stderr line and structured event in unbounded React state and rendered each into a rich DOM block (the default "nice" mode has no virtualization), so a run streaming for hours grew an unbounded live DOM tree > - The sibling `useLiveRunTranscripts` hook (used by `IssueChatThread`) already bounds its buffers and virtualizes; `LogViewer` bypassed it and managed its own uncapped state — that inconsistency is the bug > - This pull request caps the live buffers and bounds the live DOM render, aligning `LogViewer` with the already-bounded path > - The benefit is that long-lived streaming tabs no longer grow without bound, cutting multi-GB tabs back to a bounded footprint ## Linked Issues or Issue Description No public GitHub issue exists; describing inline per CONTRIBUTING.md → "Link Issues or Describe Them In-PR", following the bug report template. **What happened?** Chrome's Task Manager showed multiple long-lived Paperclip tabs (agent-run / task views) each consuming 8–16 GB of memory footprint, while each tab's JS heap stayed at only ~150–256 MB. Memory grew monotonically the longer a run streamed. **Expected behavior** A tab viewing a live agent run should hold a bounded amount of memory regardless of how long the run streams. **Steps to reproduce** Open an agent run with a long-running / high-volume stream in `AgentDetail`, leave the tab open while output streams for an extended period, and watch the tab's memory footprint climb without bound in Chrome's Task Manager. **Paperclip version or commit** `3991a19a` (branch `fix/agent-run-transcript-memory`, off `master`). **Deployment mode** Local dev (`pnpm dev`), web UI. Not adapter-specific — core UI bug in the shared transcript viewer. ## What Changed - Add `ui/src/lib/live-log-buffer.ts`: a pure `appendCapped(prev, additions, max)` helper plus caps `MAX_LIVE_LOG_LINES=5000`, `MAX_LIVE_EVENTS=2000`, and `LIVE_TRANSCRIPT_RENDER_LIMIT=1500`, with rationale documented in the module. - Add `ui/src/lib/live-log-buffer.test.ts`: 6 unit tests (append, trim-to-cap, oversized batch, exact-cap, no-mutation, referential bail-out). - `ui/src/pages/AgentDetail.tsx` (`LogViewer`): route all four live-append sites (WebSocket log / progress / event, plus the poll fallback) through `appendCapped`, and pass `limit={LIVE_TRANSCRIPT_RENDER_LIMIT}` to `RunTranscriptView` for live runs so the "nice" view mounts only the most recent blocks. - The terminated-run "Load more log" pagination is deliberately left **uncapped** (guarded by `isLive`), so no historical output is lost — older output remains on the server and reachable there. ## Verification - `vitest run src/lib/live-log-buffer.test.ts` → 6/6 pass. - Existing suites `RunTranscriptView.test.tsx`, `AgentDetail.instructions.test.tsx`, `useLiveRunTranscripts.test.tsx` → 22/22 pass. - `tsc -b` (UI) → clean. - Manual/behavioral: live runs tail the last ~1500 blocks; terminated runs still render full history via "Load more log". Follow-up planned to profile before/after with the Chrome DevTools MCP. ## Risks Low risk. Changes only bound **in-memory state for live runs**; the terminated-run paginated path is untouched (still uncapped, guarded by `isLive`). No API, schema, or persistence changes. Worst case for a live run is that only the most recent 5000 lines / 1500 rendered blocks are visible in the tab — which is the intended "tail" behavior, and full history remains on the server. ## Model Used - **Provider:** Anthropic, via the Claude Code CLI. - **Model:** Claude Opus 4.8 (`claude-opus-4-8`). - **Reasoning mode:** Extended thinking enabled. - **Capabilities used:** tool use (shell execution, file editing), sub-agent fan-out for the codebase memory sweep, and the Chrome DevTools MCP for the diagnosis phase. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work (the ROADMAP "Memory" item is about company/agent knowledge, unrelated to this browser-tab memory fix) - [x] I have searched GitHub for duplicate or related PRs and linked them above (none found among open PRs) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have considered and documented any risks above - [ ] I have updated relevant documentation to reflect my changes (N/A — no user-facing docs affected; rationale is documented inline in `live-log-buffer.ts`) - [ ] All Paperclip CI gates are green (in progress at time of writing) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (the only open P2 was this missing template, which this update resolves; awaiting re-review) - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Humans sign in through the auth layer; every authenticated request parses the session/user profile with `currentUserProfileSchema` in `packages/shared` > - The schema requires `name` to be `null` or a non-empty string, but some identity providers hand back `name: ""` for users who never set a display name > - For those users the session payload fails validation on every request, so the app treats them as unauthenticated and bounces them to `/auth` in a loop — they can never get in > - This pull request preprocesses empty/whitespace-only names to `null` before validation, so the existing `min(1).max(120).nullable()` rule still holds for real names > - Review found the sibling `email` field has the same failure mode (the DB `auth` schema declares `email` as `notNull`, so a provider that supplies no email stores `""`, which `z.string().email()` rejects); the same preprocess is applied there > - The benefit is that users whose provider reports an empty name (or email) can sign in normally instead of being locked out, with no change in behavior for anyone else ## Linked Issues or Issue Description No existing issue; described in-PR following the bug report template: **What happened?** Users whose auth provider returns `name: ""` (empty string) in the profile payload fail `currentUserProfileSchema` / `authSessionSchema` parsing (`name: z.string().min(1)...`). The parse failure makes the session look invalid and the UI redirects to `/auth` on every attempt — an endless sign-in loop. The `email` field has the same failure mode (`z.string().email()` rejects `""`). **Expected behavior:** An empty display name (or email) should be treated the same as a missing one (`null`); the user should be signed in normally. **Steps to reproduce:** Sign in with an account whose upstream identity record has an empty-string name (or set a user's `name` column to `''` directly), then load the app: session parse fails and you are bounced back to `/auth`. **Adapter(s) involved:** Not adapter-specific (core bug). **Deployment mode / version:** Any; reproduces on current `master`. ## What Changed - `packages/shared/src/validators/access.ts`: `currentUserProfileSchema.name` now runs through `z.preprocess` that coerces empty or whitespace-only strings to `null` before the existing `z.string().min(1).max(120).nullable()` validation. - `packages/shared/src/validators/access.ts`: the same preprocess is applied to `email` (review follow-up): `users.email` is `notNull` in the DB schema, so a provider without an email stores `""`, which `z.string().email()` rejects — the identical lockout loop. Empty/whitespace-only emails now coerce to `null` (the field was already nullable); malformed non-empty emails are still rejected. - `packages/shared/src/validators/access.test.ts` (new): covers empty-string → `null`, whitespace-only → `null`, real values preserved, `null` preserved, and malformed non-empty email still rejected — for both `name` and `email`, and the same cases through `authSessionSchema`. ## Verification - `vitest run src/validators/access.test.ts` in `packages/shared` — 13 tests pass. - `tsc --noEmit -p packages/shared` passes. - Manual: parse `{ id, email: "", name: "", image: null }` with `currentUserProfileSchema` — succeeds with `name: null` and `email: null` instead of failing validation. ## Risks - Low risk. The change only widens accepted input (empty/whitespace string → `null` for `name` and `email`); every previously valid payload parses identically. `updateCurrentUserProfileSchema` (user-initiated rename) is untouched and still rejects empty names. ## Model Used Claude Fable 5 (Anthropic, `claude-fable-5`, agentic coding harness via Claude Code, extended reasoning enabled). Original fix drafted with Claude Sonnet 4.6. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `/pr-gardening` skill drives a bundled agent that scans a company's issues for those linked to open GitHub PRs, then reports on their state; it relies on the server's company-search **extract** endpoint to pull PR references out of issue bodies > - Two gaps surfaced during end-to-end QA of the gardening workflow: the extract service silently ignored a per-issue match cap, so callers could not bound how many matches came back per issue, and the skill's candidate-discovery scripts fell over on large repos and on issues that referenced deleted PRs > - Left unaddressed, the gardener either truncated its scan unpredictably or aborted outright, so it could not reliably enumerate PR candidates > - This pull request honors an explicit `matchesPerIssue` limit in the extract search API and hardens the skill's candidate discovery against missing/unavailable PRs and oversized `gh` output > - The benefit is a PR-gardening workflow that scans deterministically and finishes cleanly on real-world companies ## Linked Issues or Issue Description No pre-existing public GitHub issue — describing the bug in-PR following the bug report template (`.github/ISSUE_TEMPLATE/bug_report.yml`). ### What happened? The company-search extract endpoint accepted a per-issue match limit but did not apply it, returning matches capped only by the old hardcoded constant regardless of the caller's request. Separately, the `/pr-gardening` skill's candidate-discovery scripts crashed when a scanned issue referenced a deleted PR (GitHub `Not Found (HTTP 404)` / GraphQL `Could not resolve to a PullRequest`) and could exceed the default `gh` output buffer on large result sets, aborting the whole scan. ### Expected behavior The extract API bounds matches per issue when a caller passes `matchesPerIssue` (default 20, max 200), and omitting it preserves the previous default. The gardening scripts skip PRs that are deleted/unavailable and tolerate large `gh` responses without aborting the scan. ### Steps to reproduce 1. Call the company-search extract endpoint with a `matchesPerIssue` value against an issue containing many PR references — previously the value was ignored. 2. Run the pr-gardening candidate scan against a company whose issues reference a since-deleted PR — previously the scan threw instead of skipping that PR. ### Paperclip version or commit `master` at the base of this PR (branch cut from current `origin/master`). ### Deployment mode Local Paperclip instance / self-hosted. ## What Changed - **Extract search honors `matchesPerIssue`**: added the `matchesPerIssue` field to the shared search validator/types and applied the cap in `company-search-extract` so results are bounded per issue (`packages/shared`, `server/src/services/company-search-extract.ts`, `doc/SPEC-implementation.md`). - **Hardened pr-gardening candidate discovery**: `find-candidates.mjs` / `lib.mjs` now request `matchesPerIssue=200`, treat missing/unavailable PRs (deleted PR → `isMissingPullRequestError` / `unavailable`) as skips instead of fatal errors, and raise the `gh` `maxBuffer` to 50 MB for large repos. - **Tests**: expanded `company-search-extract-{routes,service}.test.ts` for the new limit and added coverage in `pr-gardening.test.mjs`. ## Verification Re-run on a fresh worktree cherry-picked onto current `master`: - `node --test .agents/skills/pr-gardening/scripts/pr-gardening.test.mjs` → 8/8 pass - `pnpm vitest run server/src/__tests__/company-search-extract-routes.test.ts server/src/__tests__/company-search-extract-service.test.ts` → 10/10 pass ## Risks Low risk. `matchesPerIssue` is optional and backward-compatible (omitting it preserves prior behavior). The skill changes only add skip/tolerance paths and a larger buffer; no schema or migration changes. ## Model Used Claude — Opus 4.8 (`claude-opus-4-8`), extended thinking, tool use / code execution via the Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - When a run is stranded (process lost, adapter failure, a finished run with no disposition, an over-eager inactivity kill), the harness opens a *recovery action* and wakes an owner to recover it > - Recovery volume regressed sharply in one week — 3.26% of all runs vs a ~1.2% monthly norm, 5–8x the prior volume — and nobody noticed until it was ~194 actions deep, because there was no way to *see* the recovery rate > - We also could not see which causes drive recovery, nor how often a manager ends up doing the deliverable work themselves instead of handing it back to the original owner (the product goal is that managers doing the work stays rare) > - This pull request adds a recovery-observability report + API endpoint: weekly rate normalized per run, a threshold alert, the cause taxonomy live from the ledger, and the handed-back vs owner-completed ratio and per-cause routing outcomes > - The benefit is that a recovery regression like that week is caught by a threshold instead of by a human noticing it by feel, and each recovery playbook row can be verified in production ## Linked Issues or Issue Description **Feature.** **Problem or motivation** Recovery takeovers are a first-class exception path (`issue_recovery_actions`), but there is no aggregate view of them. A week where the recovery rate tripled went unnoticed until it was deep. There is no signal for (a) the per-run recovery rate over time, (b) which cause + run error code drives it, or (c) whether the recovery owner hands the task back to the original assignee or ends up doing the deliverable work themselves. **Proposed solution** A read-only report service and `GET /companies/:companyId/recovery-observability` endpoint that surfaces the weekly rate, a threshold alert, the cause taxonomy, the hand-back ratio, and per-cause routing outcomes. **Alternatives considered** Adding `handed_back` / `owner_completed` to the recovery-action outcome vocabulary and writing them at resolution time. Rejected for this change: the distinction is derivable from the recovery owner, the recorded return owner, and where the source issue actually landed, so the report works against all historical data without a backfill. **Roadmap alignment** Implements the recovery-observability line of the approved recovery-takeover plan (make regressions visible via a threshold rather than by human feel); no schema or write-path change. ## What Changed - Add `server/src/services/recovery-observability.ts`: - `recoveryObservabilityService(db).report(companyId, { weeks, thresholdPercent, now })` returns weekly rates (recovery actions / runs, Monday-anchored to match the retrospective), a `cause` + `latestRunErrorCode` breakdown, a handed-back vs owner-completed summary, and per-cause routing outcomes. - `evaluateRecoveryRateAlert(weekly, thresholdPercent)` — a pure function (default threshold 2% of runs) returning the breached weeks and whether the latest week regressed. - `classifyRecoveryHandoff(...)` — a pure classifier deriving `self_recovery` / `handed_back` / `owner_completed` from the recovery owner, return owner, and final issue landing. - Add `GET /companies/:companyId/recovery-observability` (optional `weeks` and `threshold` query params) to the existing dashboard router. - The `weeks` window is bounded (`MAX_WINDOW_WEEKS = 104`, service-authoritative and re-clamped at the route) so a large query value can't over-allocate the per-week array. - Add tests: unit coverage for the alert and the classifier, plus an embedded-Postgres integration test that seeds synthetic runs and recovery actions crossing 2% and asserts the alert fires and the hand-back ratio is computed. ## Verification - `CI=1 NODE_ENV=development npx vitest run server/src/__tests__/recovery-observability.test.ts` — 10/10 pass (includes the synthetic 2%-crossing alert case and the hand-back ratio case). - Rendered against a live database of 300+ recovery actions: the weekly rates reproduce the retrospective (e.g. 1.37% / 1.53% / 0.65% / 0.86% / 1.37% for early-June weeks), the alert fires on the two most recent weeks (3.15% and 3.05%, both over 2%), and the hand-back summary shows owner-completed ≈ 73% vs handed-back ≈ 27% — matching the observed "managers keep ~80% of takeovers". ## Risks - Low risk. Read-only: adds one GET endpoint and a service; no schema, migration, or write-path changes. The hand-back classification reads the source issue's current assignee/status, so a much-later reassignment could reclassify a historical action — acceptable for an aggregate trend view. ## Model Used - Claude, `claude-opus-4-8` (Opus 4.8), extended thinking, tool use / code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - When an agent's task is `blocked`, humans often comment on the issue thread expecting it to reopen and resume > - If the issue stays `blocked` because an unresolved (not-done) blocker still gates the reopen path, the UI said nothing — the human "sent a message and nothing happened" (a real silent-failure report) > - That silence makes the product feel broken even though the gate is working as designed > - This pull request adds Rule C copy to the existing `IssueBlockedNotice` surface: when a comment won't reopen a blocked issue, it explains why and names the deepest unresolved blocker leaf with its status > - The benefit is that the human immediately understands the task is *paused, not stuck*, and knows exactly which task to act on to unblock it ## Linked Issues or Issue Description <!-- No public GitHub issue — describing the underlying problem inline (path B), following the feature issue template. --> **Problem or motivation** A human comments on a `blocked` issue expecting it to move back to `todo` and resume. When the reopen gate keeps it blocked because an unresolved (not-`done`) blocker remains, nothing in the UI communicates this, so the user perceives a dropped message ("I sent a message and nothing happened"). **Proposed solution** Reuse the existing amber `IssueBlockedNotice` surface (the designated blocked/recovery surface — no new component). When a message won't reopen a blocked issue, the notice states that it won't reopen yet, names the unresolved blocker leaf with its status (e.g. "Still blocked by PAP-XXXXX (in progress)"), and reassures that it reopens automatically once the blocker is done. UI-only; no server change. **Alternatives considered** Adding a server signal (a reopen-suppressed reason on the comment/notice payload) was considered but rejected as unnecessary — the unresolved-blocker set is already available client-side, so the copy is derived in the component. Done-but-pending-finalize blockers are `done`, so they fall out of the unresolved set into the standard reopen (Rule B) path and are correctly not shown as reopen-suppressed. ## What Changed - `ui/src/components/IssueBlockedNotice.tsx`: added reopen-suppressed messaging for `blocked` issues that still have unresolved blockers — a lead sentence ("a message won't reopen it yet, then it reopens automatically"), the named unresolved blocker leaf with its status, and an "and N other task(s)" summarization when multiple blockers remain. - `ui/src/components/IssueBlockedNotice.test.tsx`: added/updated tests covering the single, nested-chain (deepest-leaf), multiple-blocker, and empty-blocker states, plus the not-a-reopen-case (`in_progress`) path. - `ui/storybook/stories/issue-blocked-notice.stories.tsx`: new Storybook stories rendering each notice state for copy review. ## Verification - `cd ui && tsc -b` — typecheck clean (previously failed TS2322 on the story meta; fixed by a default `args`). - Vitest: `IssueBlockedNotice.test.tsx` green (single / nested / multiple / empty / in-progress states). - Storybook stories rendered at 1440×900 for all four states; screenshots posted on the tracking issue. - UXDesigner reviewed the notice copy and signed off (no copy edits required). ## Risks Low risk. UI-only, additive copy on an already-shipped amber notice surface; no server or schema changes. The reopen behavior itself is unchanged — this only explains the existing gate. Worst case is copy wording, which had a design review. ## Model Used Claude — claude-opus-4-8 (Opus 4.8), extended thinking, with tool use / code execution in the Paperclip harness. --- - [x] I searched the GitHub PR list for similar/duplicate PRs before opening this one. --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>## Thinking Path > - Paperclip's sandbox managed runtime is responsible for provisioning the agent's execution environment — it extracts a home directory asset into the sandbox before the adapter runs. > - The sandbox runtime core was directly branching on the adapter key (`codex`) to decide which merge scripts to stage and which merge-extract command to run, coupling generic infrastructure to a specific adapter's credential-merge protocol. > - This makes it harder to add, remove, or modify per-adapter asset provisioning without touching the runtime core; it also prevents other adapters from contributing staged files or a custom extract command at all. > - The fix is to move the adapter-specific knowledge into the adapter itself: the asset descriptor gains optional `provision` (stageFiles + extractCommand) and `restore` contribution fields that any adapter can populate, and the runtime core consumes them generically. > - This pull request introduces those contribution fields, wires the Codex adapter's inbound credential-merge as a `provision` contribution, and removes the adapter-specific branching from the runtime core. > - The benefit is a clean seam: the runtime core is now adapter-agnostic for asset provisioning, the inbound behavior is unchanged (same merge matrix, same scripts), and other adapters can attach custom staged files or extract commands without modifying shared infrastructure. ## Linked Issues or Issue Description No pre-existing public GitHub issue. Describing the problem inline per the feature template: **Problem or motivation** The sandbox managed-runtime asset provisioning in `sandbox-managed-runtime.ts` branched directly on the adapter key (`codex`) to decide which merge scripts to stage and which shell command to use during asset extraction. This tight coupling prevents other adapters from customizing their provisioning without modifying the runtime core, and it means the runtime core must import and know about adapter-specific merge scripts. **Proposed solution** Add an optional `provision` contribution (array of `stageFiles` entries + an `extractCommand` string) and an optional `restore` contribution to the asset descriptor returned by adapters. The runtime core now consumes these generically — if a `provision` contribution is present, it stages those files and uses the supplied command; otherwise it falls back to the default `tar -xf` extraction. The Codex adapter populates the `provision` contribution where it previously depended on core branching. **Alternatives considered** Keeping the adapter-specific logic in the core as a documented exception; rejected because it makes the seam inextensible. **Roadmap alignment** Decoupling — removes a latent coupling between the runtime core and a specific adapter. ## What Changed - Added `provision` contribution field (`stageFiles: Array<{src, dest}>` + `extractCommand: string`) to the `SandboxManagedRuntimeAsset` descriptor type in `adapter-utils`. - Added `restore` contribution field (hook for post-restore logic, populated in a later phase) to the descriptor. - Removed adapter-key branching (`if adapterKey === 'codex'`) from the runtime core in `sandbox-managed-runtime.ts`; the core now reads `provision.stageFiles` and `provision.extractCommand` generically. - Extracted Codex-specific merge-script paths and the merge-extract command into `codex-auth-merge-scripts.ts` in `adapter-utils`; the Codex adapter's `execute.ts` now attaches them as a `provision` contribution when it builds its managed-home asset descriptor. - Updated `execution-target.ts` to pass the extended asset type through to the adapter call site so the new fields are load-bearing end-to-end. - Added seam-proving unit tests in `sandbox-managed-runtime.test.ts`: contribution-less asset uses the default path; a non-adapter asset round-trips the generic provision+restore seam; a structural assertion verifies the runtime core carries no Codex-specific string literals. - Added one test in `workspace-restore-merge.test.ts` confirming the inbound merge matrix is unaffected. ## Verification ```bash # Unit tests (20 pass): npx vitest run packages/adapter-utils/src/sandbox-managed-runtime.test.ts packages/adapter-utils/src/workspace-restore-merge.test.ts # Type-check both affected packages: cd packages/adapter-utils && npx tsc --noEmit cd packages/adapters/codex-local && npx tsc --noEmit # Structural: runtime core carries no adapter string literals grep -n 'codex\|auth\.json' packages/adapter-utils/src/sandbox-managed-runtime.ts # Expected: zero matches ``` ## Risks **Low risk.** This is a behavior-preserving refactor: the inbound provisioning output (which files get staged, which command runs) is identical to before, now driven by the adapter-supplied contribution instead of core branching. The existing inbound merge matrix tests are the regression guard. No change to which bytes cross the sandbox boundary. The SSH transport is untouched. ## Model Used - **Provider:** Anthropic - **Model:** Claude Sonnet 4.6 (`claude-sonnet-4-6`) - **Context window:** 200 K tokens - **Capabilities used:** tool use (file read/edit, bash execution, Paperclip API), extended reasoning over multi-file TypeScript refactor - **Mode:** agentic (Paperclip ACPX platform) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Harold Kim <harold@paperclip.ing>c27745e585317a7d4e4d317a7d4e4d6a86b5a49d6a86b5a49db060056720b060056720d7b9da138cd7b9da138c2e338ead8cView command line instructions
Manual merge helper
Use this merge commit message when completing the merge manually.
Checkout
From your project repository, check out a new branch and test the changes.