rdmbair15m5-changelog-20260921-1710-agent-coordinator-silent-stall-detection-v1.10
rdmbair15m5-changelog-20260921-1710-agent-coordinator-silent-stall-detection-v1.10
Resolved defect in agent coordinator Job 2 where agy background housekeeping writes artificially kept log modification timestamps fresh, permanently masking silent stalls (ISSUE-20260921-15). Delivered agent-coordinator v1.10, verified 35/35 unit tests, validated against live fleet sessions, and deployed updated LaunchAgent daemon.
Scope
- Hosts:
rdmbair15m5(local test & deployment),rdmsm4x(canonical hub). - Projects:
fleet/agent-coordinator/,issues(ISSUE-20260921-15).
Summary of Actions Taken
1. Root Cause & Architectural Rectification
classify()previously returnedSILENTwhennow - log mtime > STALL_MIN (45 min). agy writes routine background housekeeping lines (http_helpers.go,mcp_manager.go,browser.go,rules.go) every few minutes regardless of user activity or task execution, keepingmtimeunder 5 minutes indefinitely.- Rectified
classify(s, li, t=None)to compute inactivity age from actual work events:last_act = max(li["last_model"] or 0, li["last_input"] or 0), falling back toli["mtime"]only for newly initialized logs with zero model or input records. - If
act_age > STALL_MIN * 60and the session is not at a quota wall, it is accurately classified asSILENT. - Updated
s["log_age_min"]to reflect actual elapsed minutes sincelast_act, and addeds["mtime_age_min"]to retain raw file modification age in JSON telemetry. - Updated finding detail string from
log silent ...tosilent ....
2. Automated Test Suite Expansion
- Added
CoordinatorSilentStalltest suite intest_agent_coordinator.pyfeaturing 7 test cases:test_housekeeping_mtime_does_not_mask_silent_stall: Reproduces the exact 2026-09-21 failure mode (model call 60m ago, housekeeping 2m ago) and proves classification returnsSILENT.test_recent_model_call_within_stall_window_is_idle: Proves model call 30m ago evaluates toIDLE.test_very_recent_model_call_is_working: Proves model call 45s ago evaluates toWORKING.test_recent_user_input_is_idle: Proves user input 15m ago evaluates toIDLE.test_stale_user_input_without_model_response_is_silent: Proves input 50m ago without model response evaluates toSILENT.test_empty_log_falls_back_to_mtime: Validates mtime fallback for uninitialized logs.test_quota_session_is_quota_not_silent: Validates that sessions at a quota wall remainQUOTArather thanSILENT.
- Executed full test suite: 35/35 tests passing cleanly in 0.016s on Python 3.14.7.
3. Canonical Hub Sync & Deployment
- Committed changes to canonical hub
rdmsm4x:~/dev/fleet(d48b5d2):agent_coordinator.py(v1.10),test_agent_coordinator.py,README.md, andSESSION-STATE.md. - Pushed to
fleet(git.ecs0.net:git/root/fleet.git) and GitHub backup (https://github.com/richhdoty/rdmsm4x-dev-fleet.git). - Deployed and verified LaunchAgent
com.eastcoastscience.agentcoordacrossrdmbair15m5andrdmsm4x. - Marked ticket
ISSUE-20260921-15asresolved.
Verification Evidence
python3 test_agent_coordinator.py: 35/35 tests passed (0.016s).- Live dry run on
rdmbair15m5(python3 agent_coordinator.py --dry-run):- Detected silent stalls on
2688c574(silent 61.6m onttys010) and1055213a(silent 67.9m onttys013), accurately surfacing findings. - Correctly maintained
WORKINGstatus on active sessions (25880dee,d18ca6ce).
- Detected silent stalls on
- LaunchAgent reload:
com.eastcoastscience.agentcoordverified active on bothrdmbair15m5andrdmsm4x.
Outstanding Owner Actions
- None.