Summary
vyos-http-api.service can stay active while every POST /configure returns HTTP 500:
ConfigSessionError: calling validateSetPath() without config session ConfigSessionError: calling discardChanges() without config session
The process holds one long-lived ConfigSession created at startup. When the
underlying cli-shell-api session dies (failed teardown, unionfs mess, race after
a failed/partial apply), subsequent set() calls fail. The error path always
calls session.discard(), which fails the same way and becomes the ASGI 500 —
not a clean client error. Only systemctl restart vyos-http-api recreates the
session.
This is distinct from T2292 (graceful API shutdown / session teardown). Here the
service remains up and healthy from systemd's point of view.
Why it matters
- Automation that uses REST /configure (batch apply, CI, fleet tools) hits a permanent wedge: process up, data plane fine, every configure fails until an operator restarts the unit.
- Production (VyOS 1.5 rolling, 2026-07-24): observed on two routers the same day after failed/partial applies:
- goblin ~18:28–18:31 PDT — restart fixed
- borg ~18:03–18:05 PDT — restart fixed Peer routers with little configure traffic that day did not show the error.
- /config-file save may still return 200 while /configure is dead, which is confusing when operators try "save then re-apply".
Reproduction (observed class)
- Enable service https API with a long-lived process (default).
- Drive enough /configure traffic that a config session can desync (exact trigger still under investigation; seen after failed/partial apply waves).
- Observe: unit active; every /configure → 500 with "without config session"; discardChanges also without session.
- systemctl restart vyos-http-api → /configure works again.
Suggested fix
A. Harden ConfigSession (python/vyos/configsession.py):
- session_exists() via cli-shell-api inSession
- ensure_session(): if dead, best-effort teardownSession + setupSession
- discard(): on "without config session", no-op/log (do not re-raise)
B. REST /configure (api/rest/routers.py _execute_configure_op):
- Under the global lock, ensure_session (or recreate ConfigSession) before ops
- On session-lost ConfigSessionError: safe discard; recover for next request; return 503 with a clear retriable body, not opaque 500
- Optional: one retry of the same request after successful re-setup when no partial sets have been applied yet
C. call_commit / call_commit_confirm: same safe discard
Prefer A+B over a full per-request session redesign (works with GraphQL sharing
SessionState; backportable to rolling).
Environment
- VyOS 1.5 Circinus rolling (fleet image 2026.07.21 at time of report)
- HTTPS API bound to LAN INT1; key auth; batch /configure automation
Workaround
sudo systemctl restart vyos-http-api.service
then retry /configure.
Client-side: detect "without config session" in error body, restart unit via
SSH, retry once (interim until rolling ships the fix).