Bug present since the T9278 lease-event reconciliation landed
Reconciling the FRR configuration after a DHCP lease event is requested from the configuration daemon over the same socket a commit uses. The daemon serves one client at a time and assumes that client is the running commit.
A lease event arrives while a commit holds the configuration lock, and the request is then sent the very moment that lock is released - so it collides with
the commit that follows. Two ways this goes wrong:
- The request is fair-queued into the middle of the multi message handshake which hands the active and session configuration to the daemon, where it is consumed as configuration data. The daemon is left without a usable session and answers every script of that commit with a daemon error.
- The daemon is still busy reloading FRR when the next commit starts, and does not answer within the 10ms the commit waits for it.
Either way the commit falls back to executing the conf-mode scripts directly. Those never render FRR, only the daemon does. The commit succeeds, no error is shown, and FRR keeps the configuration rendered before it. Nothing recovers until the next lease event.
Steps to reproduce
On an image with an interface getting its address over DHCP:
set interfaces ethernet eth0 address dhcp commit
wait for the lease
set protocols static route 10.10.0.0/16 dhcp-interface eth0 commit
The second commit has to start while the daemon is serving the request caused by the lease event. The window is short, so this reproduces intermittently; a smoketest run hits it regularly.
Expected behavior
show ip route / vtysh -c "show running-config" contains the static route pointing at the gateway of the lease.
Actual behavior
The commit succeeds without any error, but FRR only has the configuration rendered by the previous commit:
ip route 0.0.0.0/0 192.0.2.2 eth0 tag 210 210 !
The static route is missing and never appears.