Since FRR 10.2.2 and 10.3.0, bgpd runs vpn_leak_postchange_all() each time router bgp <ASN> (the default instance) is entered, not only when the instance is created. On a router whose VRFs import from a large table (import vrf default puts the default table into the VPN RIB; VRF <-> VPN leaking imports from it directly), every entry re-imports the whole VPN RIB once per VRF and blocks bgpd for many seconds.
frr-reload sends each changed line as its own router bgp ... block, once in pass 0 and twice in pass 1, and VyOS re-sends some lines on every reload because FRR does not print them back (T9450). So a commit pays this cost tens of times even when nothing in BGP changed.
With 1.09M eBGP routes in the default VRF and three VRFs importing from it through a route-map (VyOS 1.5, FRR 10.5.2), each router bgp <ASN> entry takes 17-25 s of bgpd CPU and is logged as CPU HOG: command took 24904ms (cpu time 24902ms): router bgp <ASN>. A reload that sends 54 default-instance blocks (18 in pass 0, 36 in pass 1) then takes about 23 minutes, almost all of it in these entries (show event cpu: vtysh_read far above every other bgpd event).
Reproduced on VyOS 1.5.1 (FRR 10.5.2) with three VRFs importing from the default VRF through a route-map that accepts only a default route, and routes injected into the default VRF:
| routes in the default VRF | one router bgp <ASN> entry | same, import vrf default removed | commit that changes nothing in BGP | same, imports removed |
|---|---|---|---|---|
| 0 | 0.02 s | - | 7.7 s | - |
| 100k | 1.02 s | 0.06 s | 61.9 s | 9.9 s |
| 300k | 3.12 s | 0.15 s | 176 s | 14.8 s |
The cost per entry grows linearly with the table (about 10 us per route with three importing VRFs on a 4 vCPU VM) and disappears without the VRF imports. The VRF tables stayed empty (the import route-map accepts only a default route): the time is the walk itself.
Cause.
bgpd/bgp_vty.c, router_bgp:
- The code: if (inst_type == BGP_INSTANCE_TYPE_DEFAULT) vpn_leak_postchange_all(); runs on every entry of the default instance, although the comment above it says "If we just instantiated the default instance".
- Where: FRR 10.5.2 lines 1644-1650; FRR 10.6.1 lines 1686-1694; master (2026-10-09) has the same condition.
- 2018: ecec94950f ("Fix bgpd doing vpn_leak_postchange_all() every time "router bgp ASNUM" command is entered in vtysh") fixed this by adding an is_new_bgp check.
- 2024-12-31: 9f7177af13 ("bgpd: fix duplicate BGP instance created with unified config") replaced is_new_bgp && with bgp &&, so the check is gone. In FRR 10.3.0 and later; backported as ac31df3758 to 10.2.2 and later 10.2.x.
Possible fix.
A patch to router_bgp (FRR 10.5.2) runs vpn_leak_postchange_all() only when the command created the default instance or takes over a hidden one (VRF leak config entered before router bgp, which 3bd70bf8f3 added the pass for and which now auto-creates a hidden default instance, or a default instance kept hidden after no router bgp); 9f7177af13's lookup is kept. Tested on VyOS 1.5.1 with 300k routes and the three VRF imports:
| stock | patched | |
|---|---|---|
| one router bgp <ASN> entry | 3.12 s | 0.02 s |
| commit that changes nothing in BGP | 176 s | 9.07 s |
Leaking still follows the configuration: an injected default route reaches all three VRFs, removing one VRF's import removes it from that VRF only, re-adding brings it back; the leaks are back after systemctl restart frr and after a reboot.
Related: T9450 (VyOS re-sends lines FRR does not print back, so commits send these blocks at all) and T9445 (frr-reload's pass 1 sends every pass-0 line again, so pass 1 has twice the blocks: 54 instead of 36 per reload here).