Severity: High — can cause complete loss of management access, recoverable only from the console/IPMI.
Affected: VyOS 2026.07.21-1151-rolling, kernel 6.18.x, vyos-intel-ice 2.6.7. Believed to affect all rolling images since T9010 (2026-06-22: kernel 6.18 + Intel OOT driver refresh).
Components: interface naming & discovery (udev vyos_net_name, boot-time rescan), interface MAC handling, boot config loader. Trigger: kernel / Intel out-of-tree ice driver.
Summary
On a SuperMicro system whose Intel E823-L 25G SFP28 ports (driver ice) are unused with empty cages, those ports initialize slowly (DDP package load plus a firmware "module not present" health event). During boot they lose VyOS's interface-naming race and capture low ethN names that the configuration has pinned (via hw-id) to other, in-use ports — here, copper RJ45 ports on igb.
The consequences cascade:
- Unused ports end up named eth0/etc.; the configured copper ports are displaced and left unnamed.
- VyOS programs each interface's MAC from its configured hw-id, so the copper port's MAC is written onto the empty ice port — producing two interfaces with the same MAC, and putting the configured IP on a dead, empty cage.
- Interface naming becomes non-deterministic across reboots.
- Correcting this by pinning the displaced port to a new name fails at the CLI (Interface "ethX" does not exist!); doing it via config.boot + reboot caused a complete loss of management access, because the boot loader applies the whole configuration as a single transaction and treats a commit failure as unrecoverable.
The ice "Module is not present" log lines are expected for empty cages (a firmware health-status event) and are not themselves the defect — but the slow initialization they accompany is what exposes the naming race.
Environment
- Product: VyOS 2026.07.21-1151-rolling
- Kernel: 6.18.x
- Driver: vyos-intel-ice 2.6.7 (out-of-tree). Note: vyos-build pins ice at commit_id = v2.6.6, but the Intel-driver build step resets the source tree to origin/main before building, so the packaged version can drift from the declared pin (the running driver reports 2.6.7). Worth verifying for build reproducibility.
- NIC firmware (E823-L): fw 5.5.14 api 1.7.9 nvm 2.19 0x8000d221 1.3083.0
System / hardware
- System: SuperMicro SYS-510D-10C-FN6P
- Board: SuperMicro X12SDV-10C-SP6F
- CPU: Intel Xeon D-1747NTE (Ice Lake-D)
NIC inventory (8 ports)
| PCI | Controller | Driver | Permanent MAC | Media | In use |
|---|---|---|---|---|---|
| 0000:02:00.0 | Intel I210 [8086:1533] | igb | 3c:ec:ef:d1:87:9a | Copper RJ45 | yes |
| 0000:03:00.0 | Intel I210 [8086:1533] | igb | 3c:ec:ef:d1:87:9b | Copper RJ45 | yes |
| 0000:04:00.0 | Intel I350 [8086:1521] | igb | 3c:ec:ef:d1:87:9c | Copper RJ45 | yes |
| 0000:04:00.1 | Intel I350 [8086:1521] | igb | 3c:ec:ef:d1:87:9d | Copper RJ45 | yes |
| 0000:15:00.0 | Intel XXV710 [8086:158b] | i40e | 3c:ec:ef:d8:16:ee | SFP28 | empty |
| 0000:15:00.1 | Intel XXV710 [8086:158b] | i40e | 3c:ec:ef:d8:16:ef | SFP28 | empty |
| 0000:f4:00.0 | Intel E823-L for SFP [8086:124d] (SoC) | ice | 3c:ec:ef:d1:88:a2 | SFP28 | empty |
| 0000:f4:00.2 | Intel E823-L for SFP [8086:124d] (SoC) | ice | 3c:ec:ef:d1:88:a3 | SFP28 | empty |
The two ice ports are the Xeon-D SoC-integrated E823-L. The two i40e ports are an add-in/AIOM XXV710. Only the four igb copper RJ45 ports are cabled/used.
Steps to reproduce
- VyOS rolling (kernel 6.18 / vyos-intel-ice 2.6.x) on a host with at least one Intel 800-series (ice) port that has no media (empty SFP cage / unconnected), alongside other NIC families (igb/i40e).
- Configure interfaces ethernet ethN pinned by hw-id, with addresses on the non-ice ports.
- Reboot, ideally several times.
Expected behavior
Each configured interface deterministically binds to the physical port whose permanent MAC equals its hw-id, regardless of driver initialization order; configured addresses land on the intended ports; naming is stable across reboots.
Actual behavior
- The empty/slow ice port(s) capture low ethN names; configured copper ports are displaced and disappear from eth*.
- The copper port's hw-id is programmed onto the empty ice port, so the configured IP lands on a dead cage and two interfaces end up with the same MAC.
- Naming varies across reboots.
- Pinning the displaced port via the CLI fails: Interface "eth2" does not exist!
- Pinning via config.boot + reboot results in a complete, unreachable boot.
Evidence
dmesg (trimmed):
ice: Intel(R) Ethernet Connection E800 Series Linux Driver - version 2.6.7 ice 0000:f4:00.0: fw 5.5.14 api 1.7.9 nvm 2.19 0x8000d221 1.3083.0 [8086:124d] [15d9:124d] ice 0000:f4:00.0: Failed to apply Tx scheduling configuration, err -95 ice 0000:f4:00.0: The DDP package was successfully loaded: ICE OS Default Package version 1.3.59.0 ice 0000:f4:00.0 eth0: Module is not present. ice 0000:f4:00.2 eth2: Module is not present. ice 0000:f4:00.2 eth3: renamed from eth2
Observed naming vs. permanent MAC (key evidence) — eth0 is the empty ice port f4:00.0, but its current MAC was overwritten to the copper port's hw-id, while its permanent MAC is unchanged:
eth0 pci=0000:f4:00.0 cur=3c:ec:ef:d1:87:9a perm=3c:ec:ef:d1:88:a2 drv=ice carrier=0 eth1 pci=0000:03:00.0 cur=3c:ec:ef:d1:87:9b perm=3c:ec:ef:d1:87:9b drv=igb carrier=1 eth3 pci=0000:f4:00.2 cur=3c:ec:ef:d1:88:a3 perm=3c:ec:ef:d1:88:a3 drv=ice carrier=0 eth4 pci=0000:04:00.0 cur=3c:ec:ef:d1:87:9c perm=3c:ec:ef:d1:87:9c drv=igb carrier=1 eth5 pci=0000:04:00.1 cur=3c:ec:ef:d1:87:9d perm=3c:ec:ef:d1:87:9d drv=igb carrier=1 eth6 pci=0000:15:00.0 cur=3c:ec:ef:d8:16:ee perm=3c:ec:ef:d8:16:ee drv=i40e carrier=0 eth7 pci=0000:15:00.1 cur=3c:ec:ef:d8:16:ef perm=3c:ec:ef:d8:16:ef drv=i40e carrier=0
The copper port that config pins to eth0 (permanent MAC ...87:9a, PCI 0000:02:00.0) is absent from eth* — displaced to a temporary name. eth0 (the empty ice cage) shows cur=...87:9a (the copper hw-id) with perm=...88:a2 — the hw-id was programmed onto the wrong port, colliding with the stranded copper port's real MAC.
Configuration: interfaces ethernet pinned by hw-id for eth0, eth1, eth3, eth4, eth5, eth6, eth7 (7 of 8 ports; the second ice port, permanent MAC ...88:a2, was not pinned).
Suspected mechanism
- Slow init on empty ice ports — DDP package load plus the firmware health event (ICE_AQC_HEALTH_STATUS_ERR_MOD_NOT_PRESENT → "Module is not present") make these ports driver-ready later than the igb/i40e ports.
- The naming rules require the driver link to be ready — VyOS's persistent naming (udev rules 62-temporary-interface-rename.rules and 65-vyos-net.rules, both matching DRIVERS=="?*", invoking vyos_net_name) skips the slow ice port, which keeps its kernel name (eth0).
- udevadm settle does not cover this — it waits on the udev event queue, not per-NIC driver readiness, so boot proceeds with a mis-named interface.
- The correct port collides and is stranded at a temporary name, absent from eth*.
- hw-id is programmed as the MAC — VyOS sets each interface's MAC to its configured hw-id, comparing only against the current MAC (never the permanent MAC). The copper hw-id is written onto the empty ice port; the stranded copper port keeps its real MAC → duplicate MAC.
- The remedy is blocked, then fatal — pinning the displaced port is rejected by the CLI because the target name is not currently enumerated. Via config.boot + reboot, if the race again leaves the ice port on the low name, the pinned interface has no matching netdev, interface verification raises, and because the boot config loader commits the whole configuration as a single transaction and treats a commit exception as unrecoverable (no partial recovery unless boot debug is enabled), the entire configuration fails to apply → no management access.
Component pointers for triage: src/udev/vyos_net_name, src/etc/udev/rules.d/62-temporary-interface-rename.rules, src/etc/udev/rules.d/65-vyos-net.rules, src/helpers/vyos-interface-rescan.py, interface MAC handling in python/vyos/ifconfig/interface.py (set_mac), python/vyos/configverify.py (verify_interface_exists), and src/helpers/vyos-boot-config-loader.py (single-transaction commit).
Impact
- Configured interfaces silently bound to the wrong physical port; IPs applied to unused/dead ports.
- Duplicate MAC addresses on the host (L2 breakage).
- Non-deterministic interface naming across reboots.
- The natural user remedy (hw-id pin) can render the router completely unreachable on the next reboot, requiring console/IPMI recovery — a serious risk for remotely managed systems.
Workarounds
- Blacklisting the ice driver (the SFP28 ports are unused) removes the slow actor, restores deterministic naming, and silences the log spam. Downside: an /etc/modprobe.d blacklist does not persist across image upgrades on rolling.
- Not a safe workaround: pinning the displaced port via a hand-edited config.boot — this is what caused the management lockout.
Suggested direction (high-level)
- Base persistent naming on an attribute that is stable and independent of driver-init timing — the permanent MAC and/or PCI topology — rather than a runtime rename gated on DRIVERS=="?*".
- Treat hw-id strictly as an identity/match key; do not program it as the interface MAC. When a MAC must be restored, reset to the device's permanent hardware address, and refuse to bind a name to a device whose permanent MAC does not match the configured hw-id (fail safe rather than create a duplicate).
- Make boot-time interface validation non-fatal (warn/skip) and/or make the boot config commit best-effort, so a single interface problem cannot remove all management access.
- Wait for NIC/driver readiness (expected PCI netdevs bound) before finalizing naming, rather than relying solely on udevadm settle.
Driver-side context
Running OOT Intel ice 2.6.7 on kernel 6.18.x, introduced to rolling via T9010 (2026-06-22). Earlier images did not exhibit the reordering, which points to the newer driver's initialization timing (relative to igb/i40e) as the trigger.