Page MenuHomeVyOS Platform

Bug: Slow-initializing NIC (Intel ice / empty E823-L SFP cages) loses VyOS interface-naming race → interface reordering, duplicate MAC via hw-id, and total management lockout on reboot
Open, HighPublicBUG

Description

Severity: High — can cause complete loss of management access, recoverable only from the console/IPMI.
Affected: VyOS 2026.07.21-1151-rolling, kernel 6.18.x, vyos-intel-ice 2.6.7. Believed to affect all rolling images since T9010 (2026-06-22: kernel 6.18 + Intel OOT driver refresh).
Components: interface naming & discovery (udev vyos_net_name, boot-time rescan), interface MAC handling, boot config loader. Trigger: kernel / Intel out-of-tree ice driver.

WARNING: The natural remedy a user reaches for here — pinning the port by hw-id — can render the router completely unreachable on the next reboot. See Impact.

Summary

On a SuperMicro system whose Intel E823-L 25G SFP28 ports (driver ice) are unused with empty cages, those ports initialize slowly (DDP package load plus a firmware "module not present" health event). During boot they lose VyOS's interface-naming race and capture low ethN names that the configuration has pinned (via hw-id) to other, in-use ports — here, copper RJ45 ports on igb.

The consequences cascade:

  1. Unused ports end up named eth0/etc.; the configured copper ports are displaced and left unnamed.
  2. VyOS programs each interface's MAC from its configured hw-id, so the copper port's MAC is written onto the empty ice port — producing two interfaces with the same MAC, and putting the configured IP on a dead, empty cage.
  3. Interface naming becomes non-deterministic across reboots.
  4. Correcting this by pinning the displaced port to a new name fails at the CLI (Interface "ethX" does not exist!); doing it via config.boot + reboot caused a complete loss of management access, because the boot loader applies the whole configuration as a single transaction and treats a commit failure as unrecoverable.

The ice "Module is not present" log lines are expected for empty cages (a firmware health-status event) and are not themselves the defect — but the slow initialization they accompany is what exposes the naming race.

Environment

  • Product: VyOS 2026.07.21-1151-rolling
  • Kernel: 6.18.x
  • Driver: vyos-intel-ice 2.6.7 (out-of-tree). Note: vyos-build pins ice at commit_id = v2.6.6, but the Intel-driver build step resets the source tree to origin/main before building, so the packaged version can drift from the declared pin (the running driver reports 2.6.7). Worth verifying for build reproducibility.
  • NIC firmware (E823-L): fw 5.5.14 api 1.7.9 nvm 2.19 0x8000d221 1.3083.0

System / hardware

  • System: SuperMicro SYS-510D-10C-FN6P
  • Board: SuperMicro X12SDV-10C-SP6F
  • CPU: Intel Xeon D-1747NTE (Ice Lake-D)

NIC inventory (8 ports)

PCIControllerDriverPermanent MACMediaIn use
0000:02:00.0Intel I210 [8086:1533]igb3c:ec:ef:d1:87:9aCopper RJ45yes
0000:03:00.0Intel I210 [8086:1533]igb3c:ec:ef:d1:87:9bCopper RJ45yes
0000:04:00.0Intel I350 [8086:1521]igb3c:ec:ef:d1:87:9cCopper RJ45yes
0000:04:00.1Intel I350 [8086:1521]igb3c:ec:ef:d1:87:9dCopper RJ45yes
0000:15:00.0Intel XXV710 [8086:158b]i40e3c:ec:ef:d8:16:eeSFP28empty
0000:15:00.1Intel XXV710 [8086:158b]i40e3c:ec:ef:d8:16:efSFP28empty
0000:f4:00.0Intel E823-L for SFP [8086:124d] (SoC)ice3c:ec:ef:d1:88:a2SFP28empty
0000:f4:00.2Intel E823-L for SFP [8086:124d] (SoC)ice3c:ec:ef:d1:88:a3SFP28empty

The two ice ports are the Xeon-D SoC-integrated E823-L. The two i40e ports are an add-in/AIOM XXV710. Only the four igb copper RJ45 ports are cabled/used.

Steps to reproduce

  1. VyOS rolling (kernel 6.18 / vyos-intel-ice 2.6.x) on a host with at least one Intel 800-series (ice) port that has no media (empty SFP cage / unconnected), alongside other NIC families (igb/i40e).
  2. Configure interfaces ethernet ethN pinned by hw-id, with addresses on the non-ice ports.
  3. Reboot, ideally several times.

Expected behavior

Each configured interface deterministically binds to the physical port whose permanent MAC equals its hw-id, regardless of driver initialization order; configured addresses land on the intended ports; naming is stable across reboots.

Actual behavior

  • The empty/slow ice port(s) capture low ethN names; configured copper ports are displaced and disappear from eth*.
  • The copper port's hw-id is programmed onto the empty ice port, so the configured IP lands on a dead cage and two interfaces end up with the same MAC.
  • Naming varies across reboots.
  • Pinning the displaced port via the CLI fails: Interface "eth2" does not exist!
  • Pinning via config.boot + reboot results in a complete, unreachable boot.

Evidence

dmesg (trimmed):

ice: Intel(R) Ethernet Connection E800 Series Linux Driver - version 2.6.7
ice 0000:f4:00.0: fw 5.5.14 api 1.7.9 nvm 2.19 0x8000d221 1.3083.0 [8086:124d] [15d9:124d]
ice 0000:f4:00.0: Failed to apply Tx scheduling configuration, err -95
ice 0000:f4:00.0: The DDP package was successfully loaded: ICE OS Default Package version 1.3.59.0
ice 0000:f4:00.0 eth0: Module is not present.
ice 0000:f4:00.2 eth2: Module is not present.
ice 0000:f4:00.2 eth3: renamed from eth2

Observed naming vs. permanent MAC (key evidence) — eth0 is the empty ice port f4:00.0, but its current MAC was overwritten to the copper port's hw-id, while its permanent MAC is unchanged:

eth0 pci=0000:f4:00.0 cur=3c:ec:ef:d1:87:9a perm=3c:ec:ef:d1:88:a2 drv=ice   carrier=0
eth1 pci=0000:03:00.0 cur=3c:ec:ef:d1:87:9b perm=3c:ec:ef:d1:87:9b drv=igb   carrier=1
eth3 pci=0000:f4:00.2 cur=3c:ec:ef:d1:88:a3 perm=3c:ec:ef:d1:88:a3 drv=ice   carrier=0
eth4 pci=0000:04:00.0 cur=3c:ec:ef:d1:87:9c perm=3c:ec:ef:d1:87:9c drv=igb   carrier=1
eth5 pci=0000:04:00.1 cur=3c:ec:ef:d1:87:9d perm=3c:ec:ef:d1:87:9d drv=igb   carrier=1
eth6 pci=0000:15:00.0 cur=3c:ec:ef:d8:16:ee perm=3c:ec:ef:d8:16:ee drv=i40e  carrier=0
eth7 pci=0000:15:00.1 cur=3c:ec:ef:d8:16:ef perm=3c:ec:ef:d8:16:ef drv=i40e  carrier=0

The copper port that config pins to eth0 (permanent MAC ...87:9a, PCI 0000:02:00.0) is absent from eth* — displaced to a temporary name. eth0 (the empty ice cage) shows cur=...87:9a (the copper hw-id) with perm=...88:a2 — the hw-id was programmed onto the wrong port, colliding with the stranded copper port's real MAC.

Configuration: interfaces ethernet pinned by hw-id for eth0, eth1, eth3, eth4, eth5, eth6, eth7 (7 of 8 ports; the second ice port, permanent MAC ...88:a2, was not pinned).

Suspected mechanism

  1. Slow init on empty ice ports — DDP package load plus the firmware health event (ICE_AQC_HEALTH_STATUS_ERR_MOD_NOT_PRESENT → "Module is not present") make these ports driver-ready later than the igb/i40e ports.
  2. The naming rules require the driver link to be ready — VyOS's persistent naming (udev rules 62-temporary-interface-rename.rules and 65-vyos-net.rules, both matching DRIVERS=="?*", invoking vyos_net_name) skips the slow ice port, which keeps its kernel name (eth0).
  3. udevadm settle does not cover this — it waits on the udev event queue, not per-NIC driver readiness, so boot proceeds with a mis-named interface.
  4. The correct port collides and is stranded at a temporary name, absent from eth*.
  5. hw-id is programmed as the MAC — VyOS sets each interface's MAC to its configured hw-id, comparing only against the current MAC (never the permanent MAC). The copper hw-id is written onto the empty ice port; the stranded copper port keeps its real MAC → duplicate MAC.
  6. The remedy is blocked, then fatal — pinning the displaced port is rejected by the CLI because the target name is not currently enumerated. Via config.boot + reboot, if the race again leaves the ice port on the low name, the pinned interface has no matching netdev, interface verification raises, and because the boot config loader commits the whole configuration as a single transaction and treats a commit exception as unrecoverable (no partial recovery unless boot debug is enabled), the entire configuration fails to apply → no management access.

Component pointers for triage: src/udev/vyos_net_name, src/etc/udev/rules.d/62-temporary-interface-rename.rules, src/etc/udev/rules.d/65-vyos-net.rules, src/helpers/vyos-interface-rescan.py, interface MAC handling in python/vyos/ifconfig/interface.py (set_mac), python/vyos/configverify.py (verify_interface_exists), and src/helpers/vyos-boot-config-loader.py (single-transaction commit).

Impact

  • Configured interfaces silently bound to the wrong physical port; IPs applied to unused/dead ports.
  • Duplicate MAC addresses on the host (L2 breakage).
  • Non-deterministic interface naming across reboots.
  • The natural user remedy (hw-id pin) can render the router completely unreachable on the next reboot, requiring console/IPMI recovery — a serious risk for remotely managed systems.

Workarounds

  • Blacklisting the ice driver (the SFP28 ports are unused) removes the slow actor, restores deterministic naming, and silences the log spam. Downside: an /etc/modprobe.d blacklist does not persist across image upgrades on rolling.
  • Not a safe workaround: pinning the displaced port via a hand-edited config.boot — this is what caused the management lockout.

Suggested direction (high-level)

  • Base persistent naming on an attribute that is stable and independent of driver-init timing — the permanent MAC and/or PCI topology — rather than a runtime rename gated on DRIVERS=="?*".
  • Treat hw-id strictly as an identity/match key; do not program it as the interface MAC. When a MAC must be restored, reset to the device's permanent hardware address, and refuse to bind a name to a device whose permanent MAC does not match the configured hw-id (fail safe rather than create a duplicate).
  • Make boot-time interface validation non-fatal (warn/skip) and/or make the boot config commit best-effort, so a single interface problem cannot remove all management access.
  • Wait for NIC/driver readiness (expected PCI netdevs bound) before finalizing naming, rather than relying solely on udevadm settle.

Driver-side context

Running OOT Intel ice 2.6.7 on kernel 6.18.x, introduced to rolling via T9010 (2026-06-22). Earlier images did not exhibit the reordering, which points to the newer driver's initialization timing (relative to igb/i40e) as the trigger.

NOTE: The ice "Module is not present" messages are expected for empty cages (a firmware health-status event decoded by the driver) and are noise here, not the defect — but the slow init that accompanies them is what exposes the VyOS naming race.

Details

Version
2026.07.21-1151-rolling
Is it a breaking change?
Unspecified (possibly destroys the router)
Issue type
Bug (incorrect behavior)

Event Timeline

This is the same boot-time interface-naming race tracked by T3871, here triggered by a slow-initializing NIC (an empty-cage Intel ice E823-L, which delays bring-up with a DDP-package load plus a firmware "module not present" media check) rather than a multi-vendor probe-order difference. Related interface-ordering reports for cross-reference:

  • T3871 — Resolve unexpected interface name reordering (meta-task)
  • T4030 — SR-IOV and interface renaming bug
  • T3314 — Udev rules try to rename active interfaces in some environments
  • T1058 — hw-id is ignored when naming interfaces (closed via udevadm settle; this case shows settle is insufficient for slow drivers — it waits on the udev event queue, not per-NIC driver readiness)
  • T577 — Unconfigured Ethernet interface discovery partial failure on boot
  • T770 — Bonded interfaces get updated with incorrect hw-id in config

This appears to be addressed by open PR https://github.com/vyos/vyos-1x/pull/5350 ("boot: T3871: rework interface renaming and ordering"), which replaces per-device udev naming with a single authoritative post-settle pass — wait for configured hardware, apply hw-id bindings, then name the rest by PCIe distance + MAC — matching the direction proposed here. I have two SuperMicro X12SDV-10C-SP6F units that reliably reproduce this (empty-SFP ice ports racing igb copper, alternating which interface is displaced each boot); I am testing #5350 against them and will post results on that PR.

One thing #5350 does not touch: python/vyos/ifconfig/interface.py still programs the configured hw-id as the interface MAC when there is no mac override. With #5350 fixing naming this no longer stamps a foreign MAC, but it remains latent — restoring the port's permanent hardware address and treating hw-id strictly as an identity key would close it. Happy to raise separately.

If you need to temporarily resolve this issue, you can try increasing the number of all ethX (for example, starting the interface naming from eth100). I have a server with 8x E823-L ports and named them from eth101-eth108.

@lclements0 Retest please with the latest rollin release
Thanks.

@Viacheslav I can confirm the boot race issue is gone on the latest rolling release. Thanks!