Page MenuHomeVyOS Platform

QAT driver autoloads and self-starts independently of "set system acceleration qat"; op mode reports "not configured" while the device is active
In progress, NormalPublicBUG

Description

VyOS documents Intel QAT acceleration as an opt-in enabled with set system acceleration qat. On hardware where the QAT modules autoload, that node gates nothing. The driver binds, the device starts, and — because the Intel driver is built with --enable-qat-lkcf — it registers with the kernel crypto framework, so in-kernel users including IPsec ESP can be offloaded to it. None of that consults the configuration.

Op mode then reported the opposite of reality, because show system acceleration qat status short-circuited on the config tree before looking at the device:

def check_qat_if_conf():
    if not Config().exists_effective('system acceleration qat'):
        print("\t system acceleration qat is not configured")
        sys.exit(1)

An operator checking whether QAT was in the datapath was told it was not, on a machine where it was.

Observed on the affected system

Supermicro X12SDV, Xeon D-1700, VyOS rolling, DMVPN/IPsec hub. set system acceleration qat had never been configured.

vyos@hub:~$ show system acceleration qat status
         system acceleration qat is not configured

vyos@hub:~$ lsmod | grep -i qat
qat_200xx              20480  1
intel_qat             413696  17 qat_200xx
uio                    28672  1 intel_qat

vyos@hub:~$ lspci -nnk | grep -i -A3 'QuickAssist\|Co-processor'
01:00.0 Co-processor [0b40]: Intel Corporation 200xx Series QAT [8086:18ee] (rev 11)
        Subsystem: Intel Corporation 200xx Series QAT [8086:0000]
        Kernel driver in use: 200xx
        Kernel modules: qat_200xx

vyos@hub:~$ sudo dmesg | grep -i qat
[   15.734217] 200xx 0000:01:00.0: qat_dev0 started 6 acceleration engines

The uio dependency on intel_qat identifies this as the out-of-tree Intel driver; the in-tree QAT driver has no UIO component.

A second, related defect

While fixing the above I found that op mode and conf mode each carried their own copy of the supported QAT PCI ID list, and the two had been out of sync since 2020.

rP4a83a9ac55ca ("qat: T2968: adjust to C200xx PCI ID from Intel drivers") added 8086:18ee (QAT_200XX) to verify() in src/conf_mode/system_acceleration.py but not to detect_qat_dev() in src/op_mode/show_acceleration.py. rPffba3fc129dc ("qat: T7662: add PCI ID range for Intel C62x virtual function devices") later updated both, which left op mode with every ID except 18ee.

On the box above that means set system acceleration qat commits successfully, because conf mode recognises the device, while show system acceleration qat reports "No QAT device found", because op mode does not. Operators are told the hardware is absent on a machine where the driver is loaded and the device has started.

The attached fix moves the ID table into a single module that both read, so the two lists can no longer drift apart.

Observed vs inferred

Please keep these apart when acting on this report.

Observed on the affected box: modules loaded; device bound to driver 200xx; firmware started ("started 6 acceleration engines"); config node unset; op mode reporting "not configured".

Inferred, not observed: that QAT was registered with the kernel crypto framework on this box, and therefore that ESP was actually offloaded to it. /proc/crypto was not captured before the machine was recovered. The inference rests on --enable-qat-lkcf being set in packages/linux-kernel/build-intel-qat.sh in vyos-build.

WARNING: If QAT was not in fact registered with the crypto framework, the exposure described below does not apply to this particular incident. The reporting defect stands regardless.

That gap is itself part of the bug: there was no CLI command that would have answered the question.

Why this matters

T6177 is an AA deadlock on xfrm_state->lock that is reachable only when ESP is offloaded to the QAT driver:

  1. Inbound ESP arrives via NAPI (softirq). xfrm_input() takes a plain spin_lock(&x->lock) — deliberately not _bh, because the RX path is already in softirq context.
  2. With QAT-LKCF, crypto_aead_decrypt() returns -EINPROGRESS. The completion is delivered by the out-of-tree driver's adf_response_handler_wq on adf_pf_resp_wq_N, an ordinary highpri workqueue — process context, softirqs enabled.
  3. That path re-enters xfrm_input() at resume: and takes spin_lock(&x->lock) again, with BH on.
  4. A NIC IRQ lands on the same CPU. NET_RX re-enters xfrm_input() for the same SA and spins on a lock held by the task it just preempted.

The spinning context is __do_softirq(), so that CPU's softirq accounting freezes, RCU never observes a quiescent state, grace periods stall globally, and any NIC queue vectored to that CPU stops draining.

Upstream has ruled that this is a driver bug, not an xfrm bug. A spin_lock_bh() patch to xfrm_input() was submitted in December 2023 (Zhang Yiqun). Eric Dumazet asked for a Fixes tag and a stack trace; Steffen Klassert, the xfrm maintainer, rejected it:

This looks more like a "crypto driver" bug. xfrm_input() runs in the RX path and therefore expects to run with BHs off.

An earlier 2009 attempt (Yury Polyanskiy) was rejected by Herbert Xu on other grounds. The calling convention stands; the fix belongs in the driver's completion path.

The relevant point for this task is not the deadlock mechanism — it is that on such hardware an operator is exposed to it without having opted in, and had no way to discover that from the CLI.

Historical note: --enable-qat-lkcf was added as the fix for T2853 (2020, "Intel QAT acceleration does not work"). That fix is what created this exposure.

Decision required

Three options, in increasing order of correctness and invasiveness.

(a) Report reality in op mode. Cheap and honest; does not reduce exposure. Implemented — see the linked PR. show system acceleration qat status now reads sysfs, debugfs and /proc/crypto and distinguishes: no device; device present with no driver bound; driver bound and started with the node unset; node set. When QAT is registered with the crypto framework and the node is unset, it says so explicitly.

(b) Warn at boot or on commit when QAT hardware is bound but the node is unset. Makes the condition visible without being asked. Needs a ruling on where the warning belongs — a commit-time warning fires only on commit, a boot-time one needs an activation script — and on how to avoid nagging operators who deliberately rely on autoloaded QAT.

(c) Blacklist the QAT modules by default and have the conf-mode script load them when the node is set. This makes the documented opt-in actually true and is the only option that removes the exposure for operators who have not asked for QAT. It spans vyos-1x and vyos-build packaging.

IMPORTANT: I have deliberately not implemented (b) or (c). Changing default module loading for a hardware class is a maintainer decision: get it wrong and acceleration silently stops working for people who deliberately want it, including anyone relying on the current autoload behaviour. (c) also interacts with the qat_service init script and with whether the device config in /etc/200xx_dev0.conf should still be shipped.

Related

  • T6177 — the deadlock this exposes operators to. Fix belongs in the OOT driver's completion path (local_bh_disable() around adf_handle_response()); tracked separately against vyos-build.
  • T2853 — added --enable-qat-lkcf, the origin of the LKCF exposure.
  • T2968 — added the QAT_200XX PCI ID to conf mode only.
  • T7662 — updated both PCI ID lists but did not notice 18ee was missing from one.

Checking whether a given system is affected

This can be confirmed on any box with QAT hardware, without applying the fix:

lsmod | grep -i qat
grep -iB1 -A8 qat /proc/crypto | head -60

Modules loaded while system acceleration qat is absent from the running config puts the system in the state described above. QAT driver names additionally appearing in /proc/crypto mean the device is registered with the kernel crypto framework and can take ESP traffic.

With the linked PR applied, show system acceleration qat status reports both of these directly — which is the point of the change.

Details

Version
2026.08.14-rolling & others
Is it a breaking change?
Behavior change
Issue type
Bug (incorrect behavior)

Event Timeline

Viacheslav changed the task status from Open to In progress.Aug 21 2026, 10:27 AM
Viacheslav assigned this task to lclements0.
Viacheslav triaged this task as Normal priority.