CVE-2021-46925: net/smc: fix kernel panic caused by race of smc_sock
In the Linux kernel, the following vulnerability has been resolved:
net/smc: fix kernel panic caused by race of smc_sock
A crash occurs when smc_cdc_tx_handler() tries to access smc_sock
but smc_release() has already freed it.
[ 4570.695099] BUG: unable to handle page fault for address: 000000002eae9e88
[ 4570.696048] #PF: supervisor write access in kernel mode
[ 4570.696728] #PF: error_code(0x0002) - not-present page
[ 4570.697401] PGD 0 P4D 0
[ 4570.697716] Oops: 0002 [#1] PREEMPT SMP NOPTI
[ 4570.698228] CPU: 0 PID: 0 Comm: swapper/0 Not tainted 5.16.0-rc4+ #111
[ 4570.699013] Hardware name: Alibaba Cloud Alibaba Cloud ECS, BIOS 8c24b4c 04/0
[ 4570.699933] RIP: 0010:_raw_spin_lock+0x1a/0x30
<...>
[ 4570.711446] Call Trace:
[ 4570.711746] <IRQ>
[ 4570.711992] smc_cdc_tx_handler+0x41/0xc0
[ 4570.712470] smc_wr_tx_tasklet_fn+0x213/0x560
[ 4570.712981] ? smc_cdc_tx_dismisser+0x10/0x10
[ 4570.713489] tasklet_action_common.isra.17+0x66/0x140
[ 4570.714083] __do_softirq+0x123/0x2f4
[ 4570.714521] irq_exit_rcu+0xc4/0xf0
[ 4570.714934] common_interrupt+0xba/0xe0
Though smc_cdc_tx_handler() checked the existence of smc connection,
smc_release() may have already dismissed and released the smc socket
before smc_cdc_tx_handler() further visits it.
smc_cdc_tx_handler() |smc_release()
if (!conn) |
|
|smc_cdc_tx_dismiss_slots()
| smc_cdc_tx_dismisser()
|
|sock_put(&smc->sk) <- last sock_put,
| smc_sock freed
bh_lock_sock(&smc->sk) (panic) |
To make sure we won't receive any CDC messages after we free the
smc_sock, add a refcount on the smc_connection for inflight CDC
message(posted to the QP but haven't received related CQE), and
don't release the smc_connection until all the inflight CDC messages
haven been done, for both success or failed ones.
Using refcount on CDC messages brings another problem: when the link
is going to be destroyed, smcr_link_clear() will reset the QP, which
then remove all the pending CQEs related to the QP in the CQ. To make
sure all the CQEs will always come back so the refcount on the
smc_connection can always reach 0, smc_ib_modify_qp_reset() was replaced
by smc_ib_modify_qp_error().
And remove the timeout in smc_wr_tx_wait_no_pending_sends() since we
need to wait for all pending WQEs done, or we may encounter use-after-
free when handling CQEs.
For IB device removal routine, we need to wait for all the QPs on that
device been destroyed before we can destroy CQs on the device, or
the refcount on smc_connection won't reach 0 and smc_sock cannot be
released.
Security readout for executives and security teams
Plain-English summary
This Linux kernel issue can crash a system when a race condition frees an SMC socket while another kernel path still uses it. The business impact is availability, not data theft. Exploitation requires local access, low privileges, and high complexity, so prioritize exposed multi-user Linux systems over isolated hosts.
Executive priority
Treat this as a scheduled availability-risk patch, not an emergency breach response. Give higher priority to shared Linux servers, container hosts, and environments with untrusted local users. No active exploitation is cited in the provided sources.
Technical view
CVE-2021-46925 is a CWE-362 race in Linux SMC handling. smc_cdc_tx_handler() can touch smc_sock after smc_release() has freed it, causing a kernel page fault and panic. The fix adds refcounting for inflight CDC messages and changes QP teardown behavior so pending completion handling cannot reference freed state.
Likely exposure
Exposure is limited to Linux kernels in the affected set, particularly systems using or enabling SMC networking. The CVSS vector is local, high-complexity, low-privilege, no user interaction, with high availability impact. Distribution backports may change practical exposure, so version checks alone may be insufficient.
Exploitation context
The source bundle does not show KEV listing or active exploitation. The issue is locally reachable and complex, with impact described as kernel panic. There is no source evidence of remote exploitation, public weaponization, data exposure, or privilege escalation.
Researcher notes
The bug centers on lifetime management between smc_cdc_tx_handler() and smc_release(). The fix uses smc_connection refcounts for inflight CDC messages and waits for pending WQEs/CQEs. Validate backports by code or changelog, because distribution kernel versions may not align with upstream stable versions.
Mitigation direction
Apply Linux kernel updates that include the referenced stable fixes.
Use distribution vendor kernels with confirmed backported fixes.
Check vendor guidance where upstream version mapping is unclear.
Prioritize multi-user Linux systems and hosts with SMC in use.
Validation and detection
Inventory running Linux kernel versions across affected systems.
Confirm whether the referenced stable commits are present or backported.
Check whether SMC support is enabled or operationally required.
Review crash logs for kernel panics in smc_cdc_tx_handler paths.
Generated from the cited source records. This long-tail analysis has not been individually reviewed by a named human.
Potential ATT&CK relevance
Conservative CVE-to-ATT&CK context
These mappings and lookup hints may be relevant to the vulnerability behavior, CWE, affected product, or exposure path. Glexia-inferred context is not an official MITRE, ATT&CK, CWE, or CVE Program mapping.
ATT&CK lookup starting points
Use these exact CWE pages and searches to review the Glexia ATT&CK library from this CVE's weakness and description context.
cwe · low confidence lookup
CWE-362: Exact CWE lookup
Use the exact CWE identifier as the starting point before reviewing related ATT&CK behavior. Open the exact CWE lookup page first, then review the ATT&CK searches from that MITRE weakness context. This is a Glexia lookup hint, not an official ATT&CK mapping.
These fields come from the CVE record and ADP containers, not from Glexia's Take. They preserve time-varying source decisions such as CISA SSVC, KEV status, CVSS metrics, and provider references.
We collect every scored CVSS vector available in the official CNA and ADP containers. When more than one version is present, the table keeps the source vectors side by side instead of collapsing them into the highest score.
CWE links open Glexia weakness intelligence pages with official CWE context, developer remediation guidance, and related CVE mappings.
CWE-362 · source CWE mapping
Concurrent Execution using Shared Resource with Improper Synchronization ('Race Condition')
Concurrent Execution using Shared Resource with Improper Synchronization ('Race Condition') represents a recurring weakness pattern that can create exploitable paths when design, validation, or implementation controls are missing.