CVE-2025-38472: netfilter: nf_conntrack: fix crash due to removal of uninitialised entry
In the Linux kernel, the following vulnerability has been resolved:
netfilter: nf_conntrack: fix crash due to removal of uninitialised entry
A crash in conntrack was reported while trying to unlink the conntrack
entry from the hash bucket list:
[exception RIP: __nf_ct_delete_from_lists+172]
[..]
#7 [ff539b5a2b043aa0] nf_ct_delete at ffffffffc124d421 [nf_conntrack]
#8 [ff539b5a2b043ad0] nf_ct_gc_expired at ffffffffc124d999 [nf_conntrack]
#9 [ff539b5a2b043ae0] __nf_conntrack_find_get at ffffffffc124efbc [nf_conntrack]
[..]
The nf_conn struct is marked as allocated from slab but appears to be in
a partially initialised state:
ct hlist pointer is garbage; looks like the ct hash value
(hence crash).
ct->status is equal to IPS_CONFIRMED|IPS_DYING, which is expected
ct->timeout is 30000 (=30s), which is unexpected.
Everything else looks like normal udp conntrack entry. If we ignore
ct->status and pretend its 0, the entry matches those that are newly
allocated but not yet inserted into the hash:
- ct hlist pointers are overloaded and store/cache the raw tuple hash
- ct->timeout matches the relative time expected for a new udp flow
rather than the absolute 'jiffies' value.
If it were not for the presence of IPS_CONFIRMED,
__nf_conntrack_find_get() would have skipped the entry.
Theory is that we did hit following race:
cpu x cpu y cpu z
found entry E found entry E
E is expired <preemption>
nf_ct_delete()
return E to rcu slab
init_conntrack
E is re-inited,
ct->status set to 0
reply tuplehash hnnode.pprev
stores hash value.
cpu y found E right before it was deleted on cpu x.
E is now re-inited on cpu z. cpu y was preempted before
checking for expiry and/or confirm bit.
->refcnt set to 1
E now owned by skb
->timeout set to 30000
If cpu y were to resume now, it would observe E as
expired but would skip E due to missing CONFIRMED bit.
nf_conntrack_confirm gets called
sets: ct->status |= CONFIRMED
This is wrong: E is not yet added
to hashtable.
cpu y resumes, it observes E as expired but CONFIRMED:
<resumes>
nf_ct_expired()
-> yes (ct->timeout is 30s)
confirmed bit set.
cpu y will try to delete E from the hashtable:
nf_ct_delete() -> set DYING bit
__nf_ct_delete_from_lists
Even this scenario doesn't guarantee a crash:
cpu z still holds the table bucket lock(s) so y blocks:
wait for spinlock held by z
CONFIRMED is set but there is no
guarantee ct will be added to hash:
"chaintoolong" or "clash resolution"
logic both skip the insert step.
reply hnnode.pprev still stores the
hash value.
unlocks spinlock
return NF_DROP
<unblocks, then
crashes on hlist_nulls_del_rcu pprev>
In case CPU z does insert the entry into the hashtable, cpu y will unlink
E again right away but no crash occurs.
Without 'cpu y' race, 'garbage' hlist is of no consequence:
ct refcnt remains at 1, eventually skb will be free'd and E gets
destroyed via: nf_conntrack_put -> nf_conntrack_destroy -> nf_ct_destroy.
To resolve this, move the IPS_CONFIRMED assignment after the table
insertion but before the unlock.
Pablo points out that the confirm-bit-store could be reordered to happen
before hlist add resp. the timeout fixup, so switch to set_bit and
before_atomic memory barrier to prevent this.
It doesn't matter if other CPUs can observe a newly inserted entry right
before the CONFIRMED bit was set:
Such event cannot be distinguished from above "E is the old incarnation"
case: the entry will be skipped.
Also change nf_ct_should_gc() to first check the confirmed bit.
The gc sequence is:
1. Check if entry has expired, if not skip to next entry
2. Obtain a reference to the expired entry.
3. Call nf_ct_should_gc() to double-check step 1.
nf_ct_should_gc() is thus called only for entries that already failed an
expiry check. After this patch, once the confirmed bit check pas
---truncated---
Security readout for executives and security teams
Plain-English summary
A race condition in Linux connection tracking can make the kernel remove an incompletely initialized network-tracking entry and crash. Systems using affected kernels with netfilter connection tracking may suffer outages. The supplied evidence does not establish data theft, controlled code execution, or a reliable remote attack path.
Executive priority
Treat this as an urgent availability risk for conntrack-enabled network infrastructure, but not as confirmed compromise activity. Inventory relevant kernels immediately and schedule vendor updates promptly. The critical rating supports rapid action, although the supplied technical evidence documents a kernel crash and does not substantiate the CVSS confidentiality and integrity impacts.
Technical view
A narrow multi-CPU race can leave one CPU referencing an nf_conn object while it is deleted, recycled, and reinitialized elsewhere. The CONFIRMED flag may become visible before hashtable insertion. Garbage collection can then delete an invalid list node, crashing in __nf_ct_delete_from_lists.
Likely exposure
Exposure requires an affected Linux kernel with nf_conntrack active. Firewalls, NAT gateways, and other connection-tracking hosts are plausible priorities. The supplied version data identifies affected stable branches but appears condensed, so vulnerability managers should match exact distribution kernel builds against vendor advisories and referenced fixes.
Exploitation context
The bundle reports a crash and assigns CVSS 3.1 score 9.8, but provides no evidence of active exploitation. The CVE is not marked KEV. Trigger reliability, attacker control, and remote reachability are not demonstrated, so the documented practical impact is availability loss rather than confirmed confidentiality or integrity compromise.
Researcher notes
The fix moves IPS_CONFIRMED assignment after table insertion, enforces memory ordering, and makes garbage collection check confirmation first. The race depends on deletion, slab reuse, reinitialization, confirmation, failed insertion, and another CPU resuming. Sources explain crash mechanics but do not demonstrate controlled memory corruption or exploitation.
Mitigation direction
Update to a vendor kernel containing the referenced nf_conntrack fix, then reboot into the updated kernel.
Prioritize conntrack-enabled network infrastructure and other high-traffic Linux hosts for accelerated maintenance.
If patching is delayed, consult the Linux or distribution vendor for supported mitigations; this bundle names none.
Validation and detection
Record the running kernel build and determine whether nf_conntrack is built in, loaded, or actively used.
Match the exact kernel package against distribution advisories and the referenced stable-kernel fix commits.
After updating, verify the running kernel changed; an installed but unrebooted kernel remains ineffective.
Review kernel logs for __nf_ct_delete_from_lists, nf_ct_delete, nf_ct_gc_expired, or unexplained network-path crashes.
Generated from the cited source records. This long-tail analysis has not been individually reviewed by a named human.
Potential ATT&CK relevance
Conservative CVE-to-ATT&CK context
These mappings and lookup hints may be relevant to the vulnerability behavior, CWE, affected product, or exposure path. Glexia-inferred context is not an official MITRE, ATT&CK, CWE, or CVE Program mapping.
ATT&CK lookup starting points
Use these exact CWE pages and searches to review the Glexia ATT&CK library from this CVE's weakness and description context.
cve · low confidence lookup
CVE-2025-38472 mapping review
Open the CVE-to-ATT&CK bridge for reviewed, inferred, or future official mappings tied to this CVE.
These fields come from the CVE record and ADP containers, not from Glexia's Take. They preserve time-varying source decisions such as CISA SSVC, KEV status, CVSS metrics, and provider references.
1CVSS vectors
3Timeline events
1ADP providers
7Source links
CVSS vector scores
1 official score
We collect every scored CVSS vector available in the official CNA and ADP containers. When more than one version is present, the table keeps the source vectors side by side instead of collapsing them into the highest score.