CVE-2025-38242: mm: userfaultfd: fix race of userfaultfd_move and swap cache
In the Linux kernel, the following vulnerability has been resolved:
mm: userfaultfd: fix race of userfaultfd_move and swap cache
This commit fixes two kinds of races, they may have different results:
Barry reported a BUG_ON in commit c50f8e6053b0, we may see the same
BUG_ON if the filemap lookup returned NULL and folio is added to swap
cache after that.
If another kind of race is triggered (folio changed after lookup) we
may see RSS counter is corrupted:
[ 406.893936] BUG: Bad rss-counter state mm:ffff0000c5a9ddc0
type:MM_ANONPAGES val:-1
[ 406.894071] BUG: Bad rss-counter state mm:ffff0000c5a9ddc0
type:MM_SHMEMPAGES val:1
Because the folio is being accounted to the wrong VMA.
I'm not sure if there will be any data corruption though, seems no.
The issues above are critical already.
On seeing a swap entry PTE, userfaultfd_move does a lockless swap cache
lookup, and tries to move the found folio to the faulting vma. Currently,
it relies on checking the PTE value to ensure that the moved folio still
belongs to the src swap entry and that no new folio has been added to the
swap cache, which turns out to be unreliable.
While working and reviewing the swap table series with Barry, following
existing races are observed and reproduced [1]:
In the example below, move_pages_pte is moving src_pte to dst_pte, where
src_pte is a swap entry PTE holding swap entry S1, and S1 is not in the
swap cache:
CPU1 CPU2
userfaultfd_move
move_pages_pte()
entry = pte_to_swp_entry(orig_src_pte);
// Here it got entry = S1
... < interrupted> ...
<swapin src_pte, alloc and use folio A>
// folio A is a new allocated folio
// and get installed into src_pte
<frees swap entry S1>
// src_pte now points to folio A, S1
// has swap count == 0, it can be freed
// by folio_swap_swap or swap
// allocator's reclaim.
<try to swap out another folio B>
// folio B is a folio in another VMA.
<put folio B to swap cache using S1 >
// S1 is freed, folio B can use it
// for swap out with no problem.
...
folio = filemap_get_folio(S1)
// Got folio B here !!!
... < interrupted again> ...
<swapin folio B and free S1>
// Now S1 is free to be used again.
<swapout src_pte & folio A using S1>
// Now src_pte is a swap entry PTE
// holding S1 again.
folio_trylock(folio)
move_swap_pte
double_pt_lock
is_pte_pages_stable
// Check passed because src_pte == S1
folio_move_anon_rmap(...)
// Moved invalid folio B here !!!
The race window is very short and requires multiple collisions of multiple
rare events, so it's very unlikely to happen, but with a deliberately
constructed reproducer and increased time window, it can be reproduced
easily.
This can be fixed by checking if the folio returned by filemap is the
valid swap cache folio after acquiring the folio lock.
Another similar race is possible: filemap_get_folio may return NULL, but
folio (A) could be swapped in and then swapped out again using the same
swap entry after the lookup. In such a case, folio (A) may remain in the
swap cache, so it must be moved too:
CPU1 CPU2
userfaultfd_move
move_pages_pte()
entry = pte_to_swp_entry(orig_src_pte);
// Here it got entry = S1, and S1 is not in swap cache
folio = filemap_get
---truncated---
Security readout for executives and security teams
Plain-English summary
A race in Linux memory management can associate a swapped memory page with the wrong memory region. Confirmed outcomes include kernel BUG conditions and corrupted memory-accounting counters, potentially destabilizing a host. Although the CVSS assessment includes severe confidentiality, integrity, and availability impact, the supplied kernel description does not confirm actual data corruption.
Executive priority
Treat as an expedited kernel-maintenance issue, especially on shared compute, hosting, CI, or other systems executing untrusted local code. Immediate emergency response is not supported by the supplied exploitation evidence, but affected high-value or multi-tenant hosts should not wait for a distant routine patch cycle.
Technical view
During userfaultfd_move, a lockless swap-cache lookup can become stale while a swap entry is freed and reused. Subsequent PTE checks may accept an unrelated folio or miss one added after lookup, causing incorrect reverse mapping or accounting. The referenced fix validates the swap-cache folio after locking and handles the lookup-miss race under page-table locks.
Likely exposure
Exposure is limited to Linux systems running an affected kernel where a local, low-privileged process can exercise userfaultfd_move against swapped memory. The supplied affected-version data is ambiguous, so exact exposure must be established from kernel provenance, vendor advisories, and the referenced stable fixes rather than version numbers alone.
Exploitation context
The source describes a deliberately reproducible race requiring collisions among several uncommon events and an otherwise very short timing window. CVSS rates it local, low-privilege, low-complexity, without user interaction. It is not listed as KEV in the supplied bundle, and no cited source establishes active exploitation or a public weaponized exploit.
Researcher notes
Confirmed symptoms are BUG_ON conditions and invalid RSS accounting caused by moving or accounting the wrong folio. The original description says data corruption was uncertain. Version metadata in the supplied bundle is insufficiently precise for reliable range construction; validate commit ancestry or distribution backports. Avoid equating the CVSS impact assumptions with demonstrated exploitation outcomes.
Mitigation direction
Apply a vendor-supported kernel update containing the applicable referenced stable fix.
Compare custom kernel source against the three cited stable commits.
Prioritize multi-user hosts and systems running untrusted local workloads.
Restrict local access or affected workloads until patching where operationally appropriate.
Validation and detection
Record the running kernel version, distribution build, and source provenance.
Check vendor guidance for applicability to the exact kernel build.
Confirm the applicable fix commit is present in source or vendor changelogs.
After updating, verify the patched kernel is currently running.
Assess whether local workloads can use userfaultfd with swap-backed memory.
Generated from the cited source records. This long-tail analysis has not been individually reviewed by a named human.
Potential ATT&CK relevance
Conservative CVE-to-ATT&CK context
These mappings and lookup hints may be relevant to the vulnerability behavior, CWE, affected product, or exposure path. Glexia-inferred context is not an official MITRE, ATT&CK, CWE, or CVE Program mapping.
ATT&CK lookup starting points
Use these exact CWE pages and searches to review the Glexia ATT&CK library from this CVE's weakness and description context.
cve · low confidence lookup
CVE-2025-38242 mapping review
Open the CVE-to-ATT&CK bridge for reviewed, inferred, or future official mappings tied to this CVE.
These fields come from the CVE record and ADP containers, not from Glexia's Take. They preserve time-varying source decisions such as CISA SSVC, KEV status, CVSS metrics, and provider references.
1CVSS vectors
3Timeline events
0ADP providers
4Source links
CVSS vector scores
1 official score
We collect every scored CVSS vector available in the official CNA and ADP containers. When more than one version is present, the table keeps the source vectors side by side instead of collapsing them into the highest score.