CVE-2024-53079: mm/thp: fix deferred split unqueue naming and locking
In the Linux kernel, the following vulnerability has been resolved:
mm/thp: fix deferred split unqueue naming and locking
Recent changes are putting more pressure on THP deferred split queues:
under load revealing long-standing races, causing list_del corruptions,
"Bad page state"s and worse (I keep BUGs in both of those, so usually
don't get to see how badly they end up without). The relevant recent
changes being 6.8's mTHP, 6.10's mTHP swapout, and 6.12's mTHP swapin,
improved swap allocation, and underused THP splitting.
Before fixing locking: rename misleading folio_undo_large_rmappable(),
which does not undo large_rmappable, to folio_unqueue_deferred_split(),
which is what it does. But that and its out-of-line __callee are mm
internals of very limited usability: add comment and WARN_ON_ONCEs to
check usage; and return a bool to say if a deferred split was unqueued,
which can then be used in WARN_ON_ONCEs around safety checks (sparing
callers the arcane conditionals in __folio_unqueue_deferred_split()).
Just omit the folio_unqueue_deferred_split() from free_unref_folios(), all
of whose callers now call it beforehand (and if any forget then bad_page()
will tell) - except for its caller put_pages_list(), which itself no
longer has any callers (and will be deleted separately).
Swapout: mem_cgroup_swapout() has been resetting folio->memcg_data 0
without checking and unqueueing a THP folio from deferred split list;
which is unfortunate, since the split_queue_lock depends on the memcg
(when memcg is enabled); so swapout has been unqueueing such THPs later,
when freeing the folio, using the pgdat's lock instead: potentially
corrupting the memcg's list. __remove_mapping() has frozen refcount to 0
here, so no problem with calling folio_unqueue_deferred_split() before
resetting memcg_data.
That goes back to 5.4 commit 87eaceb3faa5 ("mm: thp: make deferred split
shrinker memcg aware"): which included a check on swapcache before adding
to deferred queue, but no check on deferred queue before adding THP to
swapcache. That worked fine with the usual sequence of events in reclaim
(though there were a couple of rare ways in which a THP on deferred queue
could have been swapped out), but 6.12 commit dafff3f4c850 ("mm: split
underused THPs") avoids splitting underused THPs in reclaim, which makes
swapcache THPs on deferred queue commonplace.
Keep the check on swapcache before adding to deferred queue? Yes: it is
no longer essential, but preserves the existing behaviour, and is likely
to be a worthwhile optimization (vmstat showed much more traffic on the
queue under swapping load if the check was removed); update its comment.
Memcg-v1 move (deprecated): mem_cgroup_move_account() has been changing
folio->memcg_data without checking and unqueueing a THP folio from the
deferred list, sometimes corrupting "from" memcg's list, like swapout.
Refcount is non-zero here, so folio_unqueue_deferred_split() can only be
used in a WARN_ON_ONCE to validate the fix, which must be done earlier:
mem_cgroup_move_charge_pte_range() first try to split the THP (splitting
of course unqueues), or skip it if that fails. Not ideal, but moving
charge has been requested, and khugepaged should repair the THP later:
nobody wants new custom unqueueing code just for this deprecated case.
The 87eaceb3faa5 commit did have the code to move from one deferred list
to another (but was not conscious of its unsafety while refcount non-0);
but that was removed by 5.6 commit fac0516b5534 ("mm: thp: don't need care
deferred split queue in memcg charge move path"), which argued that the
existence of a PMD mapping guarantees that the THP cannot be on a deferred
list. As above, false in rare cases, and now commonly false.
Backport to 6.11 should be straightforward. Earlier backports must take
care that other _deferred_list fixes and dependencies are included. There
is not a strong case for backports, but they can fix cornercases.
Security readout for executives and security teams
Plain-English summary
A Linux kernel locking race can corrupt memory-management lists when transparent huge pages are split, swapped, or moved between memory-control groups. Under load, this may trigger kernel errors, crashes, or potentially broader compromise. Exploitation requires local, low-privileged access; no user interaction is required.
Executive priority
Prioritize remediation on shared Linux hosts where untrusted or low-privileged users can execute code. Treat exposed production and container platforms as urgent, while validating vendor backports before scheduling upgrades. The high impact score warrants timely action, but the sources do not establish active exploitation.
Technical view
CVE-2024-53079 is a CWE-667 improper-locking flaw in transparent huge page deferred-split queues. Swapout or deprecated memcg-v1 charge movement can change folio memory-cgroup data before safely removing the folio from its original queue, causing the wrong lock to be used and potentially corrupting kernel lists. CVSS 3.1 is 7.8.
Likely exposure
Systems using affected Linux kernels are most relevant, particularly those exercising transparent huge pages, swapping, or memory cgroups under load. The supplied metadata identifies affected versions including 5.4, 6.6.62, 6.11.8, and 6.12, but does not clearly express complete version ranges. Confirm exposure against distribution advisories and stable-kernel commits.
Exploitation context
The CVSS vector describes a local attack requiring low privileges, with low complexity and no user interaction. The source reports observed list corruption, bad-page states, and kernel BUG conditions. KEV is false, and the supplied sources provide no evidence of active exploitation or a public exploit.
Researcher notes
The race originates from deferred-split queue membership becoming inconsistent with folio memory-cgroup ownership. Swapout could clear memcg data before unqueueing, selecting a pgdat lock instead of the original memcg lock. The fix unqueues earlier and adds safety checks; memcg-v1 movement instead attempts THP splitting or skips the move. Earlier backports may require dependencies.
Mitigation direction
Apply the appropriate vendor or stable-kernel update containing the referenced fix.
Confirm distribution backports rather than relying only on the displayed kernel version.
Prioritize multi-user, container-hosting, or otherwise locally accessible systems.
Review vendor guidance before considering configuration workarounds; none are specified in the sources.
Validation and detection
Inventory kernel versions and compare them with distribution-specific CVE advisories and backport records.
Verify the installed kernel includes the applicable referenced stable commit.
Identify systems using transparent huge pages, swap, and memory cgroups.
Review kernel logs for list corruption, bad-page-state reports, or related BUG events.
Generated from the cited source records. This long-tail analysis has not been individually reviewed by a named human.
Potential ATT&CK relevance
Conservative CVE-to-ATT&CK context
These mappings and lookup hints may be relevant to the vulnerability behavior, CWE, affected product, or exposure path. Glexia-inferred context is not an official MITRE, ATT&CK, CWE, or CVE Program mapping.
ATT&CK lookup starting points
Use these exact CWE pages and searches to review the Glexia ATT&CK library from this CVE's weakness and description context.
cwe · low confidence lookup
CWE-667: Exact CWE lookup
Use the exact CWE identifier as the starting point before reviewing related ATT&CK behavior. Open the exact CWE lookup page first, then review the ATT&CK searches from that MITRE weakness context. This is a Glexia lookup hint, not an official ATT&CK mapping.
These fields come from the CVE record and ADP containers, not from Glexia's Take. They preserve time-varying source decisions such as CISA SSVC, KEV status, CVSS metrics, and provider references.
We collect every scored CVSS vector available in the official CNA and ADP containers. When more than one version is present, the table keeps the source vectors side by side instead of collapsing them into the highest score.
CWE links open Glexia weakness intelligence pages with official CWE context, developer remediation guidance, and related CVE mappings.
CWE-667 · source CWE mapping
Improper Locking
Improper Locking represents a recurring weakness pattern that can create exploitable paths when design, validation, or implementation controls are missing.