CVE-2024-26759: mm/swap: fix race when skipping swapcache
In the Linux kernel, the following vulnerability has been resolved:
mm/swap: fix race when skipping swapcache
When skipping swapcache for SWP_SYNCHRONOUS_IO, if two or more threads
swapin the same entry at the same time, they get different pages (A, B).
Before one thread (T0) finishes the swapin and installs page (A) to the
PTE, another thread (T1) could finish swapin of page (B), swap_free the
entry, then swap out the possibly modified page reusing the same entry.
It breaks the pte_same check in (T0) because PTE value is unchanged,
causing ABA problem. Thread (T0) will install a stalled page (A) into the
PTE and cause data corruption.
One possible callstack is like this:
CPU0 CPU1
---- ----
do_swap_page() do_swap_page() with same entry
<direct swapin path> <direct swapin path>
<alloc page A> <alloc page B>
swap_read_folio() <- read to page A swap_read_folio() <- read to page B
<slow on later locks or interrupt> <finished swapin first>
... set_pte_at()
swap_free() <- entry is free
<write to page B, now page A stalled>
<swap out page B to same swap entry>
pte_same() <- Check pass, PTE seems
unchanged, but page A
is stalled!
swap_free() <- page B content lost!
set_pte_at() <- staled page A installed!
And besides, for ZRAM, swap_free() allows the swap device to discard the
entry content, so even if page (B) is not modified, if swap_read_folio()
on CPU0 happens later than swap_free() on CPU1, it may also cause data
loss.
To fix this, reuse swapcache_prepare which will pin the swap entry using
the cache flag, and allow only one thread to swap it in, also prevent any
parallel code from putting the entry in the cache. Release the pin after
PT unlocked.
Racers just loop and wait since it's a rare and very short event. A
schedule_timeout_uninterruptible(1) call is added to avoid repeated page
faults wasting too much CPU, causing livelock or adding too much noise to
perf statistics. A similar livelock issue was described in commit
029c4628b2eb ("mm: swap: get rid of livelock in swapin readahead")
Reproducer:
This race issue can be triggered easily using a well constructed
reproducer and patched brd (with a delay in read path) [1]:
With latest 6.8 mainline, race caused data loss can be observed easily:
$ gcc -g -lpthread test-thread-swap-race.c && ./a.out
Polulating 32MB of memory region...
Keep swapping out...
Starting round 0...
Spawning 65536 workers...
32746 workers spawned, wait for done...
Round 0: Error on 0x5aa00, expected 32746, got 32743, 3 data loss!
Round 0: Error on 0x395200, expected 32746, got 32743, 3 data loss!
Round 0: Error on 0x3fd000, expected 32746, got 32737, 9 data loss!
Round 0 Failed, 15 data loss!
This reproducer spawns multiple threads sharing the same memory region
using a small swap device. Every two threads updates mapped pages one by
one in opposite direction trying to create a race, with one dedicated
thread keep swapping out the data out using madvise.
The reproducer created a reproduce rate of about once every 5 minutes, so
the race should be totally possible in production.
After this patch, I ran the reproducer for over a few hundred rounds and
no data loss observed.
Performance overhead is minimal, microbenchmark swapin 10G from 32G
zram:
Before: 10934698 us
After: 11157121 us
Cached: 13155355 us (Dropping SWP_SYNCHRONOUS_IO flag)
[kasong@tencent.com: v4]
Security readout for executives and security teams
Plain-English summary
A Linux swap race can install an outdated memory page when multiple threads load the same swap entry. This can corrupt or lose data. Triggering is local and requires low privileges according to the supplied CVSS assessment. Systems using synchronous swap I/O, particularly ZRAM under concurrent workloads, warrant prompt patch review.
Executive priority
Treat this as a high-priority reliability and security update for exposed Linux systems, particularly shared or workload-dense hosts using ZRAM. Expedite normal patching and validation. It is not supported as an internet-wide emergency because access is local and the supplied sources show no active exploitation or KEV listing.
Technical view
Concurrent do_swap_page paths can allocate different pages for one swap entry. One path may free and reuse the entry while another passes pte_same because of an ABA condition, then installs stale data. ZRAM may also discard the freed entry. The fix serializes swap-in using swapcache_prepare and pins the entry until the page-table lock is released.
Likely exposure
The bundle identifies Linux versions including 4.15, 6.1.80, 6.6.19, 6.7.7, and 6.8 as affected. Its version data also contains an unexplained "0" and duplicate identifiers, so exact ranges are unclear. Exposure is most plausible where local users or workloads can create concurrent swap activity on SWP_SYNCHRONOUS_IO devices, including ZRAM.
Exploitation context
The bundle reports no KEV listing and provides no evidence of active malicious exploitation. Maintainers reproduced data loss approximately once every five minutes under deliberately constructed concurrent swap conditions. This demonstrates practical reproducibility, but not exploitation in the wild. The supplied CVSS vector describes local access, low privileges, no user interaction, and high potential impact.
Researcher notes
The core failure is an ABA race around pte_same after a swap entry is freed and reused. Impact includes stale-page installation and data loss; ZRAM introduces an additional discard-related loss case. The source reports no observed loss after hundreds of patched test rounds and characterizes performance overhead as minimal. Exact vulnerable release boundaries require vendor confirmation.
Mitigation direction
Update through the Linux distributor to a kernel incorporating the applicable cited stable fix.
Confirm vendor guidance because the supplied affected-version data does not define a reliable complete range.
Prioritize systems using ZRAM or other synchronous-I/O swap paths under concurrent workloads.
If patching is delayed, request vendor-supported mitigations; none are specified in the bundle.
Validation and detection
Inventory running kernel versions and identify systems using swap, especially ZRAM.
Confirm each installed kernel contains the applicable stable commit or distributor backport.
Review distributor advisories to resolve the bundle's ambiguous affected-version entries.
After updating, run approved swap-integrity regression tests and check for data corruption.
Generated from the cited source records. This long-tail analysis has not been individually reviewed by a named human.
Potential ATT&CK relevance
Conservative CVE-to-ATT&CK context
These mappings and lookup hints may be relevant to the vulnerability behavior, CWE, affected product, or exposure path. Glexia-inferred context is not an official MITRE, ATT&CK, CWE, or CVE Program mapping.
ATT&CK lookup starting points
Use these exact CWE pages and searches to review the Glexia ATT&CK library from this CVE's weakness and description context.
cve · low confidence lookup
CVE-2024-26759 mapping review
Open the CVE-to-ATT&CK bridge for reviewed, inferred, or future official mappings tied to this CVE.
These fields come from the CVE record and ADP containers, not from Glexia's Take. They preserve time-varying source decisions such as CISA SSVC, KEV status, CVSS metrics, and provider references.
We collect every scored CVSS vector available in the official CNA and ADP containers. When more than one version is present, the table keeps the source vectors side by side instead of collapsing them into the highest score.