CVE-2025-37964: x86/mm: Eliminate window where TLB flushes may be inadvertently skipped
In the Linux kernel, the following vulnerability has been resolved:
x86/mm: Eliminate window where TLB flushes may be inadvertently skipped
tl;dr: There is a window in the mm switching code where the new CR3 is
set and the CPU should be getting TLB flushes for the new mm. But
should_flush_tlb() has a bug and suppresses the flush. Fix it by
widening the window where should_flush_tlb() sends an IPI.
Long Version:
=== History ===
There were a few things leading up to this.
First, updating mm_cpumask() was observed to be too expensive, so it was
made lazier. But being lazy caused too many unnecessary IPIs to CPUs
due to the now-lazy mm_cpumask(). So code was added to cull
mm_cpumask() periodically[2]. But that culling was a bit too aggressive
and skipped sending TLB flushes to CPUs that need them. So here we are
again.
=== Problem ===
The too-aggressive code in should_flush_tlb() strikes in this window:
// Turn on IPIs for this CPU/mm combination, but only
// if should_flush_tlb() agrees:
cpumask_set_cpu(cpu, mm_cpumask(next));
next_tlb_gen = atomic64_read(&next->context.tlb_gen);
choose_new_asid(next, next_tlb_gen, &new_asid, &need_flush);
load_new_mm_cr3(need_flush);
// ^ After 'need_flush' is set to false, IPIs *MUST*
// be sent to this CPU and not be ignored.
this_cpu_write(cpu_tlbstate.loaded_mm, next);
// ^ Not until this point does should_flush_tlb()
// become true!
should_flush_tlb() will suppress TLB flushes between load_new_mm_cr3()
and writing to 'loaded_mm', which is a window where they should not be
suppressed. Whoops.
=== Solution ===
Thankfully, the fuzzy "just about to write CR3" window is already marked
with loaded_mm==LOADED_MM_SWITCHING. Simply checking for that state in
should_flush_tlb() is sufficient to ensure that the CPU is targeted with
an IPI.
This will cause more TLB flush IPIs. But the window is relatively small
and I do not expect this to cause any kind of measurable performance
impact.
Update the comment where LOADED_MM_SWITCHING is written since it grew
yet another user.
Peter Z also raised a concern that should_flush_tlb() might not observe
'loaded_mm' and 'is_lazy' in the same order that switch_mm_irqs_off()
writes them. Add a barrier to ensure that they are observed in the
order they are written.
Security readout for executives and security teams
Plain-English summary
A Linux x86 memory-management race can leave a processor using stale memory mappings after switching address spaces. This could undermine confidentiality, integrity, or availability for systems where a low-privileged local user can execute code. The supplied evidence does not establish remote or active exploitation.
Executive priority
Prioritize remediation on multi-user, shared-compute, and other systems allowing untrusted local code. Treat this as high priority, but below remotely exploitable or confirmed-active threats unless local access is broadly available.
Technical view
During an x86 address-space switch, should_flush_tlb() can suppress a required TLB-flush IPI after CR3 changes but before loaded_mm is updated. The correction treats LOADED_MM_SWITCHING as requiring an IPI and adds ordering protection for related state observations.
Likely exposure
Exposure is limited to x86 Linux systems running affected kernel builds. The bundle mixes commit identifiers, release numbers, and an ambiguous โ0โ entry, so version-string matching alone is unreliable. Confirm exposure against the deployed distribution or product vendor's kernel advisory and backport status.
Exploitation context
The CVSS 3.1 score is 7.8: local access, low complexity, low privileges, no user interaction, and potentially high confidentiality, integrity, and availability impact. The CVE is not listed as KEV, and the supplied sources provide no evidence of active exploitation.
Researcher notes
The vulnerable window lies between load_new_mm_cr3() and publishing next through cpu_tlbstate.loaded_mm. The fix widens TLB-flush targeting during LOADED_MM_SWITCHING and adds a barrier. The supplied version data is insufficient for definitive branch-range reconstruction; use commit ancestry or vendor backport records.
Mitigation direction
Install a vendor-supported kernel containing the applicable stable fix or backport.
Follow distribution or appliance guidance for activating the updated kernel.
Restrict untrusted local accounts and workloads where immediate remediation is unavailable.
Review Debian and Siemens advisories when their products are present.
Validation and detection
Inventory running x86 kernel builds, including vendor package and build identifiers.
Compare each build with vendor advisories and documented backport status.
Confirm the running kernel includes the applicable upstream stable correction.
After maintenance, verify systems are running the updated kernel, not merely storing it.
Generated from the cited source records. This long-tail analysis has not been individually reviewed by a named human.
Potential ATT&CK relevance
Conservative CVE-to-ATT&CK context
These mappings and lookup hints may be relevant to the vulnerability behavior, CWE, affected product, or exposure path. Glexia-inferred context is not an official MITRE, ATT&CK, CWE, or CVE Program mapping.
ATT&CK lookup starting points
Use these exact CWE pages and searches to review the Glexia ATT&CK library from this CVE's weakness and description context.
cve ยท low confidence lookup
CVE-2025-37964 mapping review
Open the CVE-to-ATT&CK bridge for reviewed, inferred, or future official mappings tied to this CVE.
These fields come from the CVE record and ADP containers, not from Glexia's Take. They preserve time-varying source decisions such as CISA SSVC, KEV status, CVSS metrics, and provider references.
1CVSS vectors
3Timeline events
2ADP providers
9Source links
CVSS vector scores
1 official score
We collect every scored CVSS vector available in the official CNA and ADP containers. When more than one version is present, the table keeps the source vectors side by side instead of collapsing them into the highest score.