CVE-2021-46939: tracing: Restructure trace_clock_global() to never block
In the Linux kernel, the following vulnerability has been resolved:
tracing: Restructure trace_clock_global() to never block
It was reported that a fix to the ring buffer recursion detection would
cause a hung machine when performing suspend / resume testing. The
following backtrace was extracted from debugging that case:
Call Trace:
trace_clock_global+0x91/0xa0
__rb_reserve_next+0x237/0x460
ring_buffer_lock_reserve+0x12a/0x3f0
trace_buffer_lock_reserve+0x10/0x50
__trace_graph_return+0x1f/0x80
trace_graph_return+0xb7/0xf0
? trace_clock_global+0x91/0xa0
ftrace_return_to_handler+0x8b/0xf0
? pv_hash+0xa0/0xa0
return_to_handler+0x15/0x30
? ftrace_graph_caller+0xa0/0xa0
? trace_clock_global+0x91/0xa0
? __rb_reserve_next+0x237/0x460
? ring_buffer_lock_reserve+0x12a/0x3f0
? trace_event_buffer_lock_reserve+0x3c/0x120
? trace_event_buffer_reserve+0x6b/0xc0
? trace_event_raw_event_device_pm_callback_start+0x125/0x2d0
? dpm_run_callback+0x3b/0xc0
? pm_ops_is_empty+0x50/0x50
? platform_get_irq_byname_optional+0x90/0x90
? trace_device_pm_callback_start+0x82/0xd0
? dpm_run_callback+0x49/0xc0
With the following RIP:
RIP: 0010:native_queued_spin_lock_slowpath+0x69/0x200
Since the fix to the recursion detection would allow a single recursion to
happen while tracing, this lead to the trace_clock_global() taking a spin
lock and then trying to take it again:
ring_buffer_lock_reserve() {
trace_clock_global() {
arch_spin_lock() {
queued_spin_lock_slowpath() {
/* lock taken */
(something else gets traced by function graph tracer)
ring_buffer_lock_reserve() {
trace_clock_global() {
arch_spin_lock() {
queued_spin_lock_slowpath() {
/* DEAD LOCK! */
Tracing should *never* block, as it can lead to strange lockups like the
above.
Restructure the trace_clock_global() code to instead of simply taking a
lock to update the recorded "prev_time" simply use it, as two events
happening on two different CPUs that calls this at the same time, really
doesn't matter which one goes first. Use a trylock to grab the lock for
updating the prev_time, and if it fails, simply try again the next time.
If it failed to be taken, that means something else is already updating
it.
Bugzilla: https://bugzilla.kernel.org/show_bug.cgi?id=212761
Security readout for executives and security teams
Plain-English summary
This Linux kernel flaw can hang a machine when kernel tracing recurses into trace_clock_global() and deadlocks. The business impact is availability, not data theft. Evidence points to local, low-privilege exposure and no cited active exploitation.
Executive priority
Treat as a normal-priority availability fix for Linux fleets, with faster action on multi-user or operationally critical hosts. It does not justify emergency response unless vulnerable systems are stability-sensitive or local users can exercise tracing paths.
Technical view
trace_clock_global() took a spin lock while tracing. A permitted single recursion could re-enter ring_buffer_lock_reserve(), call trace_clock_global() again, and block on the same lock, causing a deadlock. The kernel fix changes the update path to avoid blocking by using a trylock.
Likely exposure
Linux systems running affected kernel builds are the relevant exposure. Risk is highest where local users or operational tooling can trigger kernel tracing paths, including function graph tracing during suspend or resume testing. Network-only exposure is not supported by the provided CVSS vector.
Exploitation context
The CVSS vector is local, low complexity, low privileges, no user interaction, and high availability impact. The source bundle marks KEV as false and provides no evidence of public exploitation or in-the-wild attacks.
Researcher notes
The issue is CWE-400 resource management through a non-blocking requirement violation in tracing. The public record is clear on root cause and fix direction, but the supplied affected-version data is commit/version oriented and should be reconciled through distro advisories.
Mitigation direction
Update to a vendor-supported Linux kernel containing the referenced stable fixes.
Check Linux distribution advisories for the exact fixed package version.
Prioritize shared systems where local users can access tracing features.
Avoid relying on custom workarounds unless vendor guidance confirms them.
Validation and detection
Inventory Linux kernel versions across servers, workstations, and appliances.
Compare installed kernels against vendor advisories and referenced stable commits.
Review whether tracing or function graph tracing is enabled or delegated to users.
Confirm patched systems no longer map to affected kernel records.
Generated from the cited source records. This long-tail analysis has not been individually reviewed by a named human.
Potential ATT&CK relevance
Conservative CVE-to-ATT&CK context
These mappings and lookup hints may be relevant to the vulnerability behavior, CWE, affected product, or exposure path. Glexia-inferred context is not an official MITRE, ATT&CK, CWE, or CVE Program mapping.
ATT&CK lookup starting points
Use these exact CWE pages and searches to review the Glexia ATT&CK library from this CVE's weakness and description context.
cwe · low confidence lookup
CWE-400: Exact CWE lookup
Use the exact CWE identifier as the starting point before reviewing related ATT&CK behavior. Open the exact CWE lookup page first, then review the ATT&CK searches from that MITRE weakness context. This is a Glexia lookup hint, not an official ATT&CK mapping.
These fields come from the CVE record and ADP containers, not from Glexia's Take. They preserve time-varying source decisions such as CISA SSVC, KEV status, CVSS metrics, and provider references.
We collect every scored CVSS vector available in the official CNA and ADP containers. When more than one version is present, the table keeps the source vectors side by side instead of collapsing them into the highest score.
CWE links open Glexia weakness intelligence pages with official CWE context, developer remediation guidance, and related CVE mappings.
CWE-400 · source CWE mapping
Uncontrolled Resource Consumption
Uncontrolled Resource Consumption represents a recurring weakness pattern that can create exploitable paths when design, validation, or implementation controls are missing.