| CVE |
Vendors |
Products |
Updated |
CVSS v3.1 |
| In the Linux kernel, the following vulnerability has been resolved:
exec: Cleanup POSIX timers right after de_thread()
A per-thread CPU timer holds a reference to the PID of the thread it is
attached to and, while it is armed, its node is queued in that thread's
posix_cputimers. The task is looked up by that PID.
When a non-leader thread exec()s, de_thread() changes which task owns
that PID. pid_task(timer->it.cpu.pid, PIDTYPE_PID) then returns NULL,
but the node is still queued on tsk, which is alive. timer_lock_sighand()
takes a failed lookup to mean that the node is already dequeued, so it
has nothing to undo.
begin_new_exec() calls posix_cpu_timers_exit(me) right after
exec_task_namespaces() and that removes the leftover node, so the state
normally stays invisible. But bprm->point_of_no_return is set before
de_thread(), so if unshare_files(), set_mm_exe_file(), exec_mmap() or
exec_task_namespaces() fails, the task dies before it gets there.
exit_itimers() then frees the k_itimer while its node is still queued,
and reaping tsk later erases that freed node from the rbtree.
In short:
the non-leader thread B the parent
timer_create(CLOCK_THREAD_CPUTIME_ID)
timer_settime()
arm_timer() // the node is queued on B
execve()
de_thread(B)
exchange_tids(B, leader) // B's PID now belongs to the leader
release_task(leader)
__exit_signal(leader)
posix_cpu_timers_exit(leader) // cleans leader's queue, not B's
__unhash_process(leader) // that PID has no task anymore
exec_mmap()
mmap_read_lock_killable(old_mm)
kill(B, SIGKILL)
// -EINTR
get_signal()
do_exit()
exit_itimers()
posix_timer_delete()
posix_cpu_timer_del()
posix_timer_unhash_and_free() // freed while still queued
wait4()
release_task(B)
posix_cpu_timers_exit(B)
cleanup_timerqueue()
timerqueue_del() // use-after-free
Move the POSIX timer cleanup right after de_thread() before any of the
later failure conditions brings the task into do_exit().
[ tglx: Move the cleanup right after de_thread() ] |
| In the Linux kernel, the following vulnerability has been resolved:
fs/dax: check zero or empty entry before converting xarray entry
Calling dax_to_folio() with empty entry causes kernel panic below when
booting a VM with DAX enabled storage.
This patch checks empty entry before calling dax_to_folio() on
dax_associate_entry(), dax_disassociate_entry(), and dax_busy_page().
Commit 98c183a4fccf ("fs/dax: don't disassociate zero page entries") added
guards in the associate and disassociate paths, but the guards still come
after dax_to_folio(), and dax_busy_page() still has the same problem.
[ 0.737679] EXT4-fs (pmem0p1): mounted filesystem 79676804-7c8b-491a-b2a6-9bae3c72af70 ro with ordered data mode. Quota mode: disabled.
[ 0.737891] VFS: Mounted root (ext4 filesystem) readonly on device 259:1.
[ 0.739119] devtmpfs: mounted
[ 0.739476] Freeing unused kernel memory: 1920K
[ 0.740156] Run /sbin/init as init process
[ 0.740229] with arguments:
[ 0.740286] /sbin/init
[ 0.740321] with environment:
[ 0.740369] HOME=/
[ 0.740400] TERM=linux
[ 0.743162] Unable to handle kernel paging request at virtual address fffffdffbf000008
[ 0.743285] Mem abort info:
[ 0.743316] ESR = 0x0000000096000006
[ 0.743371] EC = 0x25: DABT (current EL), IL = 32 bits
[ 0.743444] SET = 0, FnV = 0
[ 0.743489] EA = 0, S1PTW = 0
[ 0.743545] FSC = 0x06: level 2 translation fault
[ 0.743610] Data abort info:
[ 0.743656] ISV = 0, ISS = 0x00000006, ISS2 = 0x00000000
[ 0.743720] CM = 0, WnR = 0, TnD = 0, TagAccess = 0
[ 0.743785] GCS = 0, Overlay = 0, DirtyBit = 0, Xs = 0
[ 0.743848] swapper pgtable: 4k pages, 48-bit VAs, pgdp=00000000b9d17000
[ 0.743931] [fffffdffbf000008] pgd=10000000bfa3d403, p4d=10000000bfa3d403, pud=1000000040bfe403, pmd=0000000000000000
[ 0.744070] Internal error: Oops: 0000000096000006 [#1] SMP
[ 0.748888] CPU: 0 UID: 0 PID: 1 Comm: init Not tainted 6.18.4 #1 NONE
[ 0.749421] pstate: 004000c5 (nzcv daIF +PAN -UAO -TCO -DIT -SSBS BTYPE=--)
[ 0.749969] pc : dax_disassociate_entry.constprop.0+0x20/0x50
[ 0.750444] lr : dax_insert_entry+0xcc/0x408
[ 0.750802] sp : ffff80008000b9e0
[ 0.751083] x29: ffff80008000b9e0 x28: 0000000000000000 x27: 0000000000000000
[ 0.751682] x26: 0000000001963d01 x25: ffff0000004f7d90 x24: 0000000000000000
[ 0.752264] x23: 0000000000000000 x22: ffff80008000bcc8 x21: 0000000000000011
[ 0.752836] x20: ffff80008000ba90 x19: 0000000001963d01 x18: 0000000000000000
[ 0.753407] x17: 0000000000000000 x16: 0000000000000000 x15: 0000000000000000
[ 0.753970] x14: ffffbf3154b9ae70 x13: 0000000000000000 x12: ffffbf3154b9ae70
[ 0.754548] x11: ffffffffffffffff x10: 0000000000000000 x9 : 0000000000000000
[ 0.755122] x8 : 000000000000000d x7 : 000000000000001f x6 : 0000000000000000
[ 0.755707] x5 : 0000000000000000 x4 : 0000000000000000 x3 : fffffdffc0000000
[ 0.756287] x2 : 0000000000000008 x1 : 0000000040000000 x0 : fffffdffbf000000
[ 0.756871] Call trace:
[ 0.757107] dax_disassociate_entry.constprop.0+0x20/0x50 (P)
[ 0.757592] dax_iomap_pte_fault+0x4fc/0x808
[ 0.757951] dax_iomap_fault+0x28/0x30
[ 0.758258] ext4_dax_huge_fault+0x80/0x2dc
[ 0.758594] ext4_dax_fault+0x10/0x3c
[ 0.758892] __do_fault+0x38/0x12c
[ 0.759175] __handle_mm_fault+0x530/0xcf0
[ 0.759518] handle_mm_fault+0xe4/0x230
[ 0.759833] do_page_fault+0x17c/0x4dc
[ 0.760144] do_translation_fault+0x30/0x38
[ 0.760483] do_mem_abort+0x40/0x8c
[ 0.760771] el0_ia+0x4c/0x170
[ 0.761032] el0t_64_sync_handler+0xd8/0xdc
[ 0.761371] el0t_64_sync+0x168/0x16c
[ 0.761677] Code: f9453021 f2dfbfe3 cb813080 8b001860 (f9400401)
[ 0.762168] ---[ end trace 0000000000000000 ]---
[ 0.762550] note: init[1] exited with irqs disabled
[ 0.762631] Kernel panic - not syncing: Attempted to kill init! exitcode=0x0000000b |
| In the Linux kernel, the following vulnerability has been resolved:
signal: Prevent exec() race
Hyunwoo debugged the following KASAN UAF splat:
BUG: KASAN: slab-use-after-free in __send_signal_locked+0xb27/0xba0
Write of size 8 at addr ffff888007ed80c8 by task poc/79
...
Call Trace:
__send_signal_locked+0xb27/0xba0
do_send_sig_info+0xa7/0x160
do_send_specific+0x76/0xa0
__x64_sys_tgkill+0x193/0x270
...
Allocated by task 80:
do_timer_create+0x1a4/0x1030
__x64_sys_timer_create+0x145/0x190
...
Freed by task 12:
kmem_cache_free_bulk+0x1f8/0x4a0
kvfree_rcu_bulk+0x14f/0x1c0
kfree_rcu_work+0x128/0x1a0
...
Last potentially related work creation:
kvfree_call_rcu+0x39/0x390
__flush_itimer_signals+0x211/0x320
flush_itimer_signals+0x47/0x90
begin_new_exec+0xa6b/0x28c0
It turned out that this happens with a non-leader exec() as Hyunwoo
explained:
de_thread() calls exchange_tids() before release_task(leader), so the
struct pid held by a SIGEV_THREAD_ID timer created against the leader's tid
now points to the thread which called execve(). pid_task() returns that
thread and lock_task_sighand() on it succeeds.
If the timer signal is blocked, its sigqueue stays queued on the leader's
task::pending. The next expiry of that timer can then run while
release_task() flushes the queue.
posixtimer_send_sigqueue() checks whether the sigqueue is already queued
with a plain list_empty(), which only reads list_head::next.
list_del_init() is not atomic and INIT_LIST_HEAD() stores list_head::next
before list_head::prev, so the check can pass in between. list_add_tail()
queues the entry on the task::pending of the live thread, and the
list_head::prev store from the flush then overwrites the list_head::prev
link that list_add_tail() has just set.
__flush_itimer_signals() does not undo that either. With list_head::prev
pointing at the entry itself, its list_del_init() only stores the same
values again, so the entry is not removed from the list. It is still there
after the last reference is dropped and the timer is freed by RCU, and the
list_add_tail() of a later tgkill() follows that list_head::prev into the
freed timer.
This problem surfaced with the recent commit which moved the sigqueue flush
out of the sighand lock held region.
Hyonwoo proposed to fix this by using list_del_init_careful(), but that
just papers over the problem. After some disucssions and various attempts
to solve it, Eric pointed out that there is no reason to flush
task::pending late in release_task() and it should be done in
exit_signals() already.
As nothing can collect and deliver signals which are queued in a dying
task's pending queue, there is no reason to delay it further.
But it has to be ensured that no signals can be queued into it after that
point. exit_signals() sets PF_EXITING in task::flags, which can be used as
an indicator for this.
Cure it by:
- Preventing signal queueing for task private signals (PIDTYPE_PID) when
the task has PF_EXITING set in __send_signal_locked() and in
posixtimer_send_sigqueue().
- Protecting the unlocked setting of PF_EXITING in exit_signals() for the
task group empty and the group exit case with sighand lock
- Flushing task::pending signals right there.
Optimize that by moving the whole pending list to an on-stack list head
under sighand lock and free the signals without the lock held.
There has been quite some discussion about the lockless flush and the
non-leader exec case on weakly ordered systems. The problem is that a third
party which tries to send a posix timer signal relies on the PID lookup to
find the target task and that lookup might result in the new leader when
the signal was originaly directed to the old leader. In case that the
signal was queued on the old leader then the lockless flush raised a
concern over the following situation:
old_leader new_leader third party
A: flush_list() // list_del_in
---truncated--- |
| In the Linux kernel, the following vulnerability has been resolved:
swiotlb: use the adjusted address for the highmem page lookup
swiotlb_bounce() reads the page frame number from the slot's recorded
orig_addr, then advances orig_addr by tlb_offset to reach the address
the caller asked about. The highmem branch mixes the two: the offset
within the page comes from the adjusted address, the page from the value
before it.
Once the adjustment crosses a page boundary the pair no longer describes
one location, and the whole copy lands one page below the intended one
for a positive tlb_offset, one above for a negative one. DMA_FROM_DEVICE
writes the device data over the wrong page and leaves the intended one
stale, DMA_TO_DEVICE feeds the device from a page the mapping may not
cover. Partial syncs through dma_sync_single_range_for_*() are what make
tlb_offset non-zero.
The branch test is picked the same way, so a slot recorded in lowmem can
be adjusted into highmem and the lowmem path then hands a highmem
address to phys_to_virt().
Take both from orig_addr once it is final and keep pfn in the branch
that uses it. PhysHighMem() asks the question straight from the address,
as dma-debug already does. |
| In the Linux kernel, the following vulnerability has been resolved:
RDMA/ucma: Serialize join and leave on copy_to_user failure
rdma_join_multicast() queues RoCE work that later reads the ucma_multicast
through event->param.ud.private_data, then list_add()s the CMA multicast
at the head of id_priv->mc_list. rdma_leave_multicast() matches only by
sockaddr and destroys the first hit.
ucma_process_join() used to drop ctx->mutex after a successful join and
retake it only if copy_to_user() failed. Two concurrent JOIN_MCAST calls
with the same address can therefore insert a second CMA entry before the
first thread's leave. leave then cancels the newer work and the older
worker still dereferences the ucma_multicast that the first thread frees.
Keep ctx->mutex held from rdma_join_multicast() through copy_to_user() and,
on -EFAULT, through rdma_leave_multicast() so leave cannot miss this join.
Do not leave if join itself failed: that path never published this address
on mc_list, and a leave-by-addr would destroy an earlier successful join. |
| In the Linux kernel, the following vulnerability has been resolved:
RDMA/core: fix refcount bug in iwpm_get_nlmsg_request()
iwpm_get_nlmsg_request() initializes refcount _after_ list_add_tail()
making it accessible to global list where another CPU can kref_get()
on nlmsg_request causing a refcount "addition on 0" bug. Fix this
by initializing kref _before_ list_add_tail() so refcount for
nlmsg_request can be incremented/decremented normally. In addition,
also initialize every field before list_add_tail(). |
| In the Linux kernel, the following vulnerability has been resolved:
openvswitch: avoid reallocating confirmed conntrack labels
ovs_ct_get_conn_labels() adds the labels extension when a conntrack
entry does not have one. Confirmed conntracks can be read locklessly,
so adding an extension may reallocate and free the extension block
while another CPU accesses it.
Only add the extension for unconfirmed conntracks. A confirmed
conntrack without labels now fails the caller's label operation instead
of reallocating its extension storage. |
| In the Linux kernel, the following vulnerability has been resolved:
arm64: hibernate: pass HVC_SET_VECTORS args to the resume hvc
swsusp_arch_suspend_exit() reinstalls the restored kernel's hyp stub
vectors with an hvc, but never passes the arguments. x0 is not set to
HVC_SET_VECTORS and x1 is not set to the vector address, so the stub
dispatch falls through and returns without writing vbar_el2. EL2 is
left pointing at the trans_pgd copy of the vectors, a page that
swsusp_free() releases right after resume.
Set the arguments up the same way __hyp_set_vectors() does.
Without this fix, Vladimir was able to trigger a hang when resuming from
hibernation with CONFIG_PAGE_POISONING=y and page_poison=on. |
| In the Linux kernel, the following vulnerability has been resolved:
Bluetooth: hci_codec: validate vendor codec count length
The Read Local Supported Codecs parsers consume the variable-sized
standard codec array before parsing the vendor codec count. Although the
initial reply-size check includes a vendor count byte in the fixed layout,
it does not guarantee that the byte remains after the standard codec array.
If a controller reply ends immediately after that array, calculating the
vendor codec array size reads vnd_codecs->num beyond the skb data. Use
skb_pull_data() to validate and consume each codec header before using its
count in both command variants. |
| In the Linux kernel, the following vulnerability has been resolved:
Bluetooth: hci_sync: Serialize local codec list cleanup
hci_dev_close_sync() clears hdev->local_codecs after releasing hdev->lock.
Codec list additions and both traversals in sco_sock_getsockopt() use that
lock, but the close path does not. A close and BT_CODEC query can therefore
interleave as follows:
hci_dev_close_sync() sco_sock_getsockopt()
hci_dev_lock()
fetch codec entry
hci_codec_list_clear()
kfree(entry)
read entry->id
The reader then accesses an entry which the close path has freed. KASAN
BUG: KASAN: slab-use-after-free in sco_sock_getsockopt+0xfa0/0xfe0
Read of size 1 at addr ffff8881001c3450
Call Trace:
sco_sock_getsockopt+0xfa0/0xfe0
do_sock_getsockopt+0x537/0x7b0
__sys_getsockopt+0xf2/0x170
Allocated by task 92:
hci_codec_list_add.isra.0+0x2c/0x440
hci_read_codec_capabilities+0x224/0x590
hci_read_supported_codecs+0x2c2/0x640
Freed by task 92:
kfree+0x131/0x3c0
hci_codec_list_clear+0xd8/0x160
hci_dev_close_sync+0x92a/0xfa0
Take hdev->lock around the clear operation at its existing point in the
close path. This makes the clear wait for active readers and prevents a new
traversal until the list is empty without changing teardown ordering. |
| In the Linux kernel, the following vulnerability has been resolved:
dma-buf/dma-fence: fix checking signaling bit for timeline and driver name v3
The patch "dma-buf: dma-fence: Fix potential NULL pointer dereference"
changed the check to test for the ops pointer instead of the signaled
bit to avoid a potential NULL dereference when the ops pointer has been
cleared.
The problem is now that the ops pointer is cleared only when neither the
release nor the wait callback is implemented and this isn't true for a lot
of dma_fence implementations yet. So those implementations lost the RCU
protection after signaling of the returned string resulting in potential
use after free.
Add the signaling check additional to the ops pointer check so that we
have both the protection against NULL dereference as well as the RCU
protection after signaling for the returned string.
v2: improve comments to note RCU protection and explain why we check
both signaling state and ops pointer
v3: some comment improvements suggested by Philip |
| In the Linux kernel, the following vulnerability has been resolved:
dma-buf: Fix silent overflow for phys vec to sgt
In case MMIO size is bigger than 4G and peer2peer DMA goes
through host bridge, we trigger a code path that assigns the
total linked IOVA (which is greater than 4G) to mapped_len.
Previously, `mapped_len` was declared as 32-bit `unsigned int`.
When accumulating `size_t` lengths, this leads to a silent wrap-around.
This truncation causes truncated lengths to be passed to functions
like `fill_sg_entry()`.
Fix this by changing `mapped_len` to `size_t` (64-bit). While
at it, fix similar potential overflow issues in `calc_sg_nents`
by using `check_add_overflow()` for `nents` and using
`unsigned int` for the loop iterator in `fill_sg_entry` to match. |
| In the Linux kernel, the following vulnerability has been resolved:
net: lan743x: fix RX checksum use-after-free
lan743x_rx_process_buffer() adds each non-first receive buffer to the
head skb's frag_list. On the last descriptor, lan743x_rx_trim_skb()
linearizes the head and frees the fragment skb metadata.
The checksum-success path then writes ip_summed through the local skb
pointer, which still points to the final fragment. This causes a
use-after-free write when a packet spans more than one receive buffer.
Set ip_summed on the surviving head skb instead. Multi-buffer receive
can occur after a live MTU increase because existing ring entries keep
their old buffer size until they are replenished.
A KUnit test invoking lan743x_rx_process_buffer() with a two-buffer
packet produced a one-byte KASAN use-after-free write before this change.
The same test passed after the change. The driver object also builds
with W=1. This was not tested on physical LAN743x hardware. |
| In the Linux kernel, the following vulnerability has been resolved:
scsi: core: Validate MODE SENSE lengths in scsi_cdl_enable()
scsi_cdl_enable() uses length fields returned by MODE SENSE to locate
the ATA feature mode page in a 64-byte stack buffer. A target can report
a total length shorter than its mode header and block descriptors. The
unsigned subtraction used for the MODE SELECT length can wrap, and the
separately computed buf_data can point beyond buf.
During automatic scan, enable is false, so the read-modify-write of
buf_data[4] can clear the low two bits of a target-selected
out-of-bounds stack byte. scsi_mode_select() can then copy up to 64
bytes from outside the buffer into the outgoing MODE SELECT payload,
disclosing stack contents to the target.
This is reachable while scanning a USB storage device that identifies as
an ATA device and advertises CDL support. No filesystem mount or
userspace access to the block device is required.
On upstream commit cee9395acd80 ("Linux 7.3-rc1"), a build-specific,
one-vCPU QEMU/Raw Gadget proof using QEMU-only multi-UDC allocator
sampling executed a fixed proof command inside the guest and created a
UID-0-owned marker during automatic enumeration, with KASLR and NX
enabled.
The issue was independently found during security research at Drivesec
S.r.l.
Cap the available length to the buffer size. Validate and consume the
mode header and block descriptor lengths before using the page, and
require the five bytes needed to access the CDL field. |
| In the Linux kernel, the following vulnerability has been resolved:
mm, swap: fix SWAP_USAGE_OFFLIST_BIT collision with real usage count
SWAP_USAGE_OFFLIST_BIT is embedded in the si->inuse_pages usage counter,
and is meant to sit above any value that counter can reach. However, it
is defined from BITS_PER_TYPE(atomic_t), so it is bit 30. On a system
with 4 KiB pages the flag collides with the usage count once that count
reaches 4 TiB.
swap_usage_in_pages() masks bit 30 out, so whenever the real count has
that bit set, every caller of it reads 4 TiB low:
* /proc/swaps understates Used by 4 TiB.
* A raw count of exactly 2^30 masks to zero, so try_to_unuse() takes its
"if (!swap_usage_in_pages(si)) goto success;" early exit and swapoff
tears the device down while pages are still swapped out. Nothing in
the rest of swapoff aborts the teardown, so those pages are lost.
Independently of swapoff, the collision also corrupts the counter and the
plist. On a device in normal use, a free that leaves bit 30 set in the
count makes swap_usage_sub() see the flag where there is only count, and
call add_to_avail_list(). It clears the bit with
fetch_and(~SWAP_USAGE_OFFLIST_BIT), leaving the stored count 4 TiB below
the real one, and calls plist_add() on a device that is already listed,
tripping the WARN_ON(!plist_node_empty(node)) in plist_add() and linking
the node a second time.
Change the definition of SWAP_USAGE_OFFLIST_BIT to be based on
atomic_long_t instead. Note that the usage counter field itself is of
this same type, so it is still a valid bit. |
| In the Linux kernel, the following vulnerability has been resolved:
mm/shrinker: fix bogus set_shrinker_bit() with cgroup.memory=nokmem
With cgroup.memory=nokmem, shrinker_memcg_alloc() bails out early and
never allocates an id, so shrinker->id keeps the 0 it got from the
kzalloc() in shrinker_alloc(). __list_lru_init() then copies that 0 into
lru->shrinker_id, where it looks like a valid bit index.
Nothing calls expand_shrinker_info() on nokmem either, so shrinker_nr_max
stays 0 and every memcg ends up with an empty map (map_nr_max == 0).
deferred_split_folio() hands a real memcg to __list_lru_add() regardless
of whether the lru is memcg aware, so the first THP queued in a cgroup
does set_shrinker_bit(memcg, nid, 0) and trips the bounds check:
WARNING: mm/shrinker.c:212 at set_shrinker_bit+0x7d/0x90, CPU#126
Call Trace:
<TASK>
deferred_split_folio+0x18c/0x220
map_anon_folio_pmd_nopf+0xdd/0x130
map_anon_folio_pmd_pf+0x14/0xb0
do_huge_pmd_anonymous_page+0x1a1/0x620
__handle_mm_fault+0xea9/0x10d0
handle_mm_fault+0xe5/0x320
do_user_addr_fault+0x1cc/0x870
exc_page_fault+0x81/0x1b0
asm_exc_page_fault+0x27/0x30
</TASK>
Harmless, the WARN_ON_ONCE() is what keeps the out of bounds unit[] read
from happening, but the id should not look valid in the first place.
Clear it before returning.
Two other spots could paper over this: drop the id in __list_lru_init()
when nokmem turns memcg_aware off, or make deferred_split_folio() pass
NULL like list_lru_add_obj() does. Both leave shrinker->id lying around
for the next caller, so fix it where the id is handed out. |
| In the Linux kernel, the following vulnerability has been resolved:
mm/vma: correctly unaccount on mmap_prepare() failure
__mmap_setup() accounts memory for relevant mappings via:
security_vm_enough_memory_mm()
-> __vm_enough_memory()
-> vm_acct_memory()
If __mmap_setup() fails, this indicates that this accounting did not take
place, and thus it's appropriate for __mmap_region() to jump to
abort_munmap.
However if call_mmap_prepare() fails, it also jumps there and any accounted
memory is not correctly unaccounted.
Fix this by handling each error separately. |
| In the Linux kernel, the following vulnerability has been resolved:
KEYS: encrypted: fix integer overflow of datablob_len
encrypted_key_alloc() stores datablob_len in a u16. It is computed from
multiple string and payload lengths. If the result exceeds U16_MAX, the
assignment truncates the allocation size. KASAN reports a 32760-byte
slab-out-of-bounds write when __ekey_init() copies the master key
description into the undersized buffer.
The total payload length stored in key->datalen is also a u16. Use
check_add_overflow() to reject values that do not fit either destination,
and use kzalloc_flex() for the flexible-array allocation. |
| In the Linux kernel, the following vulnerability has been resolved:
KEYS: trusted: Fix tpm2_load_cmd() boundary check
tpm2_load_cmd() does boundary checks against the ASN.1 size i.e.,
payload->blob_len. Address this by passing the decoded blob size to
tpm2_load_cmd(), and use it for the boundary checks. |
| In the Linux kernel, the following vulnerability has been resolved:
sched_ext: Close the pre-enable ops error claim window
scx_alloc_and_add_sched() publishes ops->priv before
scx_root_enable_workfn() switches the state to SCX_ENABLING. An error
claimed via scx_bpf_error_bstr() from an associated BPF program in that
window is consumed by scx_disable_workfn(), which takes the pre-enable
shortcut in scx_root_disable(). The shortcut returns without any teardown
and restores SCX_DISABLED with an unconditional scx_set_enable_state() xchg
racing the enable workfn's own transition. The enable then completes with
the claim consumed: the scheduler stays up but can never be disabled again,
and bpf_scx_unreg() frees it while still in use, resulting in a
use-after-free. Both WARN_ON_ONCE()s fire back to back:
WARNING: kernel/sched/ext/ext.c:7522 at
scx_root_enable_workfn+0xeec/0x1be0, CPU#3: scx_enable_help/276
WARNING: kernel/sched/ext/ext.c:6398 at scx_root_disable+0xb50/0xdb8,
CPU#0: sched_ext_helpe/664
scx_root_enable_workfn() switches to SCX_ENABLING before the scheduler
allocation, so ops->priv is never visible while SCX_DISABLED. The allocation
failure path restores SCX_DISABLED. |