← All reports

Race window lead to reads from deallocated memory

HighWebGPU (GPU process)Race

CVE: CVE-2026-43794 · Safari 26.6.1 · Released August 18, 2026 Impact: Processing maliciously crafted web content may lead to memory corruption Apple's description: A memory corruption issue was addressed with improved memory handling. Credit: Dung Do (@_piers2) of Calif.io

d4a150e | Bugzilla 317317

High. A size threshold silently flips a buffer upload from copy semantics to alias semantics, and nobody on either side of the process boundary held the storage past the function that borrowed it. Winning the window is genuinely hard — the memory must be freed and re-occupied before the GPU wires the pages — but the payoff is reading GPU-process memory back into a page-visible resource.

Zero-copy is the standard trick for moving large uploads to a GPU: rather than memcpy tens of megabytes into a driver-owned staging buffer, hand the driver the pages you already have and let it wire them for DMA. The cost of that trick is an ownership contract — the caller keeps the pages, and must keep them valid until the GPU is actually finished, not merely until the command buffer is committed. WebKit's Metal WebGPU backend takes this path for any transfer at or above 32 MiB, in both the GPU-process IPC receiver RemoteQueue and the Metal-side WebGPU::Queue upload path, and until this commit neither layer held an owning reference past the end of the call that borrowed the storage.

The angle: A page issuing a large enough WebGPU buffer or texture upload can get the GPU to read a freed, already-recycled range of graphics-process memory into a resource the page can then read back.

Source/WebKit/GPUProcess/graphics/WebGPU/RemoteQueue.cpp

+#if HAVE(WEBGPU_IMPLEMENTATION)
+#include <WebGPU/WebGPU.h>
+#include <WebGPU/WebGPUExt.h>
+#endif
+
+// For transfers at or above WGPU_LARGE_BUFFER_SIZE the backend uses newBufferWithBytesNoCopy and aliases `data`'s mapping; keep it alive until the GPU has consumed the bytes. Smaller transfers are copied into a Metal buffer synchronously, so `data` can be released as soon as we return.
+static void keepAliveUntilSubmittedWorkDone(WebCore::WebGPU::Queue& backing, RefPtr<WebCore::SharedMemory>&& data)
+{
+#if HAVE(WEBGPU_IMPLEMENTATION)
+ if (!data || data->size() < WGPU_LARGE_BUFFER_SIZE)
+ return;
+ backing.onSubmittedWorkDone([data = WTF::move(data)]() mutable {
+ data = nullptr;
+ });
+#else
+ // Only the Metal backend aliases the caller's storage, and it is the only WebGPU implementation.
+ UNUSED_PARAM(backing);
+ UNUSED_PARAM(data);
+#endif
+}
+
void RemoteQueue::writeBuffer(
...
- protect(m_backing)->writeBufferNoCopy(protect(*convertedBuffer), bufferOffset, data ? data->mutableSpan() : std::span<uint8_t> { }, 0, std::nullopt);
+ Ref backing = protect(m_backing);
+ backing->writeBufferNoCopy(protect(*convertedBuffer), bufferOffset, data->mutableSpan(), 0, std::nullopt);
+ keepAliveUntilSubmittedWorkDone(backing, WTF::move(data));
completionHandler(true);
}
 
void RemoteQueue::writeTexture(
...
auto convertedDataLayout = objectHeap->convertFromBacking(dataLayout);
- ASSERT(convertedDestination);
+ ASSERT(convertedDataLayout);
auto convertedSize = objectHeap->convertFromBacking(size);
ASSERT(convertedSize);
- if (!convertedDestination || !convertedDestination || !convertedSize || !data || data->size() <= WebGPU::maxCrossProcessResourceCopySize) {
+ if (!convertedDestination || !convertedDataLayout || !convertedSize || !data || data->size() <= WebGPU::maxCrossProcessResourceCopySize) {
completionHandler(false);
return;
}
 
- protect(m_backing)->writeTexture(*convertedDestination, data ? data->mutableSpan() : std::span<uint8_t> { }, *convertedDataLayout, *convertedSize);
+ Ref backing = protect(m_backing);
+ backing->writeTexture(*convertedDestination, data->mutableSpan(), *convertedDataLayout, *convertedSize);
+ keepAliveUntilSubmittedWorkDone(backing, WTF::move(data));
completionHandler(true);
}

Source/WebGPU/WebGPU/Queue.mm

-constexpr static auto largeBufferSize = 32 * 1024 * 1024;
+constexpr static auto largeBufferSize = WGPU_LARGE_BUFFER_SIZE;
...
- if (noCopy)
+ if (noCopy) {
+ if (!newData.isEmpty()) {
+ // The MTLBuffer above was created with newBufferWithBytesNoCopy and aliases newData's storage; keep that storage alive until the GPU has consumed it.
+ __block Vector<uint8_t> retainedNewData = WTF::move(newData);
+ [m_commandBuffer addCompletedHandler:^(id<MTLCommandBuffer>) {
+ retainedNewData = { };
+ }];
+ }
finalizeBlitCommandEncoder();
+ }
}

Source/WebGPU/WebGPU/WebGPUExt.h

+// Threshold above which the Metal backend uses newBufferWithBytesNoCopy in writeBuffer / writeTexture
+// and aliases the caller's storage rather than copying. Callers passing transfers >= this size MUST
+// keep the source bytes alive until the GPU has consumed them (e.g. via addCompletedHandler).
+// Value is 32 * 1024 * 1024; written as a single integer literal so Swift's clang macro importer
+// picks it up as `WGPU_LARGE_BUFFER_SIZE` rather than skipping it.
+#define WGPU_LARGE_BUFFER_SIZE 33554432

Source/WebGPU/WebGPU/RenderPipeline.mm

HashMap<String, uint64_t> entryMap;
+ // Wrapping the bump would let the array-length entry alias a user binding and defeat the bounds check.
+ auto bumpForArrayLength = [&](uint32_t webBinding) -> std::optional<uint32_t> {
+ auto checked = checkedSum<uint32_t>(webBinding, limits().maxBindingsPerBindGroup);
+ if (checked.hasOverflowed())
+ return std::nullopt;
+ return checked.value();
+ };
for (auto& entry : bindGroupLayout.entries) {
...
if (entryName.endsWith("_ArrayLength"_s)) {
- webBinding += limits().maxBindingsPerBindGroup;
+ auto bumped = bumpForArrayLength(webBinding);
+ if (!bumped)
+ return @"Binding index overflow in auto-generated layouts";
+ webBinding = *bumped;
isArrayLength = true;
}

Three distinct changes ride in this commit, and only the first is what the title describes.

The lifetime fix, in two layers. WebGPUExt.h gains WGPU_LARGE_BUFFER_SIZE (33554432, i.e. 32 MiB) as a shared macro carrying the ownership contract in its comment: at or above this size the Metal backend uses newBufferWithBytesNoCopy and aliases the caller's storage instead of copying it, so callers must keep the source bytes alive until the GPU has consumed them. The macro is deliberately written as a single integer literal so Swift's clang macro importer picks it up — Queue.swift and Queue.mm both drop their private 32 * 1024 * 1024 literals in favour of it. In Queue.mm, the noCopy branch of the texture-upload path now moves the local Vector<uint8_t> newData into a __block variable captured by an [m_commandBuffer addCompletedHandler:] block that clears it (retainedNewData = { }) only after the GPU finishes the staging command buffer; finalizeBlitCommandEncoder() moves inside the braces of the same if (noCopy). In RemoteQueue.cpp, a new static helper keepAliveUntilSubmittedWorkDone() retains the mapped RefPtr<WebCore::SharedMemory> inside a backing.onSubmittedWorkDone() callback for transfers at or above the threshold; both RemoteQueue::writeBuffer() and RemoteQueue::writeTexture() hoist Ref backing = protect(m_backing) and call it right after handing data->mutableSpan() to the backend.

The guard fix. The same RemoteQueue.cpp hunk repairs a copy-paste defect in writeTexture(): ASSERT(convertedDestination) becomes ASSERT(convertedDataLayout), and the guard if (!convertedDestination || !convertedDestination || ...) becomes if (!convertedDestination || !convertedDataLayout || ...), so the std::optional that gets dereferenced as *convertedDataLayout is actually checked before use. With data now guaranteed non-null by that same guard, the data ? data->mutableSpan() : std::span<uint8_t> { } ternaries at both call sites collapse to data->mutableSpan().

The overflow fix. In RenderPipeline.mm, Device::addPipelineLayouts() wraps both webBinding += limits().maxBindingsPerBindGroup bumps for _ArrayLength entries in a checkedSum<uint32_t> lambda that returns the error string @"Binding index overflow in auto-generated layouts" rather than wrapping. The added comment states the stake directly: wrapping would let the array-length entry alias a user binding and defeat the bounds check.

The GPU process split. WebGPU calls issued from JavaScript run in the WebContent process, but the actual Metal work happens in a separate GPU process. Each GPUQueue in content has a corresponding RemoteQueue receiver in the GPU process; messages arrive over an IPC stream, identifiers are converted back into backend objects, and the call is replayed against WebCore::WebGPU::Queue, which is implemented over Metal in Source/WebGPU/WebGPU/Queue.mm.

Getting large uploads across. Uploads bigger than WebGPU::maxCrossProcessResourceCopySize are too big to inline in an IPC message. Instead the sender allocates shared memory and transmits a WebCore::SharedMemoryHandle; the receiver turns that handle into a mapping with WebCore::SharedMemory::map(). WebCore::SharedMemory is reference counted — map() returns a RefPtr, and the mapping is torn down when the last reference goes away.

Copy versus alias. newBufferWithBytesNoCopy is a Metal API that creates an MTLBuffer aliasing caller-supplied, page-aligned host memory rather than copying it. Ownership of the pages stays with the caller; Metal wires them for GPU access. WGPU_LARGE_BUFFER_SIZE — 32 MiB — is the threshold above which WebKit's backend picks this strategy over a straight copy into a Metal-owned staging buffer.

Metal command buffers are asynchronous. Work is encoded into an encoder, the encoder is ended, the command buffer is committed, and the GPU executes it at some later point. [MTLCommandBuffer addCompletedHandler:] registers a block that runs once execution actually finishes; Queue::onSubmittedWorkDone() is WebGPU's API-level equivalent of the same signal. Separately, __block is an Objective-C storage qualifier: a __block variable captured by a block is mutable from inside it and, once the block is copied to the heap, is owned by the block's storage — which is what makes a block a viable place to park an object's lifetime.

Object heap lookups. WebGPU::ObjectHeap::convertFromBacking() translates an IPC-supplied identifier or descriptor into a backend value and returns std::optional, which is std::nullopt when the identifier isn't present. ASSERT is compiled out in release builds, so release-mode correctness rests entirely on the explicit if (!converted...) guards.

Bindings and array lengths. WGSL shaders address resources by @group/@binding index, and maxBindingsPerBindGroup is a device limit on those indices. WebKit's Metal backend synthesizes extra _ArrayLength bindings that carry the runtime length of a shader's runtime-sized arrays, which the generated Metal code uses for its bounds checks. To keep those synthetic bindings out of the user index space, auto-generated pipeline layouts offset their index by maxBindingsPerBindGroup.

The core bug is an ownership expiry: the last owning reference to storage the GPU still aliases is dropped at the end of the synchronous call, while the consumer of those bytes runs asynchronously.

  GPU process (CPU)                          GPU
  ─────────────────────────────────────      ────────────────────────
  map SharedMemory  ──► RefPtr data
  writeBufferNoCopy(data->mutableSpan())
    └─► newBufferWithBytesNoCopy(ptr,len)
    └─► encode blit, commit ───────────────►  (queued)
  return  ──► ~RefPtr ──► unmap pages
                                    ╎
   [range reclaimed / re-mapped]    ╎  ◄── window
                                    ╎
                                              wire pages, read src ──► dest

The diagram's window is bounded on the left by the destruction of the function-local RefPtr<WebCore::SharedMemory> data — which unmaps the shared memory — and on the right by the GPU actually wiring and reading those pages during the staging blit. Nothing connected the two. RemoteQueue::writeBuffer() and RemoteQueue::writeTexture() mapped the incoming handle, passed data->mutableSpan() down to the backend, and returned; one layer lower, the noCopy branch of the Metal upload path let its local Vector<uint8_t> newData fall out of scope at the end of the function, returning the heap block to the allocator. In both cases the command buffer that reads those bytes had only been encoded and committed, not completed. The commit message describes the resulting failure precisely as deallocated Vector or SharedMemory being "swapped in" over by user memory — the virtual range gets re-mapped or re-allocated for unrelated GPU-process data, the MTLBuffer still points at the same addresses, and the DMA reads whatever now lives there.

This is a lifetime bug with a timing window rather than a data race on a shared variable, and it is genuinely narrow: for the memory to be re-occupied, it must be freed and recycled before the GPU wires it, which is why the commit lands with no new test ("race window is incredibly small: memory must be freed prior to being wired to the GPU"). Winning it, though, is the whole game. The aliased MTLBuffer is the source of the staging blit, so the reclaimed range is read, not written — the primitive on a won race is an uncontrolled disclosure of whatever now occupies that range, delivered into a GPUBuffer or GPUTexture that web content can subsequently read back. That could carry heap pointers useful for defeating GPU-process ASLR, or pixel and buffer data belonging to other pages, since a single GPU process services multiple content processes and data recovered there is not necessarily same-origin with the attacking page. On a lost race the observable outcome is a GPU-process crash.

The fix pins ownership across the asynchronous window at both layers, using the completion signal as the release point:

__block Vector<uint8_t> retainedNewData = WTF::move(newData);
[m_commandBuffer addCompletedHandler:^(id<MTLCommandBuffer>) {
    retainedNewData = { };
}];

The __block capture moves the Vector into block storage, so its lifetime is now the block's lifetime, and the block outlives the function by construction. keepAliveUntilSubmittedWorkDone() does the structurally identical thing one layer up, parking the RefPtr<WebCore::SharedMemory> inside an onSubmittedWorkDone() callback — and it only does so for transfers at or above WGPU_LARGE_BUFFER_SIZE, because below that threshold the backend copies synchronously and there is nothing to keep alive. Note that the helper is the first code in RemoteQueue.cpp to even know the threshold exists; before this commit the constant lived only as a bare literal inside the Metal backend.

Reachability is unremarkable in the best way for an attacker: a plain web page issuing GPUQueue.writeBuffer or writeTexture with a sufficiently large payload drives the whole path, no compromised renderer required. What the bug buys is not itself a sandbox escape — it moves the attacker from WebContent into memory-disclosure territory inside a separate, differently-privileged sandbox that holds IOKit graphics access, the usual staging ground for kernel-facing follow-on work.

The two bundled fixes are unrelated in mechanism but share the same process boundary. The writeTexture() guard tested convertedDestination twice and never tested convertedDataLayout, then dereferenced *convertedDataLayout — so an IPC message naming an unknown data-layout backing reached a dereference of a disengaged std::optional, with the accompanying ASSERT compiled out in release. That path additionally requires an unresolvable backing identifier, which is natural for an already-compromised WebContent process rather than for a plain page. The _ArrayLength overflow is subtler: the synthetic array-length bindings are moved out of user index space by adding maxBindingsPerBindGroup to entry.webBinding — twice, in two separate loops — in plain uint32_t arithmetic. A sufficiently large webBinding wraps, and the array-length entry then lands on top of a real user binding. Since that slot supplies the runtime array size the generated shader uses for its bounds checks, aliasing it with a binding the page controls hands the page the length value those checks read.

Zero-copy upload paths dropped their last owning reference at function return while a committed MTLBuffer still aliased the pages, letting the GPU read a reclaimed range back into a page-visible resource.

The zero-copy contract here lives entirely in prose. The fix's own header comment — "Callers passing transfers >= this size MUST keep the source bytes alive" — is the enforcement mechanism, backed by a magic threshold that until this commit was duplicated as a bare literal in Queue.mm and Queue.swift, and absent entirely from RemoteQueue.cpp, which had no notion that a threshold existed. A size threshold that silently switches between copy and alias semantics is a bug generator: every call site must independently know that crossing 32 MiB changes who owns the memory, and the type system says nothing about it. Centralising the constant is a real improvement, but the API would be far harder to misuse if the no-copy entry points took an owning handle, or returned a completion token, instead of a bare std::span<uint8_t>.

The duplicated-operand typo is worth its own note, because it isn't a one-off. At this very revision, RemoteQueue::writeTextureWithCopy still carries ASSERT(convertedDestination) where it computes convertedDataLayout and still guards if (!convertedDestination || !convertedDestination || !convertedSize) before dereferencing *convertedDataLayout, and RemoteQueue::copyExternalImageToTexture computes convertedSource while guarding convertedDestination twice. The pattern propagates by copy-paste across the hand-written Remote* IPC entry points, and it is mechanically greppable.