[YARR] Fix incorrect offset when reading pattern character for Unicode backreference in JIT
CVE: CVE-2026-43740 · Safari 26.5.2 · Released June 29, 2026 Impact: Processing maliciously crafted web content may result in the disclosure of process memory Apple's description: The issue was addressed with improved memory handling. Credit: Nathaniel Oh (@calysteon), Arni Hardarson
Medium — a one-token offset expression where a constant belonged, and the wrong token is a script-controlled quantity. The commit message calls it a failed match; the address arithmetic says the same displacement walks off the front of the string buffer whenever the capture sits near position zero. Only a match/no-match oracle comes out the other side.
Regular-expression matching in JavaScriptCore compiles down to machine code, and machine code has no bounds checks unless the compiler emits them. YARR's JIT earns its speed by hoisting input-length checks out of individual terms and into a run of terms, which means every load it emits carries a compile-time bias that has to be subtracted back out at the point of use. That bias belongs to exactly one cursor — the one walking the subject string — and a backreference match runs two cursors at once.
The angle: A page can craft a regular expression whose match result depends on a character read from memory before the start of the string being matched — a script-visible oracle over adjacent heap bytes.
Source/JavaScriptCore/yarr/YarrJIT.cpp
JSTests/stress/regexp-backreference-unicode-offset.js
Patch Details
The functional change is one line in YarrGenerator::matchBackreference(). Inside the m_decodeSurrogatePairs branch — the path taken when the pattern carries the u or v flag over 16-bit text — the generated code loads the previously captured character so it can be compared against the character at the current subject position. That load was emitted as readCharacter(op.m_checkedOffset - term->inputPosition, character, patternIndex). The patch replaces the first argument with the constant 0.
The first argument of readCharacter() is a negative displacement subtracted from the index register when the load address is formed. It exists to undo YARR's speculative bias on the subject cursor. The index register supplied here is not the subject cursor — it is patternIndex, an absolute position into the capture. The non-Unicode branch, three lines above in the same if/else, already passed 0. The patch makes the two branches agree.
The commit also adds JSTests/stress/regexp-backreference-unicode-offset.js. The test runs /(.)\1c/u, /(.)\1cd/u and case-insensitive variants against non-BMP subjects 500 times, which is enough iterations to force YARR JIT compilation rather than staying in the interpreter. Its structure encodes the trigger condition directly: the file's own comment on the second case notes that /(.)\1/u — a backreference with nothing after it — "already worked".
Background
YARR. YARR is JavaScriptCore's regular-expression engine. It has two execution tiers: YarrInterpreter, a bytecode interpreter, and YarrJIT, which compiles a pattern into machine code once that pattern has been executed enough times to be worth compiling. Everything below concerns the JIT tier only.
Backreferences and the two cursors. A backreference term such as \1 matches the exact text that group 1 captured earlier in the same match. Matching it means walking that captured substring and comparing it code unit by code unit against the subject at the current position. Because the capture is a range inside the same subject buffer, the JIT keeps two cursors into that one buffer while a backreference is being matched: index, the current match position, and patternIndex, an absolute position inside the earlier capture, initialised from the capture-start slot of the output array.
Speculative input checks and m_checkedOffset. YARR does not emit a length check per term. It emits one check covering a run of terms and then lets each term index forward within the already-validated span. The consequence is that the subject cursor is, at any given term, biased ahead of that term's own logical position. m_checkedOffset records how far ahead; term->inputPosition records where the term sits. Any load keyed on the subject cursor therefore passes m_checkedOffset - inputPosition as a compensating negative offset.
readCharacter(). readCharacter(negativeOffset, resultReg, indexReg) emits a load of one character at characters + (indexReg - negativeOffset) * charSize. It is an address-forming primitive and nothing more — it performs no validation. Whatever bounds guarantee a given load enjoys comes from the surrounding checkInput machinery, not from readCharacter() itself.
m_decodeSurrogatePairs. This flag is set when the pattern uses the u/v flag over 16-bit text, so that a lead+trail surrogate pair is treated as a single code point rather than two independent units. That path routes the character load through the shared tryReadUnicodeChar() helper — which is why it uses the standard resultReg and then moves the value into patternCharacter, rather than loading straight into the destination as the non-Unicode path does.
Non-BMP escapes. \u{10000} denotes a code point encoded as two UTF-16 code units. A single . in a u-flag pattern therefore consumes two units of the subject, which is what makes the regression test's subjects two units per "character".
Analysis
The root cause is a coordinate-space mismatch: a bias belonging to the subject cursor was applied to an index that is already absolute.
Before: After:
patternIndex (absolute capture pos) patternIndex (absolute capture pos)
│ │
├─ delta = checkedOffset ├─ offset = 0
│ - inputPosition │
▼ ▼
load @ chars + (patternIndex-delta) load @ chars + patternIndex
│ │
├─ delta ≤ patternIndex → wrong char └─► correct captured char
└─ delta > patternIndex → OOB read
(before buffer start)
In the diagram, delta is op.m_checkedOffset - term->inputPosition. That quantity is zero only when nothing follows the backreference in the pattern, because then the speculative check span ends where the term ends and m_checkedOffset equals inputPosition. This is precisely why /(.)\1/u matched correctly while /(.)\1c/u did not, and why the regression test pairs those two patterns as the contrast case. Add terms after the backreference and delta grows by the number of code units those terms consume — one per literal ASCII character, two per non-BMP escape. delta is therefore a direct function of the pattern text, which is to say it is chosen by whoever wrote the regexp.
The two outcomes in the left column of the diagram are separated by a single comparison. When delta <= patternIndex, the displaced address still lands inside the subject string, and the generated code merely compares the current input character against the wrong captured character — the incorrect-match symptom the commit message describes. When delta > patternIndex, the index underflows past zero and the load reads from memory preceding the string's character buffer. Nothing stops it. The checkInput guards that YARR emits constrain the subject cursor's forward travel; they say nothing about patternIndex, and readCharacter() does not re-derive a bound from its arguments:
// readCharacter(negativeCharacterOffset, resultReg, indexReg)
// → load @ characters + (indexReg - negativeCharacterOffset) * charSize
// No comparison, no branch. The address is formed and the load is issued.
Reaching the underflow needs two things the pattern author controls together: enough terms after the backreference to push delta up, and a capture that begins near enough to the start of the subject that patternIndex stays below it. Both are ordinary pattern-and-input construction from script; nothing here requires exotic heap state. What comes back is not a value the script can read — the loaded code unit is consumed only by the equality comparison that decides whether the backreference matches. The observable is therefore a match/no-match answer about a byte in front of the string buffer: an oracle. Repeated over varying delta, that oracle reads out StringImpl header words and whatever heap contents sit adjacent, which is the shape of an ASLR defeat or the reconnaissance half of a later corruption chain. This weakens memory-safety confinement inside the WebContent process; it corrupts nothing, and any disclosure stays in the renderer's address space, so cross-process impact would still require a separate memory-corruption bug plus a separate sandbox escape. The fix restores the invariant that the pattern character is loaded from exactly patternIndex, in the coordinate space patternIndex actually lives in.
A pattern-controlled bias meant for the subject cursor was subtracted from an absolute capture cursor, letting a regexp match outcome depend on memory before the string buffer.
Insight
The correct code was already sitting three lines above: the non-Unicode branch passed 0 while the Unicode branch passed a bias expression. Paired specialised implementations of one logical operation — Unicode vs non-Unicode, 8-bit vs 16-bit, inlined vs helper-call — are a recurring source of YARR bugs precisely because the rarer branch drifts from the common one and is exercised by far fewer tests. Worth noting too: the security consequence is invisible in the commit message. The author frames it purely as "incorrectly failing to match", and only the address arithmetic reveals that the same wrong offset walks off the front of the buffer when the capture is near position zero. Correctness-framed YARR JIT offset fixes deserve a second read for exactly that reason.
Audit directions
-
One offset, two coordinate spaces. The invariant is that every bias argument must be paired with the cursor it was derived from. Narrow: grep
Source/JavaScriptCore/yarr/YarrJIT.cppforreadCharacter(,readCharacterDontDecode(andnegativeOffsetIndexedAddress(call sites whose index-register argument is not the defaultindex(patternIndex,endIndex,matchPos) and confirm the offset argument is the constant0; a non-constantm_checkedOffset-derived expression paired with a non-default index register is the match tell. Wider: the class shows up wherever a codebase carries a speculative-bias variable alongside absolute cursors — checkYarrInterpreter.cpp's backreference and Boyer-Moore search index math, and JSC's string-search fast paths, for expressions mixing acheckedOffset/lookaheadbias with an offset loaded from an output or capture array; in search results, look for arithmetic combining a compile-time bias with a runtime-loaded position. Widest: this is the general "two coordinate spaces, one offset" class present in any matcher or codec with a lookahead cursor plus a back-reference cursor — V8's irregexp backreference code, PCRE2's JIT, and any decoder tracking both a read pointer and a dictionary pointer. The invariant to carry over: if a variable's name encodes whose coordinate space it belongs to, any arithmetic mixing two such variables is a bug candidate. -
Loads that inherit their bounds guarantee from a different register. The invariant is that the register whose value forms the load address must be the register that was bounds-checked. Narrow: audit each
checkInput/checkNotEnoughInputsite inYarrJIT.cpp, enumerate which register the guard constrains, then list every load between that guard and the next one whose base index is a different register — thepatternIndex-based loads insidematchBackreference()andbacktrackBackReference()are the starting set. Wider: the same class appears in any JIT that hoists or shares bounds checks across a region — DFG/FTL array-access lowering where aCheckInBoundson one index is followed by loads using a derived or aliased index, and Wasm bounds-check elision over reused index temporaries; the tell is a guard and a load naming different index variables. Widest: the reusable principle is that a bounds check proves a property of a value, not of a memory region, and it applies to any compiler doing check elision or hoisting (V8 Turbofan, SpiderMonkey Ion, LLVM's elision assumptions). Carry the question: which SSA value was proven in range, and is that the exact value used to form the address? -
Divergence between paired specialised implementations. The invariant is that paired fast/slow or wide/narrow implementations must agree on argument semantics, not merely on results for common inputs. Narrow: walk the
#if ENABLE(YARR_JIT_UNICODE_EXPRESSIONS)/m_decodeSurrogatePairsbranches inYarrJIT.cppside by side with their non-Unicode counterparts and flag every call where the two branches pass structurally different arguments to the same helper — a constant on one side and a computed expression on the other is the match tell, and is exactly the shape of this bug. Wider: apply the same side-by-side diffing to WebKit's other duplicated-specialisation surfaces — 8-bit vs 16-bitStringImplcode paths, SIMD vs scalar fallbacks inWTFtext search, inlined vs thunk-call variants of the same JIT helper; look for helper calls whose argument lists differ in shape between siblings. Widest: any codebase maintaining N specialised copies of one algorithm — encoding-specialised parsers, vectorised/scalar kernel pairs, per-architecture backends — carries this class, and the portable audit move is differential testing that drives every specialisation over one shared input corpus and compares outputs. -
Match outcomes that depend on out-of-string data. Such paths convert silently into disclosure oracles, which is what turned this correctness bug into a CVE. Narrow: for each generated comparison in
matchBackreference()and the character-class matchers, trace the provenance of both compared operands and check whether either address can be formed from a runtime-loaded position with no intervening guard; the tell is a comparison whose success or failure is script-observable — through the match result, capture contents, orlastIndex— with one input coming from an unguarded load. Wider and beyond WebKit: the same question applies to any pattern engine that exposes match success to untrusted input while performing unchecked loads on backreference or dictionary cursors.