← All reports

[YARR] Fix incorrect offset when reading pattern character for Unicode backreference in JIT

CVE: CVE-2026-43740 · Safari 26.5.2 · Released June 29, 2026 Impact: Processing maliciously crafted web content may result in the disclosure of process memory Apple's description: The issue was addressed with improved memory handling. Credit: Nathaniel Oh (@calysteon), Arni Hardarson

Severity: Medium | Component: JSC YARR JIT | 2693828 | Bugzilla 308046

Medium — a one-token offset expression where a constant belonged, and the wrong token is a script-controlled quantity. The commit message calls it a failed match; the address arithmetic says the same displacement walks off the front of the string buffer whenever the capture sits near position zero. Only a match/no-match oracle comes out the other side.

Regular-expression matching in JavaScriptCore compiles down to machine code, and machine code has no bounds checks unless the compiler emits them. YARR's JIT earns its speed by hoisting input-length checks out of individual terms and into a run of terms, which means every load it emits carries a compile-time bias that has to be subtracted back out at the point of use. That bias belongs to exactly one cursor — the one walking the subject string — and a backreference match runs two cursors at once.

The angle: A page can craft a regular expression whose match result depends on a character read from memory before the start of the string being matched — a script-visible oracle over adjacent heap bytes.

Source/JavaScriptCore/yarr/YarrJIT.cpp

@@ -2011,7 +2011,7 @@ class YarrGenerator final : public YarrJITInfo {
else {
// For reading Unicode characters, use the standard resultReg so we can call the standard tryReadUnicodeChar()
// helper instead of emitting an inlined version.
- readCharacter(op.m_checkedOffset - term->inputPosition, character, patternIndex);
+ readCharacter(0, character, patternIndex);
m_jit.move(character, patternCharacter);
}
#else

JSTests/stress/regexp-backreference-unicode-offset.js

+for (var i = 0; i < 500; i++) {
+ // Basic case: backreference followed by a literal character.
+ // This triggers checkedOffset != inputPosition for the backreference term.
+ var result = /(.)\1c/u.exec("\u{10000}\u{10000}c");
+ shouldBe(result !== null, true, "...");
+
+ // Without trailing term (checkedOffset == inputPosition), this already worked.
+ var result2 = /(.)\1/u.exec("\u{10000}\u{10000}");
+ shouldBe(result2 !== null, true, "...");
+}

The functional change is one line in YarrGenerator::matchBackreference(). Inside the m_decodeSurrogatePairs branch — the path taken when the pattern carries the u or v flag over 16-bit text — the generated code loads the previously captured character so it can be compared against the character at the current subject position. That load was emitted as readCharacter(op.m_checkedOffset - term->inputPosition, character, patternIndex). The patch replaces the first argument with the constant 0.

The first argument of readCharacter() is a negative displacement subtracted from the index register when the load address is formed. It exists to undo YARR's speculative bias on the subject cursor. The index register supplied here is not the subject cursor — it is patternIndex, an absolute position into the capture. The non-Unicode branch, three lines above in the same if/else, already passed 0. The patch makes the two branches agree.

The commit also adds JSTests/stress/regexp-backreference-unicode-offset.js. The test runs /(.)\1c/u, /(.)\1cd/u and case-insensitive variants against non-BMP subjects 500 times, which is enough iterations to force YARR JIT compilation rather than staying in the interpreter. Its structure encodes the trigger condition directly: the file's own comment on the second case notes that /(.)\1/u — a backreference with nothing after it — "already worked".

YARR. YARR is JavaScriptCore's regular-expression engine. It has two execution tiers: YarrInterpreter, a bytecode interpreter, and YarrJIT, which compiles a pattern into machine code once that pattern has been executed enough times to be worth compiling. Everything below concerns the JIT tier only.

Backreferences and the two cursors. A backreference term such as \1 matches the exact text that group 1 captured earlier in the same match. Matching it means walking that captured substring and comparing it code unit by code unit against the subject at the current position. Because the capture is a range inside the same subject buffer, the JIT keeps two cursors into that one buffer while a backreference is being matched: index, the current match position, and patternIndex, an absolute position inside the earlier capture, initialised from the capture-start slot of the output array.

Speculative input checks and m_checkedOffset. YARR does not emit a length check per term. It emits one check covering a run of terms and then lets each term index forward within the already-validated span. The consequence is that the subject cursor is, at any given term, biased ahead of that term's own logical position. m_checkedOffset records how far ahead; term->inputPosition records where the term sits. Any load keyed on the subject cursor therefore passes m_checkedOffset - inputPosition as a compensating negative offset.

readCharacter(). readCharacter(negativeOffset, resultReg, indexReg) emits a load of one character at characters + (indexReg - negativeOffset) * charSize. It is an address-forming primitive and nothing more — it performs no validation. Whatever bounds guarantee a given load enjoys comes from the surrounding checkInput machinery, not from readCharacter() itself.

m_decodeSurrogatePairs. This flag is set when the pattern uses the u/v flag over 16-bit text, so that a lead+trail surrogate pair is treated as a single code point rather than two independent units. That path routes the character load through the shared tryReadUnicodeChar() helper — which is why it uses the standard resultReg and then moves the value into patternCharacter, rather than loading straight into the destination as the non-Unicode path does.

Non-BMP escapes. \u{10000} denotes a code point encoded as two UTF-16 code units. A single . in a u-flag pattern therefore consumes two units of the subject, which is what makes the regression test's subjects two units per "character".

The root cause is a coordinate-space mismatch: a bias belonging to the subject cursor was applied to an index that is already absolute.

  Before:                              After:
  patternIndex (absolute capture pos)  patternIndex (absolute capture pos)
    │                                    │
    ├─ delta = checkedOffset             ├─ offset = 0
    │          - inputPosition           │
    ▼                                    ▼
  load @ chars + (patternIndex-delta)  load @ chars + patternIndex
    │                                    │
    ├─ delta ≤ patternIndex → wrong char  └─► correct captured char
    └─ delta >  patternIndex → OOB read
                                (before buffer start)

In the diagram, delta is op.m_checkedOffset - term->inputPosition. That quantity is zero only when nothing follows the backreference in the pattern, because then the speculative check span ends where the term ends and m_checkedOffset equals inputPosition. This is precisely why /(.)\1/u matched correctly while /(.)\1c/u did not, and why the regression test pairs those two patterns as the contrast case. Add terms after the backreference and delta grows by the number of code units those terms consume — one per literal ASCII character, two per non-BMP escape. delta is therefore a direct function of the pattern text, which is to say it is chosen by whoever wrote the regexp.

The two outcomes in the left column of the diagram are separated by a single comparison. When delta <= patternIndex, the displaced address still lands inside the subject string, and the generated code merely compares the current input character against the wrong captured character — the incorrect-match symptom the commit message describes. When delta > patternIndex, the index underflows past zero and the load reads from memory preceding the string's character buffer. Nothing stops it. The checkInput guards that YARR emits constrain the subject cursor's forward travel; they say nothing about patternIndex, and readCharacter() does not re-derive a bound from its arguments:

// readCharacter(negativeCharacterOffset, resultReg, indexReg)
//   → load @ characters + (indexReg - negativeCharacterOffset) * charSize
// No comparison, no branch. The address is formed and the load is issued.

Reaching the underflow needs two things the pattern author controls together: enough terms after the backreference to push delta up, and a capture that begins near enough to the start of the subject that patternIndex stays below it. Both are ordinary pattern-and-input construction from script; nothing here requires exotic heap state. What comes back is not a value the script can read — the loaded code unit is consumed only by the equality comparison that decides whether the backreference matches. The observable is therefore a match/no-match answer about a byte in front of the string buffer: an oracle. Repeated over varying delta, that oracle reads out StringImpl header words and whatever heap contents sit adjacent, which is the shape of an ASLR defeat or the reconnaissance half of a later corruption chain. This weakens memory-safety confinement inside the WebContent process; it corrupts nothing, and any disclosure stays in the renderer's address space, so cross-process impact would still require a separate memory-corruption bug plus a separate sandbox escape. The fix restores the invariant that the pattern character is loaded from exactly patternIndex, in the coordinate space patternIndex actually lives in.

A pattern-controlled bias meant for the subject cursor was subtracted from an absolute capture cursor, letting a regexp match outcome depend on memory before the string buffer.

The correct code was already sitting three lines above: the non-Unicode branch passed 0 while the Unicode branch passed a bias expression. Paired specialised implementations of one logical operation — Unicode vs non-Unicode, 8-bit vs 16-bit, inlined vs helper-call — are a recurring source of YARR bugs precisely because the rarer branch drifts from the common one and is exercised by far fewer tests. Worth noting too: the security consequence is invisible in the commit message. The author frames it purely as "incorrectly failing to match", and only the address arithmetic reveals that the same wrong offset walks off the front of the buffer when the capture is near position zero. Correctness-framed YARR JIT offset fixes deserve a second read for exactly that reason.