Skip to content

ORE v2 (5/n): 6-bit block scheme + v2 wire format - #82

Open
coderdan wants to merge 26 commits into
feat/ore-v2-simdfrom
feat/ore-v2-bit6
Open

coderdan wants to merge 26 commits into
feat/ore-v2-simdfrom
feat/ore-v2-bit6

Conversation

@coderdan

@coderdan coderdan commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #81. Plan §4 + §6 (docs/plans/2026-06-12-ore-v2-architecture.md); crypto review brief docs/reviews/2026-06-14-ore-v2-crypto-review-brief.md.

Crypto review status ✅

The §6 H-construction gate (A1) is resolved (2026-06-15): the scheme uses the BHKR fixed-key-AES σ-MMO H(x, r) = LSB(π(σ(x) ⊕ r) ⊕ σ(x) ⊕ r), with π = AES-128_{K₀} (public K₀) and σ(x) = 2x in GF(2¹²⁸).

  • The known fixed-key-MMO attacks (GKWY; the half-gates multi-instance attack, eprint 2019/1168) require known, Free-XOR-correlated hash inputs and a recoverable global offset. ORE has independent secret PRF inputs and no global offset, so they do not port.
  • The tight tweak-as-key variant (2019/1168 Thm 2) is declined: rekeying per evaluation breaks the keyless-comparator / performance requirement and fixes a degradation ORE doesn't suffer.
  • The AES-hashing cryptanalysis (eprint 2025/792) targets collision/preimage/one-wayness — not the 1-bit correlation-robustness we rely on — and reaches only round-reduced AES (7/10 collision).
  • The orthomorphism σ is cheap defense-in-depth (security by matching the named BHKR/Zahur construction, not a usage argument). Full write-up: review brief A1.

Wire format is now frozen — Bit6 vectors pinned in tests/compat_w6_vectors (mirrors the Bit8 compat_vectors contract: pinned left + full bytes, comparison/order fixtures, plus a signed i64 order vector).

Also lands the constant-time PRP hardening (A4): LemireFyPrp (fixed-draw Fisher–Yates + Lemire reduction, replacing the rejection-sampled shuffle — closes a plaintext-dependent timing channel and a modulo bias), #[repr(C, align(64))] + an N ≤ 64 compile guard for the secret-indexed key-gen writes, and an oblivious comparator block read (ct_select_byte, closing the cache-line / MemJam channel). See review brief A4.

Numbers (Apple M1 Max, u64)

Bit8 (after #80) Bit6 (this PR)
encrypt ~25 µs ~8.9 µs
compare ~234 ns ~198 ns
ciphertext 408 B 295 B (incl. 4 B v2 header)

The σ-MMO's GF(2¹²⁸) doubling is vectorized to a single u128 word op (fused with the nonce XOR in the ~704-eval/u64 hot loop), so encrypt is back to the pre-orthomorphism shape-i level (~8.9 µs vs ~12.3 µs for the scalar byte-loop). Byte-identical, still constant-time.

What

  • OreAes128Bit6 / OreAes128Bit6ChaCha20 — 64-element domain, 64 RO evals/block, 8-byte right blocks, MSB-first 6-bit decomposition (order-preservation quickchecked).
  • v2 wire header (version ‖ scheme_id ‖ count) via OreCipher::WIRE_HEADER (None for legacy — Bit8 bytes unchanged, its vectors still pass). Cross-scheme / cross-shape comparisons return None; corrupt/truncated headers fail parsing.
  • H = BHKR σ-MMO (FixedPiZ2Hash) — A1 resolved.
  • PRP = LemireFyPrp (fixed-draw FY + Lemire reduction), constant-time hardened (A4).
  • Domain separation: block count bound into the PRF inputs; fixes the repeated-PRP-seed TODO for this scheme.
  • OreEncrypt for primitives ≤ 64 bits with compile-time block-count assertions; u128/i128/Decimal exceed the 14-block packed-prefix cap (documented; they arrive with the chained prefix, PR 6).
  • Wire vectors pinned (tests/compat_w6_vectors).
  • API change: OreEncrypt blanket impls (T: OreCipher) became scheme-specific — coherence requires it once block count ≠ byte count. Downstream code generic over OreCipher must name a scheme.

Not in this PR

  • chrono/decimal Bit6 impls (chrono fits trivially; Decimal can't).
  • The chained-prefix / variable-length scheme + string encryption (PR 6; behind the A2/A3 crypto gate).

@coderdan coderdan changed the title ORE v2 (5/n): 6-bit block scheme + v2 wire format — DO NOT MERGE before H review ORE v2 (5/n): 6-bit block scheme + v2 wire format Jun 15, 2026
coderdan added a commit that referenced this pull request Jun 15, 2026
Make explicit that the chained scheme has exactly one secret key (k),
producing both branches via the branch tag; branch-tag domain separation
under a good PRF is equivalent to independent per-branch keys but cheaper
(one key schedule + one subkey pair). Note H's pi is public (not secret) and
the nonce is not a key; contrast with fixed-N (#82, two keys) and the
init(k1,k2) API (k2 redundant here); flag the single-key choice for explicit
sign-off.
coderdan added a commit that referenced this pull request Jun 16, 2026
…ebug (#82 review)

Code-review follow-ups on the Bit6 / v2-wire PR:

- compare_raw_slices: reject a degenerate count=0 header. No OreEncrypt path
  produces zero blocks; without this, two crafted 0-block ciphertexts (header
  + nonce only) compare Equal because the scan loop never runs. (C2)
- Deduplicate encode_right_block: make the bit2 helper generic over the hash
  (encode_right_block<W: BlockWidth, H: Hash>) and have the Bit6 scheme call it
  instead of keeping a near-verbatim copy. (D1)
- Add width::ct_bit and route all four get_bit sites through it: extract the
  target bit with constant shift amounts + a constant-time select, instead of
   (a shift by a secret amount — constant-time on x86_64/
  aarch64 but not guaranteed on every target). Pairs with ct_select_byte for a
  fully oblivious, data-independent bit read. (D3 mitigation)
- Replace #[derive(Debug)] on OreAes128 / OreAes128Bit6 with an explicit opaque
  Debug impl (finish_non_exhaustive) so key material can never be rendered.
  (Used a manual impl rather than vitaminc::OpaqueDebug — vitaminc is a 0.2.0
  pre-release and ore-rs is a published crate; see PR discussion.)

No wire-format change: bit2 and bit6 compat vectors remain byte-identical.
coderdan added a commit that referenced this pull request Jun 16, 2026
…ebug (#82 review)

Code-review follow-ups on the Bit6 / v2-wire PR:

- compare_raw_slices: reject a degenerate count=0 header. No OreEncrypt path
  produces zero blocks; without this, two crafted 0-block ciphertexts (header
  + nonce only) compare Equal because the scan loop never runs. (C2)
- Deduplicate encode_right_block: make the bit2 helper generic over the hash
  (encode_right_block<W: BlockWidth, H: Hash>) and have the Bit6 scheme call it
  instead of keeping a near-verbatim copy. (D1)
- Add width::ct_bit and route all four get_bit sites through it: extract the
  target bit with constant shift amounts + a constant-time select, instead of
  "byte >> (bit % 8)" (a shift by a secret amount — constant-time on
  x86_64/aarch64 but not guaranteed on every target). Pairs with ct_select_byte
  for a fully oblivious, data-independent bit read. (D3 mitigation)
- Replace #[derive(Debug)] on OreAes128 / OreAes128Bit6 with an explicit opaque
  Debug impl (finish_non_exhaustive) so key material can never be rendered.
  (Used a manual impl rather than vitaminc::OpaqueDebug — vitaminc is a 0.2.0
  pre-release and ore-rs is a published crate; see PR discussion.)

No wire-format change: bit2 and bit6 compat vectors remain byte-identical.
@coderdan

Copy link
Copy Markdown
Contributor Author

Review: constant-time & zeroization

Ran the Trail of Bits constant-time-testing and zeroize-audit skills (plus a multi-agent code review) over this PR. Summary of the two security dimensions below; the code-review fixes are in 7458510.

Constant-time

The Bit6 path delivers the intended hardening (verified against the legacy bit2 baseline):

  • CT-1 (encrypt-timing via rejection sampling): fixed — LemireFyPrp is fixed-draw (Lemire multiply-high, no rejection loop, no %).
  • CT-4 (compare-side cache leak): fixed — all four get_bit sites read the right-block byte through the oblivious ct_select_byte.
  • CT-2 / CT-3 (secret-indexed permute lookup / FY swap): mitigated to a MemJam sub-cache-line residual by #[repr(C, align(64))] + N = 64 (one cache line) + const assert!(N <= 64); alignment confirmed empirically (size 128 / align 64).
  • New (this PR): the post-ct_select_byte bit extract used byte >> (bit % 8) — a shift by a secret amount, constant-time on x86_64/aarch64 but not guaranteed on every target. Now routed through a new width::ct_bit (constant shift amounts + constant-time select), so the bit read is data-independent on all targets.

No new timing channels introduced; the σ-MMO hash is branch-free over constant-time AES.

Zeroization

Clean — no new gaps. All new key-derived material zeroizes on drop, and the compiler/IR phase confirms the wipes survive -O2 (crate-wide store volatile rises 2→6→384 across O0/O1/O2; zero OPTIMIZED_AWAY):

  • LemireFyPrp — Zeroize + manual Drop + ZeroizeOnDrop; keystream wiped after the permutation is built; AES key schedule via aes's own ZeroizeOnDrop (survives O2 under --cfg aes_armv8).
  • OreAes128Bit6 — derive(ZeroizeOnDrop) over the PRF keys; SeedBuf + per-block template/work scratch all wiped.
  • FixedPiZ2Hash — correctly not flagged (PI_KEY is a fixed public key; the new param is a nonce).

The only finding is the pre-existing ZA-0001 (legacy Aes128Prng lacks ZeroizeOnDrop) — not introduced here; fixed on #84 and inherited on rebase. The benign #[derive(Debug)] on the cipher was replaced with an explicit opaque Debug impl in this PR.

Full write-ups: docs/reviews/2026-06-16-zeroize-audit-pr82.md and docs/reviews/2026-06-16-constant-time-analysis-pr79.md.

@coderdan

coderdan commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Added 4be9e6b, which restores the #81 self-review hardening (951c7cb) that e164730 overwrote when this branch was rebased over #81. That change predates today's rebases; it was already in the previously pushed 92833dd.

  • lsb_mask_256 uses real assert_eq! again. The NEON path gathers blocks through raw pointers assuming exactly 256 of them, so with only a debug_assert a shorter slice reads out of bounds in release, from a safe fn. Current callers check the length first, so nothing triggered it.
  • The scalar fallbacks sit under cfg(not(target_arch = "aarch64")) again instead of #[allow(unreachable_code)], which Rust 1.78 does not honour on a statement (clippy failed on aarch64). The new gt_mask_xor_64 gets the same shape.

Also on #79: 08a308e derives BlockWidth::DOMAIN from BITS (stable clippy flagged BITS as never read), so this PR's Bit6 drops its explicit DOMAIN. #80, #81 and #83 were rebased over both with no conflicts. fmt, clippy (stable and 1.78, aarch64 and x86_64) and tests (stable and 1.78) pass at every tip.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

PRF input collisions, malformed-symbol acceptance, and inaccurate security premises remain unresolved.

6 open findings
What changed in this PR

Adds an opt-in 6-bit ORE scheme with smaller ciphertexts and less per-block AES work, while preserving legacy Bit8 wire bytes.

Changes:

  • Introduces Bit6 encryption, v2 headers, and pinned compatibility vectors.
  • Adds σ-MMO hashing, fixed-draw permutations, and oblivious comparator reads.
  • Documents security decisions, benchmarks, and the future CMAC accumulator.
File Description
packages/​ore-rs/​tests/​compat_w6_vectors/​vectors.rs Pins Bit6 ciphertext bytes.
packages/​ore-rs/​tests/​compat_w6_vectors.rs Tests wire compatibility and ordering.
packages/​ore-rs/​src/​scheme/​width.rs Adds Bit6 types and oblivious reads.
packages/​ore-rs/​src/​scheme/​bit2/​block_types.rs Hardens legacy bit access.
packages/​ore-rs/​src/​scheme/​bit2.rs Shares encoding and hardens comparisons.
packages/​ore-rs/​src/​scheme/​bit2_w6/​block_types.rs Defines eight-byte right blocks.
packages/​ore-rs/​src/​scheme/​bit2_w6.rs Implements Bit6 encryption and comparison.
packages/​ore-rs/​src/​scheme.rs Exposes the Bit6 module.
packages/​ore-rs/​src/​primitives/​simd.rs Adds 64-lane indicator processing.
packages/​ore-rs/​src/​primitives/​prp.rs Adds fixed-draw Fisher–Yates permutations.
packages/​ore-rs/​src/​primitives/​hash.rs Implements fixed-key σ-MMO hashing.
packages/​ore-rs/​src/​lib.rs Adds scheme-specific wire-header support.
packages/​ore-rs/​src/​encrypt.rs Makes legacy encryption implementations scheme-specific.
packages/​ore-rs/​src/​decimal.rs Keeps Decimal encryption on Bit8.
packages/​ore-rs/​src/​ciphertext.rs Adds header-aware serialization and parsing.
packages/​ore-rs/​src/​chrono.rs Keeps date encryption on Bit8.
packages/​ore-rs/​Cargo.toml Registers Bit6 benchmarks.
packages/​ore-rs/​benches/​bit6.rs Benchmarks Bit6 encryption and comparison.
docs/​reviews/​2026-06-14-ore-v2-crypto-review-brief.md Records cryptographic decisions and review gates.
docs/​plans/​2026-06-15-ore-v2-cmac-accumulator-spec.md Specifies the future chained-prefix design.
docs/​plans/​2026-06-12-ore-v2-architecture.md Updates architecture and leakage guidance.
docs/​benchmarks/​2026-06-13-bit6-prp-results.md Records permutation benchmark results.

🧠 Review effort: Balanced


Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.

Comment thread packages/ore-rs/src/scheme/bit2_w6.rs
Comment thread packages/ore-rs/src/scheme/bit2_w6.rs Outdated
Comment thread packages/ore-rs/src/ciphertext.rs
Comment thread packages/ore-rs/src/scheme/bit2_w6.rs
Comment thread docs/reviews/2026-06-14-ore-v2-crypto-review-brief.md Outdated
Comment thread docs/reviews/2026-06-14-ore-v2-crypto-review-brief.md Outdated
@coderdan
coderdan requested a review from freshtonic October 8, 2026 23:41
coderdan added a commit that referenced this pull request Oct 9, 2026
…ator's block

Review of #82 (A4): permute reads inverse[x[n]] with the plaintext
symbol, and invert reads permutation[j]. The tables' alignment defends
only cache-line granularity; sub-line timing can reveal which part of
the table was read, which for permute is plaintext bits. Both PRPs now
read every entry and select the wanted one in constant time
(ct_select_byte). Encryption calls permute once per block, so the cost
is up to 256 selects per block against ~10 µs of PRP setup.

The Bit6 comparator had the compare-side twin of this. Its prefix scan
is constant-time, but the resolution step then loaded a[l], f[l] and
right[l] by the latched index l. The left blocks are cache-resident
after the scan whatever l is; the right blocks are never touched by it,
so the right[l] load hits or misses according to l, and a dudect run
measured that as a timing signal (2 % of the spread on a 23-block
chained ciphertext, at the edge of detection on Bit6's five lines).
The scan now latches all three under the choice "this is the first
differing block" (ct_assign_bytes), so no load after it is indexed by
l. One extra 8-byte read and about 25 masked byte copies per block.
The brief's compare-side section records the measurement and the fix.

Same output, so no ciphertext byte changes; the legacy and Bit6
vectors pass unchanged.
coderdan added a commit that referenced this pull request Oct 9, 2026
Review of #82, two findings that change Bit6 bytes, done together so the
vectors are regenerated once (Bit6 is not frozen; legacy bytes are
unchanged):

- Bit6 ran PRF1/PRF2 under the caller's raw keys, which the legacy
  scheme also uses, and the count in byte 15 did not separate the
  inputs: a legacy N=15 left ciphertext can reproduce the PRF1 input of
  any Bit6 masking key, so under shared keys a legacy query acts as a
  Bit6 query token. Bit6 now keys PRF1/PRF2 with AES_k1("ORE.v2.bit6.prf1")
  and AES_k2("ORE.v2.bit6.prf2"). The labels' last byte (ASCII '1'/'2')
  is outside every legacy PRF input (byte 15 is 0..=14 for PRF1, 0 for
  PRF2), so no legacy ciphertext publishes a derived key.
- Seeds bound only the block count, so a zero-padded prefix collided with
  a shorter one and every position of an all-zero value shared one
  permutation (the pinned zero vector had eleven identical xt bytes).
  Seeds now carry the block index in byte 14, free because N <= 14.

Regression tests: a search over legacy N=15 encryptions for a Bit6 tag
under shared keys, and distinct symbols across an all-zero value's
positions. Both fail on the previous construction. The plan's section 4
claim that byte 15 separates the schemes is corrected.
coderdan added a commit that referenced this pull request Oct 9, 2026
…ad paths

Review of #82:
- A1 justified the σ-MMO hash by saying H's inputs are secret and never
  learnt. They are not: every left tag f[n] is the RO key at xt[n], and
  the comparator hashes it. Restated in the brief, hash.rs and the plan:
  unrevealed RO keys are secrets shared across ciphertexts with public
  nonces as offsets, so GKWY's multi-instance shape applies, and security
  rests on the BHKR bound (about p*C/2^128) rather than on input secrecy.
  The conclusion holds for realistic parameters. Marked pending sign-off.
- A4 said permute/invert reads were already constant-time. permute read
  inverse[x[n]] by plaintext symbol; both are now oblivious scans (the
  previous commit), and A4 lists the reads with the writes for the
  high-assurance tier.
@coderdan

coderdan commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Force-pushed (rebase, with lease): the Bit6 comparator now latches the first differing block inside its constant-time scan.

A dudect run on the stacked verification PR (#100) found that the resolution step after the scan loaded a[l], f[l] and right[l] by the latched index l. The scan never touches the right blocks, so the right[l] load hit or missed the cache according to l, which a timing-only attacker can read. On Bit6 the signal was at the edge of detection (t ≈ 10 over 151 M samples, tau 0.0008, the whole ciphertext spans five cache lines); on the chained scheme it was 2 % of the timing spread.

The scan now copies the three values under the choice "this is the first differing block" (width::ct_assign_bytes, set for exactly one block), and the resolution step works from the copies, so no load after the scan is indexed by l. Folded into e033aad (harden(prp): read permute/invert tables obliviously; latch the comparator's block) as the compare-side twin of that commit's oblivious table reads. The brief's compare-side section, which had accepted the direct index on the grounds that l is leaked to the ciphertext holder anyway, now records the measurement and the fix.

No ciphertext bytes change; the Bit6 vectors pass unchanged. Clippy clean on stable, 1.78, aarch64 and x86_64-apple-darwin.

Replaces Bit6's rejection-sampled Knuth shuffle with LemireFyPrp: a
Fisher-Yates shuffle driven by a fixed count of wide (64-bit) draws,
each reduced to range by Lemire multiply-high ((x * range) >> 64).

Why it matters beyond speed: the Knuth shuffle's rejection sampling does
a seed-dependent number of draws, and the seed is PRF(plaintext prefix),
so PRP construction time was weakly plaintext-dependent — a timing side
channel. Fixed-count draws make construction time seed-independent and
branch-free, closing it. Uniformity is provable: each Lemire reduction is
within range/2^64 of uniform, so the permutation is within < 2^-55 of a
uniformly random permutation (the object Lewi-Wu models) — a pure
statistical term, no new assumption.

Scope:
- Bit6 only. Bit8 (legacy) stays on the Knuth shuffle: its byte-exact
  output is wire-frozen, so the timing channel there is documented in the
  plan rather than fixed (a fix would change ciphertext bytes).
- Seed-keyed shape (i): drops into the existing Prp::new(seed) with no
  architectural change. Bit6 u64 encrypt 11.5us -> 8.6us.
- The remaining win to ~3.3us (pre-scheduled stream, shape ii — no
  per-block AES key schedule) is deferred to PR 6, where the CMAC
  accumulator can emit the PRP keystream as a branch family under the same
  crypto review. Documented in the plan + bench doc.

Tests: permutation validity + bidirectional round-trip over 32 seeds,
determinism, short-key rejection, and indicator-mask-vs-reference
quickcheck (guards the gt_mask_xor_64 kernel). Plan open-question 1 and
bench results updated.

Part of the ORE v2 program (docs/plans/2026-06-12-ore-v2-architecture.md, PR 5).
Draft analysis for the CipherStash research blog covering the
rejection-sampling timing channel in PRP generation, the wide-draw/Lemire
fix and its security framing (2^-55 statistical distance vs Lewi-Wu's
random-permutation model), the swap-or-not rejection on proof grounds, and
the hardware-AES build-flag finding. Cross-references the vitaminc issue.
Restructured to the CipherStash standard (cipherstash-js-suite/prompts/
_shared/writing-guidelines.md): definition -> why it matters -> how it
works -> example -> related, with you/we voice, Note/Tip/Warning callouts,
a meta description, title options, and a short 'Why this matters for
CipherStash' framing. Genre-adapted for a crypto-internals post (Rust
snippets rather than the TS default; no sales CTAs).
…profile)

Redone against cipherstash-js-suite .claude/skills/blog-writing-voice
(the canonical voice skill, which post-dated the stale local checkout I
first wrote to). Narrative detective-story arc, hook opening instead of a
definition, recurring 'old code I was proud of' motif bookended, first-
person Dan voice, evocative headers, sparing em-dashes, US English, italic
closing aphorism, and the :wq sign-off.
The post now lives in the marketing site content
(cipherstash/cipherstash-js-suite#548); it doesn't belong in the library
repo.
Lewi-Wu is a left/right scheme: a comparison is only evaluated between a
left and a right ciphertext, and a right ciphertext in isolation reveals
nothing about order. With right-only-at-rest storage (the default
deployment), an offline attacker recovers nothing -- not order, and a
fortiori not common-prefix length. The first-differing-block disclosure
surfaces only at query time (to the legitimate operator) or to an online
adversary observing query traffic.

Rewrite the plan's string-leakage discussion to lead with this left/right
asymmetry and the three threat tiers, and narrow the security-checklist
item accordingly. The product decision (B1) is thus about acceptable
query-time/online leakage, not at-rest leakage.
Standalone sign-off brief a reviewer can act on without reading the full
plan or codebase. Covers: A1 (1-bit hash H instantiation -- shipped
FixedPiZ2Hash), A2 (CMAC cached-state accumulator for PR 6), A3 (PRP
keystream as a CMAC branch family, shape ii), A4 (secret-indexed
Fisher-Yates swap).

Sequencing: A1+A4 gate Bit6 vector pinning (both change ciphertexts),
A2+A3 gate PR 6.

Flags a discrepancy found in the shipped code: LemireFyPrp is not
repr(align(64)), so the one-cache-line argument the plan claims for the
secret-indexed permutation.swap(i,j) does not hold as written -- either
add the alignment or take the constant-time fallback.
A4 hardening from the crypto review brief. None of these change ciphertexts
(compat + comparison vectors unchanged).

- prp.rs: LemireFyPrp is now repr(C, align(64)) so each [u8;N] table is one
  64-byte cache line (permutation@0, inverse@64) -- covers both secret-indexed
  key-gen writes (the FY swap and the inverse fill) for the one-cache-line
  argument. Add a const assert!(domain <= 64) in impl_lemire_fy_prp! so a
  larger instantiation, which would span multiple lines and lose the property,
  is a compile error.

- width.rs: add oblivious ct_select_byte(block, idx) -- scans the whole block
  and constant-time-selects the byte, so the access address is independent of
  the secret index.

- bit2/bit2_w6 comparators: route all four get_bit sites through
  ct_select_byte. The right-block byte read was indexed by a[l] (the secret
  permuted symbol); the oblivious read closes that cache-line channel and, by
  touching every byte, the sub-line (MemJam) channel too -- chosen over mere
  alignment for that reason. Cost is <=32 byte-ops per comparison.

- review brief: add the MemJam analysis (4K-aliasing, 4-byte granularity,
  Intel-wide scope incl. SGX, ARM/AMD out, SMT co-residence required),
  the oblivious-swap-FY vs swap-or-not wire-compatibility distinction, and the
  consequence that A4 no longer gates Bit6 vector pinning (the fixes are
  byte-stable; A1/H is the sole remaining gate).
Add a "Block width is a leakage decision" subsection to plan section 5(b)
and rewrite open question 3.

Lewi-Wu leaks the first-differing-block index, so larger blocks leak less
(Bit8 < Bit6 < CLWW): a u64's first differing bit is localised to an 8-bit
window at Bit8 vs a 6-bit window at Bit6. That online prefix-leakage axis is
the one the library cannot fix; it pulls against Bit6's encrypt-side
one-cache-line CT advantage, which is only a cost difference (full oblivious
CT is available at both widths, ~12x cheaper at Bit6).

Conclusion: width is a per-domain/per-deployment policy keyed on target data
and threat model, not a global default. Default numerics to Bit8 (lower
leakage + wire-compat); Bit6 is opt-in for at-rest-dominated, size/perf, or
encryptor-hostile deployments. Supersedes the earlier 'Bit6 as default' lean.
Record mantissa+exponent / log-domain encoding as an encoding-layer option
for wide-dynamic-range numerics (currency, measurements, high-range decimals):
bounds block count, uniform relative precision, deliberate magnitude-band
leakage; composes with the variable-block machinery. Caveats: Benford
leading-digit skew (relative != flat), magnitude-band is a conscious leak, and
parameters must be fixed per-domain as public (Parameter-Hiding ORE, Cash et
al. 2018).

Explicitly scope it as NOT a low-entropy / narrow-domain mitigation: it is
order-preserving, so it cannot touch the order floor, and for a narrow domain
like DOB the exponent is near-constant (increases high-order skew). DOB-class
fields are mitigated by coarsening the plaintext to the queried granularity,
not by re-encoding. Keeps the two ideas from being conflated later.
Keep the fixed public-key AES construction (option 3) but upgrade it with the
BHKR/Zahur orthomorphism:

  H(x, r) = LSB( pi(sigma(x) XOR r) XOR sigma(x) XOR r )

with pi = AES-128 under public key K0 and sigma(x) = 2x in GF(2^128) (the
'multiply by x' / CMAC-subkey doubling; reduction constant 0x87). FixedPiZ2Hash
(Bit6 only) updated in both the scalar comparator path and the bulk
scalar/SIMD encryption path; added gf128_double plus tests
(fixed_pi_scalar_matches_bulk, gf128_double_reduction). Legacy Bit8 keeps
Aes128Z2Hash unchanged.

Justification (review brief A1, plan section 6):
- The known fixed-key-MMO attacks (GKWY; the half-gates multi-instance attack
  of eprint 2019/1168) require *known* hash inputs and a recoverable global
  Free-XOR offset. ORE has independent *secret* PRF inputs and no global
  offset, so the O(p*C/2^k) degradation does not arise.
- The tight tweak-as-key variant (2019/1168 Thm 2) is declined: rekeying per
  evaluation breaks the keyless-comparator / performance requirement and fixes
  a degradation ORE does not suffer.
- eprint 2025/792 cryptanalyses collision/preimage/one-wayness (not the 1-bit
  correlation-robustness we rely on) and only round-reduced AES (7/10
  collision on AES-MMO/MP), leaving full AES-128's margin intact.
- The orthomorphism is cheap defense-in-depth: security holds by matching the
  named BHKR/Zahur construction rather than by a usage argument.

This changes Bit6 right ciphertexts; A1 was the gate holding Bit6 vector
pinning and is now cleared. Docs mark A1 RESOLVED.
Freeze the OreAes128Bit6 (bit2_w6) wire format now that A1 (the H
construction) is resolved — BHKR σ-MMO. Mirrors compat_vectors.rs (Bit8):
deterministic TestRng nonce, pinned left + full ciphertext bytes, and
comparison/order fixtures over the pinned bytes. Adds a signed-int (i64)
order vector since Bit6 is a new scheme; cross-checks confirm the v2 header
(0x02/0x02 + u16 block count) and the orderable sign-flip equivalence
(i64::MIN ↔ 0u64, i64::MAX ↔ u64::MAX).

Regeneratable via:
  cargo test --test compat_w6_vectors -- --ignored --nocapture generate
Replace the per-byte carry loop in the BHKR sigma-MMO with a single u128
word op: 2x = (x << 1) ^ ((x >> 127) * 0x87) over the big-endian field
element. In the bulk encrypt path, fold the doubling and the nonce XOR into
one u128 pass per block (the ~704-evals/u64 hot loop).

Byte-identical output (sigma is unchanged), so the pinned Bit6 vectors and
the scalar<->bulk consistency test still pass; still constant-time
(branch-free, no secret-dependent control flow).

Recovers the A1 sigma-MMO regression: Bit6 u64 encrypt ~12.3us -> ~8.9us
(Apple M1 Max), back to the pre-orthomorphism shape-i level.
The detailed, review-ready spec for the PR 6 accumulator (the A2 gate),
replacing the sketch in the review brief. Covers:

- Construction: CMAC (NIST SP 800-38B) over an injective prefix-block +
  final-block encoding; cached CBC state = incremental CMAC. Subkeys reuse
  the (vectorized) gf128_double.
- Injective message encoding (byte layouts) + injectivity argument.
- Per-block algorithm; the left tag f[n] = ro(n, xt[n]) is the RO_KEY branch
  at the permuted symbol (not a separate output), preserving mask
  cancellation at compare time.
- NO total-length binding -- required so prefix-sharing strings of different
  lengths compare correctly; length comparability enforced at the comparator.
- One dedicated KDF'd key unifies the old prf1/prf2 via branch tags
  (RO_KEY/PRP_STREAM); PRF2 subsumed.
- Shape-(ii) PRP (A3): keystream derived as CMAC tags, no per-block key
  schedule (needs a LemireFyPrp::from_stream ctor).
- Security: reduction to CMAC PRF + the three auditable claims (injectivity,
  incremental faithfulness, zeroization) + birthday budget; the chain state
  is never published, so the cascade/GGM interaction does not arise.
- Test plan + open questions for the reviewer.

Linked from review brief A2 and plan section 5(b).
The final-block signature F(branch, n, s) and section 3's 'branch
separation' referenced 'branch' before it was named. Add an explicit
definition at the top of section 4 (RO_KEY / PRP_STREAM output families,
carried as the byte-0 branch tag) and point section 3 at it.
Make explicit that the chained scheme has exactly one secret key (k),
producing both branches via the branch tag; branch-tag domain separation
under a good PRF is equivalent to independent per-branch keys but cheaper
(one key schedule + one subkey pair). Note H's pi is public (not secret) and
the nonce is not a key; contrast with fixed-N (#82, two keys) and the
init(k1,k2) API (k2 redundant here); flag the single-key choice for explicit
sign-off.
…ebug (#82 review)

Code-review follow-ups on the Bit6 / v2-wire PR:

- compare_raw_slices: reject a degenerate count=0 header. No OreEncrypt path
  produces zero blocks; without this, two crafted 0-block ciphertexts (header
  + nonce only) compare Equal because the scan loop never runs. (C2)
- Deduplicate encode_right_block: make the bit2 helper generic over the hash
  (encode_right_block<W: BlockWidth, H: Hash>) and have the Bit6 scheme call it
  instead of keeping a near-verbatim copy. (D1)
- Add width::ct_bit and route all four get_bit sites through it: extract the
  target bit with constant shift amounts + a constant-time select, instead of
  "byte >> (bit % 8)" (a shift by a secret amount — constant-time on
  x86_64/aarch64 but not guaranteed on every target). Pairs with ct_select_byte
  for a fully oblivious, data-independent bit read. (D3 mitigation)
- Replace #[derive(Debug)] on OreAes128 / OreAes128Bit6 with an explicit opaque
  Debug impl (finish_non_exhaustive) so key material can never be rendered.
  (Used a manual impl rather than vitaminc::OpaqueDebug — vitaminc is a 0.2.0
  pre-release and ore-rs is a published crate; see PR discussion.)

No wire-format change: bit2 and bit6 compat vectors remain byte-identical.
e164730 (13 June) predates the #81 self-review commit 951c7cb (16
June). When this branch was rebased over it, e164730's version of
simd.rs won, silently undoing two of 951c7cb's fixes:

- lsb_mask_256's length checks went back to debug_assert. The NEON path
  gathers blocks through raw pointers assuming exactly 256 of them, so a
  shorter slice reads out of bounds in a release build: UB reachable
  from a safe fn. Today's callers check the length first, but the fn
  was unsound. They are real asserts again.
- The scalar fallbacks went back behind #[allow(unreachable_code)],
  which Rust 1.78 does not honour on a statement, so clippy failed on
  aarch64. Every dispatcher, including the new gt_mask_xor_64, now ends
  in the NEON call on aarch64 and puts the AVX2 check and scalar
  fallback under cfg(not(aarch64)); scalar::gt_mask_xor is compiled only
  where it is used, as 951c7cb had it.

Bit6 also drops its explicit DOMAIN, now derived from BITS.
Fisher-Yates swapped at the draw-derived index and filled the inverse at
the permutation's values. Both offsets stay inside one 64-byte line, so
the cache cannot see them, but dudect on Apple M4 measured the build time
depending on which permutation is built, with no co-tenant: the swaps
carry the signal and the inverse fill alone shows none. That fits
store-to-load forwarding and memory disambiguation on byte-granular
accesses at secret offsets, which the one-cache-line argument does not
cover.

The builder now keeps both tables in registers and never uses a secret
value as an address: perm[j] is fetched by a table lookup on a broadcast
j and written back by compare-and-select, and the inverse is kept by
exchanging the values i and j in it at each step, which needs neither
swapped entry. Step i is public and j <= i, so only the chunks covering
0..=i are searched. Same 63 Lemire draws and swap sequence, so the
tables, and every ciphertext, are byte-identical; the Bit6 vectors pass
unchanged.

Dispatch: NEON on aarch64 (baseline, no runtime check), SSSE3 (pshufb)
on x86_64 when the CPU reports it, a u64 SWAR form with exact
byte-equality masks everywhere else. The textbook builder survives only
as a test-only reference, with a byte-at-a-time subtle_ng form; property
and edge-stream tests pin every builder's tables to it.

Each draw is decoded from an 8-byte copy of the keystream into a u64;
both copies are wiped, as the caller wipes the stream itself. Raised in
the review of #90 against the #82 zeroize audit.
…ength

Review of #82:
- The shared parser copied xt after checking only lengths, so a Bit6
  ciphertext with an xt byte of 64..=255 parsed; comparing it then read
  the right block at an invalid index (a debug panic, a silent zero in
  release). OreCipher gains SYMBOL_DOMAIN (256 by default, 64 for Bit6),
  and Left/CipherText parsing rejects any xt outside it.
- The Bit6 raw comparator now returns None for such symbols in either
  input before scanning.
- gt_mask_xor_64 asserts its output length in release, as
  gt_mask_xor_256 does.

Regression tests patch a pinned vector's first xt byte to 64; both fail
on the previous code. No ciphertext byte changes.
…ator's block

Review of #82 (A4): permute reads inverse[x[n]] with the plaintext
symbol, and invert reads permutation[j]. The tables' alignment defends
only cache-line granularity; sub-line timing can reveal which part of
the table was read, which for permute is plaintext bits. Both PRPs now
read every entry and select the wanted one in constant time
(ct_select_byte). Encryption calls permute once per block, so the cost
is up to 256 selects per block against ~10 µs of PRP setup.

The Bit6 comparator had the compare-side twin of this. Its prefix scan
is constant-time, but the resolution step then loaded a[l], f[l] and
right[l] by the latched index l. The left blocks are cache-resident
after the scan whatever l is; the right blocks are never touched by it,
so the right[l] load hits or misses according to l, and a dudect run
measured that as a timing signal (2 % of the spread on a 23-block
chained ciphertext, at the edge of detection on Bit6's five lines).
The scan now latches all three under the choice "this is the first
differing block" (ct_assign_bytes), so no load after it is indexed by
l. One extra 8-byte read and about 25 masked byte copies per block.
The brief's compare-side section records the measurement and the fix.

Same output, so no ciphertext byte changes; the legacy and Bit6
vectors pass unchanged.
Review of #82, two findings that change Bit6 bytes, done together so the
vectors are regenerated once (Bit6 is not frozen; legacy bytes are
unchanged):

- Bit6 ran PRF1/PRF2 under the caller's raw keys, which the legacy
  scheme also uses, and the count in byte 15 did not separate the
  inputs: a legacy N=15 left ciphertext can reproduce the PRF1 input of
  any Bit6 masking key, so under shared keys a legacy query acts as a
  Bit6 query token. Bit6 now keys PRF1/PRF2 with AES_k1("ORE.v2.bit6.prf1")
  and AES_k2("ORE.v2.bit6.prf2"). The labels' last byte (ASCII '1'/'2')
  is outside every legacy PRF input (byte 15 is 0..=14 for PRF1, 0 for
  PRF2), so no legacy ciphertext publishes a derived key.
- Seeds bound only the block count, so a zero-padded prefix collided with
  a shorter one and every position of an all-zero value shared one
  permutation (the pinned zero vector had eleven identical xt bytes).
  Seeds now carry the block index in byte 14, free because N <= 14.

Regression tests: a search over legacy N=15 encryptions for a Bit6 tag
under shared keys, and distinct symbols across an all-zero value's
positions. Both fail on the previous construction. The plan's section 4
claim that byte 15 separates the schemes is corrected.
coderdan added a commit that referenced this pull request Oct 9, 2026
- The MMO validation claimed a pl/pgSQL comparator was validated; only
  the σ-MMO hash was implemented and checked (12 vectors). The doc now
  scopes its result to the hash and lists what a comparator still needs
  end-to-end vectors for.
- "LSB" in both the SQL and Rust is bit 0 of byte 0 of the block (bit
  120 of the big-endian field element), as in the frozen legacy hash.
  The SQL and summary now name that coordinate.
- The #82 and #83 zeroize audits marked items clean that left
  key-derived temporaries unwiped (from_stream's per-draw copy; the CMAC
  mixed values). Both audits are amended, pointing at the fixes on #82
  and #83, and say plainly that the wipes are source-level only.
coderdan added a commit that referenced this pull request Oct 9, 2026
…ad paths

Review of #82:
- A1 justified the σ-MMO hash by saying H's inputs are secret and never
  learnt. They are not: every left tag f[n] is the RO key at xt[n], and
  the comparator hashes it. Restated in the brief, hash.rs and the plan:
  unrevealed RO keys are secrets shared across ciphertexts with public
  nonces as offsets, so GKWY's multi-instance shape applies, and security
  rests on the BHKR bound (about p*C/2^128) rather than on input secrecy.
  The conclusion holds for realistic parameters. Marked pending sign-off.
- A4 said permute/invert reads were already constant-time. permute read
  inverse[x[n]] by plaintext symbol; both are now oblivious scans (the
  previous commit), and A4 lists the reads with the writes for the
  high-assurance tier.
…ad paths

Review of #82:
- A1 justified the σ-MMO hash by saying H's inputs are secret and never
  learnt. They are not: every left tag f[n] is the RO key at xt[n], and
  the comparator hashes it. Restated in the brief, hash.rs and the plan:
  unrevealed RO keys are secrets shared across ciphertexts with public
  nonces as offsets, so GKWY's multi-instance shape applies, and security
  rests on the BHKR bound (about p*C/2^128) rather than on input secrecy.
  The conclusion holds for realistic parameters. Marked pending sign-off.
- A4 said permute/invert reads were already constant-time. permute read
  inverse[x[n]] by plaintext symbol; both are now oblivious scans (the
  previous commit), and A4 lists the reads with the writes for the
  high-assurance tier.
…om-key benchmarked and rejected

A dudect run (2026-10-09) measured the Bit6 PRP builder's time depending
on which permutation it builds, on a core with no co-tenant, and isolated
it to the Fisher-Yates swaps; the inverse fill showed none. That fits
store-to-load forwarding on byte-granular accesses at secret offsets
within one L1-resident line, which the one-cache-line argument does not
cover. A4 now records it: the oblivious builder (NEON / SSSE3 / SWAR,
byte-identical tables) is the default on every target, and on aarch64 it
is faster than the indexed builder it replaced, so the high-assurance
tier, the MemJam posture call and the alignment-only default are marked
superseded rather than deleted. The architecture plan's tier discussion
and the Bit6 benchmark note (new addendum, mains-power numbers) follow.

Sort-by-random-key (vitaminc's construction) was prototyped inside
LemireFyPrp::from_stream and benchmarked as a candidate for the tier on
2026-10-08: +135% to +215% on the builder, +34% to +102% on Bit6 encrypt
(+26% to +98% on chained), and it changes every ciphertext. Rejected; it
stays rejected now that the FY form is oblivious, since it removes the
secret address by the same means, costs more and changes the wire format.
Its no-retry rule is adopted for every builder: a bad draw is resolved by a
fixed rule, never by consuming more stream, because left and right halves
are produced in separate calls and must agree.
@coderdan

coderdan commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Force-pushed (with lease), docs only: the restated A1 argument is now recorded as signed off (2026-10-09) in the review brief (§A1 note and the Gate 1 checklist) and in the architecture plan's three "pending sign-off" references. Folded into 5630a93 (docs(review): restate A1 on a correct exposure model). No code change; #83 and #100 rebased on top.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants