Skip to content

lora: make lora_strength scale DoRA/BoRA magnitudes, not just the direction - #102

Open
Cortexelus wants to merge 1 commit into
mainfrom
lora-strength-scales-magnitude
Open

Cortexelus wants to merge 1 commit into
mainfrom
lora-strength-scales-magnitude

Conversation

@Cortexelus

Copy link
Copy Markdown
Collaborator

Problem

lora_strength scales the low-rank direction update, but a DoRA/BoRA magnitude vector replaces the base row/column norm outright. So as strength → 0 the direction converges to W0/‖W0‖ while the magnitude stays fully learned, leaving (mag/‖W0‖)·W0 — whereas strength == 0 short-circuits to exactly W0.

That is a discontinuity at zero. "Almost no LoRA" is not almost no change, and most of the slider's useful range does nothing.

Measurement

Rank-16 dora-rows adapter, ‖W − W0‖ / ‖W0‖ on the layer under test:

strength before after
0.0 0.00000% 0.00000%
1e-05 1.78760% 0.00000%
1e-03 1.78761% 0.00204%
0.01 1.78793% 0.04403%
0.1 1.82682% 0.41863%
0.5 2.59486% 2.08700%
1.0 4.16283% 4.16283%

Before: 43% of the total effect was already present at strength 1e-5, and the first 0.1 of the slider moved almost nothing. After: near-linear.

Fix

_blend_magnitude interpolates each learned magnitude toward the base weight's corresponding norm. Applied to all four magnitude-bearing adapters: dora, dora-xs, bora, bora-xs. The last two interpolate row and column magnitudes separately, since at strength 0 the row stage must reproduce W0 before the column stage can.

Compatibility

  • strength 1.0 is bit-identical to before (the table's last row), so trained adapters and any strength-1 inference are unaffected.
  • _blend_magnitude returns early at strength 1.0, so the common path — all of training — does not pay for the extra base-norm reduction.
  • Anyone currently using a strength slider with a DoRA adapter will see different output at strengths other than 0 and 1. That is the point of the change, but it is a behaviour change and worth noting in release notes.

…ection

lora_strength scaled the low-rank direction update but left the magnitude
vector replacing the base row/column norm outright. So as strength -> 0 the
direction converged to W0/||W0|| while the magnitude stayed fully learned,
leaving (mag/||W0||)*W0 — whereas strength == 0 short-circuited to exactly W0.
A discontinuity at zero: "almost no LoRA" was not almost no change.

Measured on a rank-16 dora-rows adapter (the sa3-medium seconds_total embedder),
||W - W0|| / ||W0||:

    strength      before      after
         0.0     0.00000%   0.00000%
       1e-05     1.78760%   0.00000%
       1e-03     1.78761%   0.00204%
        0.01     1.78793%   0.04403%
         0.1     1.82682%   0.41863%
         0.5     2.59486%   2.08700%
         1.0     4.16283%   4.16283%

Before, 43% of the total effect was already present at strength 1e-5 and the
first 0.1 of the slider did almost nothing. After, the response is near-linear
and strength 1 is bit-identical to the old behaviour, so trained adapters are
unaffected.

_blend_magnitude interpolates each learned magnitude toward the base weight's
corresponding norm. Applied to all four magnitude-bearing adapters: dora,
dora-xs, bora and bora-xs — the last two interpolate row and column magnitudes
separately, since at strength 0 the row stage must reproduce W0 before the
column stage can. It returns early at strength 1.0, so the common path (all of
training) does not pay for the extra base-norm reduction.
@Cortexelus

Copy link
Copy Markdown
Collaborator Author

Tested

Ran a full suite against this branch and against all three LoRA PRs merged together (they conflict in model.py — _blend_magnitude and _delta/baked_vnorm_row land at the same insertion point; both sides are purely additive, so the resolution is to keep both).

The discontinuity this fixes, measured on the one conditioner layer a real adapter touches (conditioners.seconds_total.embedder.embedding.1), as ‖W − W₀‖/‖W₀‖:

strength before after
0 0% 0%
1e-5 1.78760% 0.00000%
1.0 4.16283% 4.16283%

Full slider sweep (0 → 10, both a DiT layer and the conditioner layer):

0.0:0.00%  1e-05:0.00%  0.001:0.00%  0.1:0.42%  0.25:1.04%
0.5:2.09%  1.0:4.16%    2.0:8.27%    5.0:20.14%  10.0:37.81%

Asserted: exact base at strength 0, monotonic across the whole range, near-linear away from 0 (max/min slope ratio 1.006), amplification above 1.0 still usable to the slider's max of 10, round-trip back to 1.0 bit-exact, and strength 1.0 unchanged from the pre-PR value — so no existing output moves.

Also confirmed the wrapper's set_lora_strength dispatches to both model.model and model.conditioner, which is what makes the conditioner layer reachable by the slider at all.

54 assertions across the strength and gradio groups, all passing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant