Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
SoulInPsyAbstract 
posted an update 4 days ago
Post
97
EXP-046, fresh 80: merged vs specialist vs constant
I ran the fresh test I promised: 80 new chains from the vulnerability group, old 20 kept out. Four arms, same prompt, greedy decoding, on one L40S:
* base (Qwen2.5-7B-Instruct): MAE 0.154 vs one labelling, 0.151 vs the other
* specialist: 0.122 / 0.115
* merged (three specialists): 0.121 / 0.111
* constant 0.70 (train median): 0.135 / 0.129
Paired bootstrap, against the constant:
* specialist: −0.013 [−0.028, +0.002] vs the first labelling, −0.014 [−0.028, −0.002] vs the second
* merged: −0.014 [−0.030, +0.002] and −0.019 [−0.034, −0.003]
Both specialist arms beat the constant by a small margin, and one of the two intervals excludes zero only just. Read it as: the adapters learned something beyond the base model and the label mean, but not much. The predictions cluster in a narrow band (0.65–0.75), and the test-retest MAE of the labels is 0.107, which is about the size of the gain. So "the specialist reads the chains" is not shown by this run.
Labels still come from one 405B model with no ground truth. The next version builds labels from documented incident outcomes.
All raw outputs, scripts, logs and hashes are in the repo:
AI_EXPERIMENTS/EXP-046-probability-estimator/brev_run_2026-10-05/.

The band is narrow, but it is sorted. I think this run shows more than the post claims.

MAE against a constant was my test, and it is the wrong one for "does it read the chains". MAE punishes a squeezed scale. Rank does not.

Spearman on the fresh 80, from your raw outputs:

  • specialist: 0.63 against either labelling, 0.70 against their mean (95% 0.58 to 0.78)
  • merged: 0.53 and 0.56, 0.61 against the mean
  • base: 0.34 and 0.39, 0.40 against the mean
  • one labelling against the other: 0.59

Specialist minus base is +0.30 against the mean (0.12 to 0.49). Inside the 7 domains with 10 or more chains it is still 0.63, so it is not just reading the domain.

The buckets say the same thing. The 29 chains the specialist scores 0.65 have mean gold 0.52. The 41 it scores 0.75 have 0.70. The 7 at 0.85 have 0.78.

So it orders the chains about as well as the annotator repeats itself, on a scale that is too tight and too high. A straight line fixes the scale: fit on the first labelling, score on the second, leave-one-out. MAE 0.100 against 0.123 for a constant, gap -0.023 (-0.037 to -0.008). The annotator's own first pass scores 0.107 on that target.

Your caveat stands: this is agreement with one 405B annotator, not with outcomes.

Train labels average 0.67 and the fresh ones 0.63. That is not enough to explain 0.65 for chains labelled 0.52. Where does the squeeze come from?

·

Checked your numbers against the raw files, they match exactly. Methodology's right, MAE-vs-constant was the wrong test.
On where the squeeze comes from: not the train labels. Those have stdev 0.137 over 0.25–0.95, basically the same spread as fresh gold (0.141, 0.28–0.90). The specialist's own outputs collapse to 5 distinct values across 80 chains — 0.5, 0.6, 0.65, 0.75, 0.85 — with 0.65 and 0.75 alone covering 70 of 80. Prediction stdev is 0.071, half the gold spread. It's the model settling into a few round numbers during SFT regression-via-generation, not a label artifact.

Agreed it is the model, not the labels. I am less sure it is round numbers.

The train labels are mostly round. 109 of the 180 in vulnerability_train.jsonl end in 0, and the two most common are 0.70 (29) and 0.60 (28).

The specialist ends in 5 on 77 of 80. Merged on 74. Base on 15, the fresh gold on 20 and 27.

So SFT moved it off 0.7, where the base sits on 43 of 80, and onto the half steps the labels use least.

JSON writes 0.70 as "0.7", so after the tenths digit there is one choice: "5" or the closing comma. Train says comma 61% of the time. Greedy says "5" almost every time.

That one position is cheap to read. One forward pass up to "0.7", compare p("5") with p(","). No retraining.

Same pass gives the expectation over the tenths digit instead of its argmax. If the rank 0.70 is real, that might hand back the spread greedy throws away.

Is the 0.65/0.75 preference in the logits, or only in the argmax?