- BarScript (experimental preview) β C1.3, denominator augmentation, seed 1234
- 0. Status: the champion comparison is
inconclusive - 1. What it does
- 2. KNOWN LIMITATIONS
- 2.1 The difficulty ladder is 25β32 coverage points less nested than authored
- 2.2 The note-type channel carries almost no information
- 2.3 The note budget is 9 % short, and it is the whole of the apparent timing regression
- 2.4 Raising density squeezes drumrolls out β this is the compromise a user can hit
- 2.5 Absolute bucket calibration is still wrong on the density axis
- 2.6 It fails its own span-quality gate and triggers a pre-registered kill condition
- 2.7 One span in seven is a micro-span
- 2.8 The deployment grid has a two-role conflict that no checkpoint dissolves
- 2.9 Playability floor: gaps no human hand can play
- 2.10 Pattern proxies regress against C1 on hard+oni
- 2.11 The note-type marginal, and one C1 finding that does NOT reproduce
- 2.12 Fine-lattice, tuplet and syncopation accuracy are much worse than average
- 2.13 Smaller, but on the record
- 2.14 The evidence base is thinner than it looks
- 2.15 Serving regime: what
motif_constraint: autoresolves to, and why it matters less here
- 3. Serving contract
- 4. Weights are bfloat16
- 5. Provenance
- 0. Status: the champion comparison is
BarScript (experimental preview) β C1.3, denominator augmentation, seed 1234
Audio log-mel + a bar grid β a Taiko no Tatsujin chart, decoded under a finite-state grammar over bar-scoped plan, skeleton and realization stages.
BarScript is the release name of this charting model; it is also the name of
the bar-scoped multi-stage token encoding it decodes into. It is a preview
under evaluation, published so it can be tried, not as a finished model: the
champion comparison below is inconclusive, and Β§2 lists measured defects β one
of which fires a pre-registered kill condition. The seven published SoftChart
1.x models are a separate, unaffected line.
| source checkpoint | runs/sc2_c1_v13_s1234/best.pt |
| sha256 | 44980bf64b1611ea73c1433c41adceff7596cfeece30dfaad5d228eae4a14e6e |
| seed / step | 1234 / 96 000 of 100 000 |
| best validation CE | 0.5769365892960475 |
| parameters | 9 107 291 |
| shipped weights | bfloat16 safetensors, 18.2 MB (fp32 checkpoint preserved in the run dir) |
| training data | JacobLinCool/taiko-1000-parsed-clean, revision b72da4616d643018e81f372cea06ce51349285e0 |
| label spec | barscript_labels_v1, frozen, shipped as spec.json (sha1 cd876768β¦) |
| training flags vs C1 | --cal-axes big span density Β· --balance-exclude-big Β· --denom-augment 0.5 --denom-augment-mode lattice |
| external pretraining | none |
This card is written for THIS checkpoint. Every number below was measured on
C1.3's own generations under the serving code this package ships
(generate.py md5 73b6e49bβ¦, min_onset_gap_sec = 0.020). Nothing is
inherited from the C1 card. Where an item on the C1 card has no C1.3
measurement, it is marked not measured on this arm rather than carried over.
Read the KNOWN LIMITATIONS section before using any number above it. This model has real, measured defects, one of which fires a pre-registered kill condition, and its difficulty ladder is roughly 25 coverage points less nested than an authored one.
0. Status: the champion comparison is inconclusive
scripts/sc2_gates.py champion, three pairings on one serving code base
(experiments/sc2_eval/FINAL_CHAMPION.md Β§1):
| pairing | champion_verdict |
reason code |
|---|---|---|
| C1.2 vs C1 | inconclusive |
no_seed_replicate |
| C1.3 vs C1 | inconclusive |
no_seed_replicate |
| C1.3 vs C1.2 | inconclusive |
no_seed_replicate |
MIN_SEEDS_FOR_VERDICT is 2 in both arms and C1.3 has one seed. The rule
fires before any metric is read. No arm has been declared champion.
What the clause ledger says, which is one-directional: C1.3 passes the
required Β§4 clause plan_token_calibration, which C1 fails on both its
seeds. No clause is passed by C1 and failed by C1.3. The two a2 span clauses
fail on all four arms and belong to the campaign, not to this checkpoint.
The shipping rationale, the two named compromises and the decision rule for
choosing a different checkpoint are in
experiments/sc2_eval/SHIP_DECISION.md.
Gate items on the 71-chart hard+oni population, this arm against the C1 seed-1234 arm under identical code:
| C1 | C1.3 | |
|---|---|---|
| fail / pass / na | 9 / 10 / 6 | 8 / 11 / 6 |
| newly passing | β | control_own_axis_density, control_named_regression_density_big_count |
| newly failing | β | a2_span_placement_fit (knife-edge, Β§2.6) |
1. What it does
- Input β 128-bin log-mel at 22 050 Hz (
preprocessor_config.jsonfreezes the exact decode/STFT/mel contract), plus a bar grid with explicit measure edges, exact rational meters and per-bar lattice denominators. - Conditions β course, authored level, density bucket, and the split-axis
knobs
big_rate/span_rate/stream/sync. - Output β a BarScript token stream decoded under
BarscriptFSM, exported as TJA on the 96-slot lattice. - Family decoding β
generate_chart_familydecodes a song's courses as one nested ladder, hardest first, each easier course biased toward its harder sibling's onsets. Off by default; the campaign's headline population was decoded independently (Β§2.1).
Capabilities this checkpoint carries (config.json:capabilities): aux,
beat_head (hi-res), hierarchical_ctx, axis_knobs, sync_token,
stage_emb, span_duration_head, tempo_head, rich_section_stats. It
carries no style, sibling, ctx, plan-prefix, slot, dual,
align, mask_infill, func_time, global_ctx or complexity conditioning;
requests on those axes reach nothing.
What works
Evidence marks: β n β₯ 71 charts with a measured C1 seed band; β n = 19β71 charts or 8β12 sweep cells, one seed; β n < 19 or a CI covering zero.
| axis | this checkpoint | reference | mark |
|---|---|---|---|
| Density knob β median within-song Ο 0.850, sign p 0.0078, 8/8 cells positive, +1.413 nps firstβlast, 12 % monotone adjacency | C1 moves the chart +0.024 / +0.151 nps over the same range and fails the gate; C1 seed-to-seed ΞΟ 0.003 (p 1.000) | 8 songs Γ 5 buckets, oni | β |
big_rate knob β Ο 1.000, p 0.00049, 91.7 % monotone, and calibrated: true β mean absolute bucket error 0.533, within Β±1 bucket 93.3 % |
first axis in the whole campaign to clear calibrated; C1 0.533 vs 1.200 / 0.967 |
12 songs Γ 5 buckets, oni | β |
span_rate knob β Ο 0.810, p 0.0117, calibrated: true, mean abs err 1.483; and unlike C1 it no longer drags hit density with it |
C1's span knob flags onset_nps / hit_nps / n_onsets as interference; C1.3 flags only the definitional span_count |
12 songs Γ 5 buckets, oni | β |
Span budget β 739 spans against the authored 762 (0.970Γ); span_rate_gen_per_min gap to authored 0.049 against a C1 seed band of 0.198 |
C1 1.160β1.244Γ, C1.2 1.259Γ | 97 charts | β |
| FSM guarantees β 0 unclosed, 0 orphan ends, 0 swallowed hits over 739 generated spans | holds on all four arms | 97 charts | β |
Placement does not degrade β precision 0.6477 (C1 band 0.6480β0.6502), median |offset| 0 ms, exact_slot_lift_over_null 1.1755 (above the C1 band), long-song drift tests significant 2 vs C1's 6 / 5 |
see Β§2.3 for why apparent recall drops | 71 charts | β |
| Accent over-emission reduced furthest β pooled hard+oni big share 0.0873 against C1's 0.1082 / 0.1027 (3.3 seed bands) and an authored 0.0589 | still outside the pre-registered stop window 0.045β0.075 | 71 charts | β |
Cheaper β 1 618.6 tokens/chart, 19.22 tokens/bar, max_window_tokens_p99 449 |
C1 1 731.2 / 20.51 / 515 | 71 charts | β |
Deployment grid repaired β on a forced /16 BPM grid, notes Γ· authored 1.130 / 1.112 / 1.067 / 1.027 and onset F1 0.333 / 0.374 / 0.466 / 0.580 |
C1 1.688 / 1.461 / 1.297 / 1.210 and F1 0.318 / 0.376 / 0.477 / 0.599; C1.3 beats C1.2 on 12 of 12 deployable cells | 13 songs Γ 4 courses, one seed | β |
Empty-bar declaration restored β /16 empty bars 12.82 / 8.60 / 4.55 / 3.14 % against C1.2's 5.96 / 3.23 / 1.08 / 0.33 %, sign-significant on 3 of 4 courses |
authored 15.69 / 12.30 / 9.08 / 6.27 % | 13 songs | β |
2. KNOWN LIMITATIONS
Nothing in this section is softened. Where a limitation is invisible to the gate suite, that is said.
2.1 The difficulty ladder is 25β32 coverage points less nested than authored
Coverage = the fraction of the easier chart's notes that have a note in the harder chart within tolerance. Two decode modes, two different numbers, and both belong on the record.
| adjacent pair | authored charts | independent decode (package default) | family decode Ξ² = 2 (authored grid) | family decode Ξ² = 2 (/16 deploy grid) |
|---|---|---|---|---|
| easy β normal | 0.970 | 0.658 | 0.730 | 0.726 |
| normal β hard | 0.979 | 0.667 | 0.797 | 0.811 |
| hard β oni | 0.985 | 0.660 | 0.822 | 0.834 |
Independent-decode figures: FINAL_CHAMPION.md Β§8.1, n = 13 / 13 / 35 songs, one
seed, same serving code as this package (β/β). Family-decode figures:
C13_PREDICTION.md Β§3.3, n = 13 songs, one seed (β).
Five things make this worse than the headline:
nesting_coverageandhand_agreementFAIL the family gate on every arm measured, C1 included.difficulty_monotonicitypasses on all of them.- C1.3 is slightly worse than C1 here. C1's independent-decode coverage on the same population is 0.690 / 0.704 / 0.737; C1.3 is β0.032 / β0.037 / β0.077 against C1 seed bands of 0.010 / 0.021 / 0.025. The move is small against the ~0.30 gap to authored that every arm shares, but it is in the wrong direction.
- Hand agreement at coinciding hits is near chance on hardβoni: C1.3 0.539 against a marginal-chance null of 0.500 and an authored 0.753. When two generated courses agree that a note belongs somewhere, which drum they pick is near-independent across courses.
- The
/16figure overstates the model's nesting. The same checkpoint on a/96grid drops to 0.649 / 0.730 / 0.752, worse on 11β12 of 13 songs (p = .003 / .022 / .022). A coarse lattice manufactures agreement by leaving few places to disagree (C13_PREDICTION.mdΒ§5.2). family_bias = 2.0/family_hand_bias = 1.5are PROVISIONAL β described ingenerate.pyas logit offsets on a decoder never trained at that setting, and never calibrated against the authored target.
2.2 The note-type channel carries almost no information
71 charts, 22 260 matched slots (FINAL_CHAMPION.md Β§5.2, β).
| quantity | this model | its own floor | corpus Β§2.1 floor |
|---|---|---|---|
| 4-way accuracy | 0.4776 | majority 0.5559 | majority 4-way 0.563 |
| β margin over majority | β0.0783, 95 % CI [β0.1022, β0.0530] | ||
| hand accuracy | 0.5493 | always-don 0.5975 | always-don 0.620 |
| β margin | β0.0482, CI [β0.0721, β0.0229] | ||
| MI(generated; authored), 4-way | 0.01028 bits | β | 0.79 % of H(authored) = 1.300 bits |
| MI, hand | 0.00534 bits | β | 0.55 % of 0.972 bits |
- The model sits below its own majority floor on both endpoints, and below the campaign's corpus baselines (56.3 % / 62.0 %), which are a fixed reference line from a separate 59 961-hit census, not this population.
- The honest caveat, which cuts the other way. The published floors are argmax predictors scored against a temperature-1.0 top-p-0.95 sample. Against the sampler-appropriate i.i.d. floor (Ξ£ pΒ² = 0.4582 4-way, 0.5190 hand) this model is above baseline by +0.0194 and +0.0303 β the largest 4-way excess of any arm in the campaign. Both readings belong on the record.
- MI is flat at ~0.010β0.012 bits on all four arms (seed band 0.0008). The accuracy differences between arms are marginal-matching, not information.
- Accuracies are conditional on coverage 0.6442 (precision side) / 0.5842 (recall side); more than a third of generated hits have no authored partner, and coverage is 9 % lower than C1's, so this population is smaller and differently selected than C1's.
- No inter-charter agreement ceiling exists for this split β no (song, course) carries two independent authored charts β so 100 % is not a legitimate target and is not used as one.
- Greedy re-decoding was not repeated on this arm. The C1-era finding that
greedy closes part of the gap at an unacceptable
motif_reusecost is not measured on this arm and must not be quoted for it.
2.3 The note budget is 9 % short, and it is the whole of the apparent timing regression
| C1 (s1234 / s4321) | C1.3 | |
|---|---|---|
onset_precision |
0.6480 / 0.6502 | 0.6477 |
| note budget (gen Γ· authored notes) | 1.0046 / 0.9912 | 0.9069 |
onset_recall |
0.6510 / 0.6445 | 0.5874 |
frac_of_authored_within_jnd |
0.6463 / 0.6397 | 0.5840 |
type_coverage_recall |
0.6465 / 0.6401 | 0.5842 |
precision Γ budget equals onset_recall to four decimals on every arm, by
construction. Those three "timing" rows are one quantity, and it is the note
budget, not placement: precision is inside the C1 seed band, median offset is
0 ms, exact_slot_lift_over_null is above the band, and significant drift
tests fall from 6 / 5 to 2.
The shortfall itself is a real regression and its cause is unknown. It
appears on both C1.2 and C1.3, so it belongs to the --cal-axes /
--balance-exclude-big pair rather than to the denominator augmentation. It is
the largest practical regression in this release.
2.4 Raising density squeezes drumrolls out β this is the compromise a user can hit
Median spans/min over eight oni songs, per requested density bucket
(authored 3.885; FINAL_CHAMPION.md Β§4.3, 8 songs Γ 5 buckets, β):
| requested density | 0 | 3 | 7 | 11 | 15 |
|---|---|---|---|---|---|
| C1 (s1234) | 3.629 | 3.055 | 4.153 | 2.818 | 2.971 |
| C1.3 | 3.550 | 4.413 | 2.172 | 2.727 | 1.279 |
Ο(density β span_per_min) = β0.759, sign p 0.0078, firstβlast β2.443 /min.
At the top of the density request C1.3 delivers 33 % of the authored span
rate; C1 delivers 76β79 %, C1.2 44 %.
- Do not advertise the density knob at its top setting until this is fixed.
- There is a named suspect:
cal_spannever left 0.1667 on either challenger, so this is plausibly a calibration term that never trained rather than an intrinsic cost of the knob. - The mirror image is a gain: C1.3's
span_rateknob no longer drags hit density, which C1's does.
2.5 Absolute bucket calibration is still wrong on the density axis
| axis | exact bucket | within Β±1 | mean abs error | mean signed | calibrated |
|---|---|---|---|---|---|
| density | 5.0 % | 25.0 % | 4.125 | β1.075 | false |
big_rate |
53.3 % | 93.3 % | 0.533 | +0.267 | true |
span_rate |
β | β | 1.483 | β | true |
The density knob orders the request correctly and lands in the wrong bucket, under-shooting by ~1 bucket on average β more than C1.2 does. Rank control is real; absolute density targeting is not delivered.
Note the cost that comes with the big_rate calibration win: C1.3 has the
shortest big_rate ladder of any arm (firstβlast +0.159 against C1's
0.183 / 0.224), so the top of that request now delivers less accent than C1's
did.
2.6 It fails its own span-quality gate and triggers a pre-registered kill condition
71 charts hard+oni (467 closed spans) and 26 charts easy+normal (272):
| axis | hard+oni | easy+normal |
|---|---|---|
duration_low_tail |
fail (rate verdict passes; balloon 0.250 beats = 0.75Γ the shortest authored balloon, roll 0.250 = 0.60Γ) | fail (balloon 0.250 beats = 0.20Γ the shortest authored balloon; 45/159 balloons below authored support) |
forced_close |
fail β 7 / 467 = 1.499 % (0 unclosed, 7 clamp-forced) | fail β 7 / 272 = 2.574 % |
over_span_max |
fail β one balloon 7.72 s = 1.19Γ the 6.5 s ceiling | fail β one balloon 7.58 s = 1.17Γ |
placement_fit |
fail β AUC 0.560 (CI [0.529, 0.593]) against a 0.56 floor; authored reference 0.668 | pass β AUC 0.604 |
orphans / swallowed / rate / duration |
pass | pass |
a2_kill_condition_not_triggered |
fail β the kill condition FIRED | fail |
The pre-registered kill condition (SOFTCHART2_DESIGN.md Β§4-A2) reads: rate is
calibrated but placement/length quality is bad β concede that the bar-scoped
commitment mechanism is insufficient and escalate to an independent span-state
objective. It fires on all four arms, including C1. Shipping does not
discharge it.
placement_fit is a knife-edge, not a finding. C1.2 passes at 0.561 and
C1.3 fails at 0.560 with CIs overlapping every other arm across their whole
width. One thousandth of AUC decides the verdict. It is reported, not used.
2.7 One span in seven is a micro-span
All generated spans, 97 charts, FINAL_CHAMPION.md Β§2 (β):
| n spans | p1 | p5 | p50 | p90 | min | share < 0.25 s | |
|---|---|---|---|---|---|---|---|
| authored | 762 | 0.177 | 0.273 | 0.963 | 2.419 | 0.083 | 3.8 % |
| this model | 739 | 0.089 | 0.150 | 0.818 | 2.258 | 0.072 | 14.7 % |
Per course: easy 6.6 %, normal 8.1 %, hard 13.7 %, oni 25.2 % β authored 3.0 / 3.3 / 2.1 / 7.1 %.
- C1.3 does not inherit C1.2's worsening (16.4 %); 14.7 % is exactly at the top of the C1 seed band (13.6β14.7 %). It is not an improvement either.
- Mechanism is known and unfixed on every arm. Duration bucket 0 is open at the bottom; 123 of C1.3's 739 spans sit in it, at a median 0.5 beats against the authored 0.917.
- The shipped floor does not fix it.
span_min_sec/span_min_beatsfired 263 times (span_close_floor) plus 33 carries across 97 charts and the < 0.25 s share is still 14.7 %. - The mitigation is unbuilt.
span_bucket0_profileisnullin this package becausescripts/build_span_bucket0_profile.pyhas never been run on the train split. This is the shortest available fix and it is not in the box. - The gate cannot see it directly.
a2_span_durationcompares p50/p90/p99 only and passes;duration_low_tailcatches the support violation but not the rate (its rate verdict passes on this arm). - On a
/16deployment grid the same quantity is 14.7 / 12.4 / 30.6 / 21.8 % β hard is untouched by the augmentation and is the worst cell in the release.
2.8 The deployment grid has a two-role conflict that no checkpoint dissolves
BAR_DENOM is simultaneously an input the grid supplies and a trained output
token the model reads as a density announcement. bar_denoms="supplied", the
shipped default, teacher-forces it.
13 songs Γ 4 courses, one seed, --cond authored, C13_PREDICTION.md:
| easy / normal / hard / oni | authored grid | forced /16 |
/96 |
/24 |
|---|---|---|---|---|
| notes Γ· authored | 0.917 / 0.919 / 0.875 / 0.899 | 1.130 / 1.112 / 1.067 / 1.027 | 0.944 / 0.950 / 0.943 / 1.027 | 1.102 / 1.070 / 0.950 / 0.921 |
| onset F1 | 0.613 / 0.565 / 0.595 / 0.647 | 0.333 / 0.374 / 0.466 / 0.580 | 0.303 / 0.325 / 0.393 / 0.535 | 0.342 / 0.370 / 0.442 / 0.542 |
| genuine triplets (den 3/6/12/24) % | 1.13 / 2.48 / 6.34 / 9.43 | 0.00 (arithmetically impossible) | 3.45 / 5.78 / 10.57 / 11.46 | 14.13 / 20.31 / 28.57 / 38.70 |
| ultra-fine (den 48/96) % | ~0 | 0.00 | 11.53 / 13.22 / 15.62 / 6.60 | 0.00 |
| nesting (family Ξ² = 2) | 0.730 / 0.797 / 0.822 | 0.726 / 0.811 / 0.834 | 0.649 / 0.730 / 0.752 | 0.714 / 0.769 / 0.779 |
Authored triplet rates for reference: 4.20 / 5.72 / 7.43 / 9.95 %; ultra-fine ~0.00 %.
The recommended deployment configuration is --grid bpm --grid-bpm-denom 16
with the package default bar_denoms="supplied". Its two costs, stated
plainly: triplets are exactly 0 and cannot be otherwise, and hard micro-spans
stay at 30.6 % against an authored 2.1 %.
Four consequences a deployer must carry:
- Deployment costs 10β46 % of onset F1 relative to the authored grid:
easy β46 %, normal β34 %, hard β22 %, oni β10 %. Every
--grid authorednumber in this card is therefore optimistic for a user who supplies only BPM. p90 |offset|is 0 ms on the authored grid and 14.6β36.7 ms on the bpm grids. That residual belongs to the bar edges; no denominator choice touches it.- The model's off-dyadic rate is a property of the model, not the lattice.
Given any 3-divisible grid it puts 2β3Γ the authored share off the dyadic
grid; the grid only decides whether that mass gets called "triplets" (
/24) or "ultra-fine jitter" (/96). This is a training-side defect and no deployment grid fixes it. - The gate suite cannot score the deployment path.
robustness.py:629-642raisesContractErroron uniform-BPM arms, so every gate timing metric isnullfor the path the model ships on; the/16figures above come fromtiming.py, a second implementation that reproduces the gate on the authored grid.
The deploy_bpm_grid profile in this package is NOT the recommendation. It
sets bar_denoms="model", bar_denom_mask="div96", per DEPLOY_GRID_FIX.md
Β§7 β a recommendation made explicitly conditional on C1.3 failing, which it did
not. That arm was measured on C1 only: density 1.06 / 1.08 / 0.97 / 1.01 and
nesting 0.635 / 0.706 / 0.778. It has never been run on this checkpoint. The
profile is kept because it is implemented and because the arm is the obvious
next experiment, not because it is advised here.
One contradiction inside this package, stated so nobody has to find it.
config.json:serving_profiles.deploy_bpm_grid.basis still quotes
DEPLOY_GRID_FIX.md Β§7 verbatim, including the words "Recommended now", and
still carries C1's nesting figures. That string is generated by
scripts/sc2_package.py, which is read-only for this release. This section
supersedes it. The generated string is left intact rather than silently
diverging from the shipped code.
2.9 Playability floor: gaps no human hand can play
Recomputed on this arm's own events, 71 charts hard+oni, generated vs authored
on the same songs (producer experiments/sc2_eval/ship_c13/playability.py,
records .../records/playability_all_arms.json; β):
| this model | authored | |
|---|---|---|
| inter-onset gaps < 40 ms | 102 across 29 charts | 2 across 1 chart |
| β as a share of adjacent pairs | 0.296 % (102 / 34 484) | 0.0053 % (2 / 38 032) |
| gaps below the series record 29.4118 ms | 19 | 0 |
| shortest gap | 20.8 ms (the serving floor) | 39.7 ms |
| peak burst in any 1 s window | 17 notes | 15 notes |
C1 on the same population: 151 gaps (0.395 %), shortest 16.7 ms, peak 22 notes/s. C1.3 is better on every row and still 56Γ the authored rate.
min_onset_gap_sec = 0.020is a degeneracy guard, not a fix: it is set below the fastest thing the series has shipped (29.4118 ms, TAIKO-TONGUE-TWISTER oni, BPM 170, 48th notes) precisely so it cannot refuse a chart the domain writes. It fired 125 times across 97 charts with 0 fail-opens. A mask cannot fix a distribution.- The
generate.pyconstant comment quotes the generated sub-40 ms rate as ~0.35 % against ~0.045 % authored. That authored figure does not reconcile with the 0.0053 % measured here, and the two have never been put on the same denominator. The excess is 8Γ on the comment's accounting and 56Γ on this one. - Good news that replaces a C1-card claim. The C1 card reported the model
placing hits in 19.6 % of sub-0.25 s scaffolding bars. Under today's serving
floors this arm places 0 hits in all 112 such bars across 97 charts, exactly
as the authored side does;
degenerate_bar_hits_blockedfired 12 times with 0 fail-opens.
2.10 Pattern proxies regress against C1 on hard+oni
71 charts, one code version, C1 band = |s1234 β s4321| (FINAL_CHAMPION.md Β§3):
| metric (lower is better) | C1 mean | C1 band | C1.3 | Γband | relative |
|---|---|---|---|---|---|
motif_reuse_gap |
0.3630 | 0.0042 | 0.4222 | +14.1 | +16.3 % |
compression_gap |
0.0976 | 0.0006 | 0.1116 | +23.3 | +14.3 % |
motif_ref_marginal_js |
0.0331 | 0.0046 | 0.0452 | +2.6 | +36.6 % |
ioi_js_per_chart |
0.0295 | 0.0006 | 0.0352 | +9.7 | +19.3 % |
ul_4gram_js_per_chart |
0.2402 | 0.0055 | 0.2530 | +2.3 | +5.3 % |
class_4gram_js_per_chart |
0.1337 | 0.0080 | 0.1357 | +0.2 | +1.5 % |
The Γband column overstates the case β several bands are under 1 % of their own
level β so the relative column is the honest one. On motif_reuse the bootstrap
CIs are nonetheless disjoint: C1 [β0.392, β0.334] / [β0.388, β0.334] against
C1.3 [β0.447, β0.396]. The regression is real and not seed noise.
Four bounds, all measured, none of them a dismissal:
- It does not happen on easy+normal. Every pattern proxy there is inside the
C1 seed band,
motif_reuse_gapis marginally better than C1's mean, andclass_4gram_jsis better on both challengers. - The decomposition puts the loss on rhythm and hand, while the accent layer improves: rhythm 0.1251 β 0.1640, hand 0.1774 β 0.2102, accent 0.0606 β 0.0480.
ioi_js_per_chartis where the augmentation shows up β it reads the exact rational IOI lattice, and it is C1.3's worst pattern metric relative to C1.2 (+9.7 band against +1.5).- The proxy reads mostly a channel carrying ~0.010 bits about the authored type (Β§2.2), and it has never been validated against a listener. It is quoted because Β§4 names it, not because it is strong.
Realized motif_reuse is 0.3165 against an authored 0.7387 β under half
the authored repetition.
2.11 The note-type marginal, and one C1 finding that does NOT reproduce
Big notes are 1.48Γ the authored rate pooled hard+oni (0.0873 vs 0.0589), worst on the easiest course:
| course | authored | C1.3 | C1 (s1234 / s4321) |
|---|---|---|---|
| easy | 18.80 % | 26.22 % (1.40Γ) | 32.68 % / 29.65 % |
| normal | 11.75 % | 17.41 % (1.48Γ) | 19.21 % / 20.77 % |
| hard | 7.38 % | 10.61 % (1.44Γ) | 13.11 % / 12.73 % |
| oni | 4.88 % | 7.40 % (1.52Γ) | 9.29 % / 8.58 % |
The pre-registered stop window (0.045β0.075 pooled hard+oni) is not met.
The C1 card's Β§2.8 claim does not reproduce here and is corrected. On C1 the
realized class marginal was 13Γ closer to the training loss weights than to the
corpus. On C1.3 it is the other way round: JS(gen β authored) = 0.00588
against JS(gen β class-weight prediction) = 0.00746
(JS(authored β weights) = 0.01452). --balance-exclude-big moved the marginal
off the loss weights and toward the corpus. The log-log fit of realized share
against class weight collapses from slope 0.378 (Ο 0.436) on C1 to slope 0.078
(Ο 0.156) on C1.3.
2.12 Fine-lattice, tuplet and syncopation accuracy are much worse than average
robustness.py timing strata, 71 charts, exact-slot rate (overall 0.5840):
| stratum | n authored | exact-slot | C1 (s1234) |
|---|---|---|---|
| lattice step 11.6β23.2 ms | 114 | 0.342 | 0.456 |
| lattice step 23.2β46.4 ms | 790 | 0.443 | 0.465 |
| lattice step β₯ 92.9 ms | 24 228 | 0.594 | 0.651 |
| positions with a denominator divisible by 3 | 3 480 | 0.345 | 0.496 |
| syncopation band 0 β band 5 | 20 846 β 1 669 | 0.615 β 0.434 | 0.672 β 0.514 |
| local IOI 25β50 ms | 189 | 0.402 | β |
| BPM > 250 | 2 570 | 0.499 | β |
Every stratum is lower than C1's, which is the note-budget effect of Β§2.3 acting on a per-stratum recall, but the shape is the finding: tuplet positions and fast lattice steps are 0.24 below the overall rate, and accuracy decays monotonically with syncopation.
2.13 Smaller, but on the record
| finding | number | n | source |
|---|---|---|---|
long_song_no_drift fails |
2 of 21 metrics significant at BH q β€ 0.05 (dens_signed_err, plan_dens_bucket_tvd) β C1 fails 6 / 5 |
71 charts | robustness.json β |
| Harness-level generation failures | 6 of 71 song-courses refused (#BRANCHSTART and non-representable meters); 1 chart exported gen_tja = null β identical to C1 on the same population |
71 charts | index.json β |
| Slot-export off-lattice events | 8 across 97 charts (C1: 5) β the documented contract path, not a crash | 97 charts | decoder counters β |
escape_hatch fires |
91 times across 97 charts (C1: 84) | 97 charts | decoder counters β |
| Generation wall time | 9.29 s/chart, 4.33 s per audio-minute β contended, the GPU was shared, quoted as a declaration only | 71 charts | β |
| Complete-bar rest placement | not measured on this arm (C1: Jaccard 0.474, recall 0.623) | β | β |
| Greedy-decode contrast | not measured on this arm | β | β |
| Micro-span seed sensitivity | not measured on this arm (C1: 24.0 % vs 32.4 % across two sampling seeds on 19 oni charts) | β | β |
2.14 The evidence base is thinner than it looks
- One seed. Every number in this card is n = 1 in the training seed. The C1 seed band quoted throughout is C1's, used as a reproducibility scale; it is a range over n = 2 and carries no confidence statement.
- easy and normal are evaluated on 13 songs. The 71-chart population is
hard+onionly; the easy/normal population is 26 charts from 13 songs, and its seed band is 4β5Γ wider than hard+oni's. - Every difficulty-family and deployment-grid number is 13 songs, one seed (β). Every knob sweep is 8β12 songs Γ 5 buckets (β).
- At these sizes a gate
passcarries little information. Worked example from this campaign: thebig_share_of_hitsgate flips between two C1 training seeds while the failing seed's point estimate is smaller. At n = 36 that gate cannot rank checkpoints, and no pass/fail on it should be quoted as evidence for any arm. - Nothing here is audio-referenced except
a2_span_placement_fit. Every other quantity is chart-vs-chart on the exact rational lattice or on hit order. No timing window and no game judgement parameter appears anywhere in this path. best_valis comparable across C1 / C1.2 / C1.3 (one training cache) and not comparable to C1.1 or to any arm on a different cache. C1.3's 0.5769 is 2.9 % worse than C1.2's 0.5605; that was the registered trade and the deployment side won it.- No plan-neutral arm was generated, so the
plan_neutral_fallbackΒ§4 clause isnaand the ledger is incomplete by one required clause.
2.15 Serving regime: what motif_constraint: auto resolves to, and why it matters less here
motif_constraint: "auto" resolves to OFF for this checkpoint: the training
cache's barscript_md5 (634e3dc3β¦) differs from the serving barscript.py
(9afc715aβ¦), so the gold sequences it trained on never satisfied the MOTIF hard
constraint. train_serve_matched = true β OFF is the train/serve-matched choice
and it is the right default.
Unlike the C1 package, this card needs no correction for it. Every number in
this card was measured with the constraint OFF, i.e. under exactly this
package's resolved default; the 20 evaluation runs all record
serving_fsm.motif_constraint: off, verified: true. In particular
motif_ref_marginal_js here is 0.0452 (hard+oni) and 0.0114
(easy+normal), both measured OFF.
Two things that still belong on the record:
- The ON/OFF blast radius was measured on C1 only, where switching the
constraint on moved
motif_ref_marginal_js0.0311 β 0.0189 (β39 %). On the C1 package the shipped default therefore makes the correct value 0.0311, not the 0.0189 in the older tables. That contrast has never been measured on C1.3, so no ON-constraint number should be quoted for this checkpoint. - Serving under the constraint ON would be train/serve mismatched for this checkpoint and is not a supported configuration.
3. Serving contract
Every value below is in config.json:serving, with its justification in
config.json:serving_basis. Pass them explicitly β the package's
serving_kwargs() does β so the recorded contract is the one that reaches the
decoder.
| parameter | value | one-line basis |
|---|---|---|
min_onset_gap_sec |
0.020 | Degeneracy guard, not a corpus percentile. Must stay strictly below the series record of 29.4118 ms (BPM 170, 48ths). Two earlier corpus-derived values (0.0395, 0.0300) were both wrong. See Β§2.9. |
degenerate_bar_sec |
0.25 | Authored charts place zero hits in any bar under 0.375 s (3 010 bars). Largest round threshold with zero counterexamples and 1.5Γ margin; identical to the gate's threshold. Blocks hits only, never span geometry. |
span_min_sec / span_min_beats |
0.0833 / 1β3 | Authored population minima. Fail-open, counted. Does not fix Β§2.7. |
span_bucket0_profile |
null | Not built β see Β§2.7. |
family_bias / family_hand_bias / family_mode |
2.0 / 1.5 / bias |
Provisional, never calibrated β see Β§2.1. Family decode is opt-in. |
motif_constraint |
auto β resolves off here |
Train/serve matching by barscript.py md5 β see Β§2.15. |
bar_denoms / bar_denom_mask |
supplied / lattice |
The evaluated regime, and the recommended deployment regime on a uniform /16 grid. The deploy_bpm_grid profile exists but is not recommended for this checkpoint β see Β§2.8. |
greedy / temperature / top_p |
false / 1.0 / 0.95 | As evaluated. |
plan_temperature |
null (follows temperature) |
Never exercised in evaluation; shipped unset rather than tuned. |
_has_sync must be forced on at load. train.py records
sync_token=False while the training prefix carries the SYNC slot, so a loader
that trusts the recorded flag drops the slot and every sync bucket reaches the
decoder as an identical prefix. The package loader does this and records why in
config.json:serving_prefix_fix. This is a workaround; the fix belongs in
train.py.
Recommended deployment configuration
--grid bpm --grid-bpm-denom 16 # uniform /16 grid
bar_denoms = "supplied" # the package default
motif_constraint = auto -> off # resolved at package time
family decode: optional; Ξ² = 2.0 / 1.5 is provisional
Expected behaviour under it, 13 songs Γ 4 courses, one seed: notes within 13 / 11 / 7 / 3 % of authored; onset F1 0.333 / 0.374 / 0.466 / 0.580; zero triplets; hard micro-spans ~30 %.
4. Weights are bfloat16
Training ran bf16 autocast and CUDA inference runs bf16 autocast, so fp32
storage carried no information the forward pass could use β the argument the
v1.5 release made and verified. Measured cast cost: max absolute delta
0.00711 on frontend.2.weight (0.349 % of that tensor's max) over 9 107 291
float elements.
This is an argument about the compute dtype, not a proof that the two checkpoints decode identically. Sampling is chaotic in the logits, so individual charts can differ. The paired fp32-vs-bf16 evaluation has not been repeated for BarScript. The fp32 checkpoint is preserved in the source run directory and is the reference for any bit-level comparison.
5. Provenance
| serving code | generate.py md5 73b6e49b00fb60f3229c8539bc7a1029; all 18 src/softchart/*.py md5s in config.json:code.module_md5 |
| evaluation harness | sc2_generate_eval_v2_2_serving_floors, contract sc2_eval_v2 |
| evaluation runs | 20 run indices, 1 028 charts, one serving_floors signature, comparable: true, 6/6 floor counters present |
| headline population | 71 charts hard+oni (36 songs) + 26 charts easy+normal (13 songs), --cond authored, --grid authored, seed 1 |
| deployment population | 13 songs Γ 4 courses Γ 10 arms, --cond authored, one serving tree, one signature |
| label spec | spec.json, field-for-field equal to the cache manifest's copy (spec_parameters_match_cache_manifest: true) |
| uploaded | no (release_manifest.json: "uploaded": false) |
Reports: experiments/sc2_eval/SHIP_DECISION.md,
experiments/sc2_eval/FINAL_CHAMPION.md,
experiments/sc2_eval/C13_PREDICTION.md,
experiments/sc2_eval/C12_EVAL.md, experiments/sc2_eval/DEPLOY_GRID_FIX.md.
- Downloads last month
- -