15 KiB
Composition strategy pass — from a7dc34f / overnight-review
Human rejected the prior macro-form, splices and arrival preparation. Preserve engineering foundation; do not protect SA3 out of sunk cost. All physical audio stays muted; session locks retained. Private Desktop/RoadScore work only. No fine-tuning, invalid graph snapshots, public edits, or speculative Chestnut ports.
Priorities: one aggressive instrumental K-pop/game-score SA3 control; screen 3–5 alternatives cheaply; run ACE-Step1.5 locally if feasible and at most one justified additional candidate; generated transitions and actual outro; deterministic tempo-aware gesture family and causal turn-signal/curve/nav integration; regress Mac/native/stored replay; one listening page and precise evidence.
Development hardware: Mac M1 Max32GB,32GPU cores;~43GiB disk free at start. Alternative downloads/installations stay isolated and bounded. A musical winner requires human listening; claims in model documentation are not our measured results.
Checkpoint during execution
- Fresh SA3 control: nine trained-weight outputs; five deliberately extreme roles share one fresh identity. Three transitions are generated from existing context with no end anchor. All remain human-listening candidates.
- ACE-Step1.5 turbo + MLX planner/DiT/VAE successfully generated a 90s structured piece, five reference-conditioned roles and a 40s transition. First90s call108.03s (RTF1.20); not a Chestnut measurement. Initialization248.42s includes model acquisition. Peak process RSS8.10GiB understates total unified-memory pressure; system swap reached17.38GiB. No musical winner declared.
- YuE2 is the single additional executed candidate, using official MPS support with native symbolic planning. Bounded to90s semantic material and20min runtime. No Chestnut port attempted.
- Native arrival run
normal_1789575069:253.9s, eight fresh jobs, zero fallbacks/underflows. First outro-conditioned material203.964s; cadence trigger239.2s. Navigation clearing reset the initial outro counter, despite35.236s of archived outro-conditioned playback. Original evidence retained; accounting fix now separates delivered musical history from navigation intent. Concluding musical behavior is still unverified. - New source-derived percussion layer: finite turn-signal/curve/nav/arrival phrases, estimated pulse phase, sample-addressed scheduling, cancellation recorded for stored replay, source-dependent level/headroom. Synthetic turn-signal example115ms onset, zero scheduling lateness. Estimated beat/bar/key values are not ground truth. -61 tests pass including real-log future-mutation causality regression. Initial test invocations had environment import failures; corrected test environment uses native runtime plus analysis libraries.
Physical output remains muted throughout. Alternative models are isolated under experiments; resident Chestnut worker stays exclusive and unchanged. No MusicGen edits.
Model screen and porting decision
The musical winner is not selected. Listen to the new single review page. Runtime/role labels cannot establish that a chorus is convincing or that transitions sound natural. No candidate has earned Tier3 on human musical evidence; no speculative Chestnut port was started.
| Candidate | Interface and evidence | Execution / decision |
|---|---|---|
| SA3 Small-Music control | Native tinygrad DiT + decoder, cached CPU text conditions, audio-prefix inpainting. One fresh identity, five extreme roles and three explicitly generated transitions. | Demonstrated Chestnut runtime retained. Control roles:28.05s unique audio in about19.3s. Current live continuations:26.006s unique in about19.9–20.3s,RTF0.77–0.78. Clear macro-form remains a listening question. |
| ACE-Step1.5 turbo | Official hybrid planner + flow/diffusion model; instrumental tags, tempo/key metadata, reference audio; upstream repaint/cover interfaces. We tested full-form generation and reference-conditioned roles, not live repaint. | Seven actual local Mac outputs.90s/108.03s,RTF1.20;30s roles take63.15–91.40s,RTF2.11–3.05, including reference encoding;40s transition73.33s,RTF1.83. Preserve for human evaluation as a possible longer-horizon composer. |
| YuE2-3B | Native symbolic ABC melody/chord planning, autoregressive semantic audio tokens, non-autoregressive acoustic flow synthesis, VAE. Instrumental request through style/structure text, not a verified dedicated no-vocal mode. | Actual MPS end-to-end output81.279s/374.55s,RTF4.61. Neither score nor semantic stream hit our token cap. However37.7s of100ms windows are below−50dBFS, much of the latter half; the symbolic plan choseF minor despiteD minor conditioning. Keep original evidence; no further investment this pass. |
| MAGNeT | Documented masked-token generation,300M/1.5B family,10/30s clips, EnCodec decode. No documented full-song structural interface in the inspected guide. | Cheap screen only. Does not present an obvious solution to this pass's missing macro-form; no model downloaded/executed. This is an interface-based screen, not a listening rejection. |
| LeVo2 / SongGeneration | Full-song model claims surfaced in search, but the official repository/raw source returned unavailable during this run. | Cheap screen blocked on source availability; exact failure saved inlevo_access.txt. Do not infer model quality or claim an executed benchmark. |
Primary sources: ACE-Step1.5 official implementation, YuE official implementation, MAGNeT official guide, SongGeneration official repository checked. Local model revisions, sizes and YuE integrity hashes are retained in the experiment/results. No vendor source was patched; only isolated research runners were added.
Why not port the alternatives yet?
ACE's inspected turbo configuration has24 decoder layers,2048 hidden width,16 query/8 KV heads, alternating sliding/full attention, rotary embeddings, RMS normalization, gated feedforwards, and extra lyric/timbre conditioning. Its local packages occupy4.463GiB for the core,3.498GiB for the1.7B planner,1.125GiB for the text embedding model and0.314GiB for the VAE. These are on-disk footprints, not peak allocations. The narrow matrix/attention/convolution primitives are familiar, but this is a much larger conditioning/decoder stack than the SA3 port. Complete simultaneous residency is not established on the8GB Chestnut; host3.5GB makes staging/offload difficult. Quantization, serial model lifetimes or a separate local planner would need engineering and measurement. A2B core does not automatically imply an impossible port, nor does familiar attention imply real-time performance.
YuE's inspected backbone has28 layers,2048 hidden width,16 query/8 KV heads,6144 feedforward width and184704 vocabulary. Audio uses64-channel latents at25Hz and a convolutional VAE. Its official MPS path actually ran; a general PyTorch compatibility layer was unnecessary. Nevertheless it combines sequential planning/semantic generation with32 acoustic flow steps. The unquantized backbone file is7,261,441,640 bytes plus530,512,720-byte VAE, before activations/KV cache. Mac MPS driver allocation at the end was18,031,886,336 bytes (~16.79GiB), not a measured peak. The configured16GiB budget does not enforce a hard MPS cap in this implementation. It is not a demonstrated viable8GB live backend.
ACE peak process RSS was8,291.9MiB, while unified allocations plus other system use caused substantial swap (17.38GiB observed). YuE RSS was1,275.6MiB despite the large MPS driver allocation: RSS alone is particularly misleading here. Alternative runs were serialized, never run concurrently with each other. Their model operations executed on Mac GPU via MLX/MPS; Python/tokenization, file handling and some conditioning/offload work execute on CPU. ACE uses its upstream mixed MLX/PyTorch path. YuE's operator fallback environment was enabled, so this is not proof every operator ran on GPU. Neither is a Chestnut benchmark.
SA3 worker process peak RSS ~175–178MiB and tracked post-decode allocation1,111.85MiB are not whole-bench RAM or verified peak VRAM. Its DiT and decoder execute on Chestnut; CPU handles already-cached text conditioning, noise, replay, resampling, mixing and gestures. Text encoding was performed separately on CPU before the runs. Existing process-cold127–185s and resident15–24s launcher evidence remains historical; this pass's native curve resident launch14.62s is recorded. Alternative initialization timings are not apples-to-apples: ACE248.42s includes downloads; YuE7.42s resolves/checks local files, with lazy model loading included in the generation wall timer. No warm YuE repeat was run merely to improve a number.
Minimum viable architecture recommendation, conditional on listening
Keep the model-independent route/replay → causal input → final PCM → host/archive pipeline. Split musical control by time horizon:
- Immediate / next beat: source-derived deterministic percussion gestures. Turn signal uses an eighth-note motif whose first entrance is quantized to the next sixteenth. Curve preparation uses a one-bar fill and a predicted-peak accent. Navigation uses a small phrase, with a cooldown. No new guessed chord/key is introduced.
- Tens of seconds ahead: cached text conditions and prefix-conditioned generated transitions. Current navigation can select a chorus/bridge transition or concluding material far enough ahead to survive generation and existing buffered audio. We do not pretend the neural model can create a new chorus six seconds before a curve.
- Longer-form composition: if ACE's actual examples win human listening, investigate its longer-form material as a local prebuffered composer or arrangement guide. SA3/Chestnut would still generate meaningful continuing material, and deterministic gestures would own precise timing. This hybrid is a candidate architecture, not a demonstrated coherent multi-model composition system.
The live prototype now uses continuous generated windows, not the previously rejected independent section bank. It still overlaps retained context at decoder boundaries; this cannot guarantee identical tempo, harmony or groove across windows. Three context-included transition outputs expose the underlying transition for listening rather than hiding it with seam tuning. The alternative runners own model-specific preparation and decoding;backends.py records their actual capabilities/results. They are not selectable live workers. The existing transport/score format remains final PCM, avoiding alternative-model imports in the replay process.
The earlier rock example's human impact cannot be re-proved while muted. The earlier dense-source report does document stable107–109BPM pulse, denser transient vocabulary and no premature silence in its continuous run. A plausible lesson is stronger attack contrast and headroom for event punctuation, not simply more distortion. The present source-derived hats/fills/crashes test that hypothesis; this is explicitly an inference awaiting listening.
What the new controls do and do not establish
- A persistent identity is a source prefix and repeated musical target, not a verified recurring motif.
- A source-derived drum bank is stylistically related by material and approximate tempo. Filtered source attacks can still contain pitched leakage. We deliberately avoid claiming a new in-key bass/synth part from unreliable key estimates.
- Gesture start frames are exact in the captured sample stream. The source's beat/downbeat inference is uncertain and can drift across subsequent generated windows. Input polling and host output buffering add latency (Mac target buffer250ms); synthetic115ms scheduling is not an acoustic115ms measurement.
- Curve payoff now follows revisions to the currently delivered model forecast until it is within one beat. Already-started audio is never moved. Curvature forecast and physical steering onset are different quantities; no future recorded steering drives runtime.
- Outro conditioning begins well before the final cadence in the demonstrated arrival run. This is necessary evidence, not proof of musically conclusive phrasing. Original counter failure and the corrected history/intent separation are documented rather than rewriting old records.
- Physical speakers, subjective form, instrumental adherence, groove/key continuity and a convincing ending all remain human verification items. No fine-tuning, model port, or subjective acceptance was claimed.
Replay validation and known capture failures
Native arrival253.9s:8 fresh jobs,0 fallbacks/renderer underflows/output flags;81 repeated turn-signal phrase starts,7 curve preparation/payoff pairs,3 navigation cues and1 arrival preparation. Native curve176.3s:6 fresh jobs,0 fallbacks/underflows/output flags;27 repeated signal phrase starts and4 curve pairs. These counts are phrases, not distinct signal activations. Both had zero sample scheduling lateness. Repeated signal phrases are intentionally queued up to one beat ahead; their queue wait is not the first-signal response latency.
A native video-recording attempt failed the existing1x replay-clock guard after20.8s audio. Retrying without recording passed. A Mac live recording attempt had a persistent~94ms reported DAC offset beginning at42.5s, despite zero PortAudio flags/starvation/recovery. It was stopped at159.4s rather than presented as a synchronized pass. Recording load is a plausible contributor, not a proven root cause. Both failures and original logs remain infailures.json; no timing tolerance was weakened. A full no-recording Mac rerun is being validated, followed by separate stored-score capture.
Stored overlay now derives road phase/lead only from decisions whose archived sample time has been reached; cancelled future gestures stay cancelled. The normal camera/path/lane UI remains the existing replay UI. Capture screenshots now wait for a ready onroad display instead of preserving the preparation screen. No host PCM scheduling or jitter logic changed; its friendly style label was the only host-audio edit.
Final regression: full community routenormal_1789576104 completed571.6s,21 accepted fresh jobs, median unique-audioRTF0.7795, zero host starvation/late samples/output flags and max host alignment0.324ms. Live outro counter correctly recorded21.653s before cadence. Normal camera/path/lane drawing remained active. Mac stored capturenormal_1789576767 verified contiguous samples with zero output flags and0.583ms maximum clock error; VFR review uses10,978 recorded UI timestamps over186.46s including startup, with149ms worst frame gap. Native stored arrivalnormal_1789576781 verified contiguous samples with zero flags and9.90ms maximum clock error. Neither stored run invoked generation. The review video is the archived fresh score, not a new generation presented as live.
Shutdown: owned resident worker stopped; power snapshots match exactly. Normal bench UI running, realIsOnroad=false. Mac system mute and both session locks remain; every test stream was muted. No Bluetooth output or speaker output was enabled. ALSA mixer queries were unavailable even through the permitted read attempt, so no hardware-mixer-register verification is claimed; native muted PCM and session-policy evidence are retained. Public StarPilot tree was not edited; MusicGen remains untouched.