mirror of
https://github.com/firestar5683/StarPilot.git
synced 2026-10-05 05:44:03 +08:00
76 lines
15 KiB
Markdown
76 lines
15 KiB
Markdown
# Composition strategy pass — from a7dc34f / overnight-review
|
||
|
||
Human rejected the prior macro-form, splices and arrival preparation. Preserve engineering foundation; do not protect SA3 out of sunk cost. All physical audio stays muted; session locks retained. Private Desktop/RoadScore work only. No fine-tuning, invalid graph snapshots, public edits, or speculative Chestnut ports.
|
||
|
||
Priorities: one aggressive instrumental K-pop/game-score SA3 control; screen 3–5 alternatives cheaply; run ACE-Step1.5 locally if feasible and at most one justified additional candidate; generated transitions and actual outro; deterministic tempo-aware gesture family and causal turn-signal/curve/nav integration; regress Mac/native/stored replay; one listening page and precise evidence.
|
||
|
||
Development hardware: Mac M1 Max32GB,32GPU cores;~43GiB disk free at start. Alternative downloads/installations stay isolated and bounded. A musical winner requires human listening; claims in model documentation are not our measured results.
|
||
|
||
## Checkpoint during execution
|
||
|
||
- Fresh SA3 control: nine trained-weight outputs; five deliberately extreme roles share one fresh identity. Three transitions are generated from existing context with no end anchor. All remain human-listening candidates.
|
||
- ACE-Step1.5 turbo + MLX planner/DiT/VAE successfully generated a 90s structured piece, five reference-conditioned roles and a 40s transition. First90s call108.03s (RTF1.20); not a Chestnut measurement. Initialization248.42s includes model acquisition. Peak process RSS8.10GiB understates total unified-memory pressure; system swap reached17.38GiB. No musical winner declared.
|
||
- YuE2 is the single additional executed candidate, using official MPS support with native symbolic planning. Bounded to90s semantic material and20min runtime. No Chestnut port attempted.
|
||
- Native arrival run `normal_1789575069`:253.9s, eight fresh jobs, zero fallbacks/underflows. First outro-conditioned material203.964s; cadence trigger239.2s. Navigation clearing reset the initial outro counter, despite35.236s of archived outro-conditioned playback. Original evidence retained; accounting fix now separates delivered musical history from navigation intent. Concluding musical behavior is still unverified.
|
||
- New source-derived percussion layer: finite turn-signal/curve/nav/arrival phrases, estimated pulse phase, sample-addressed scheduling, cancellation recorded for stored replay, source-dependent level/headroom. Synthetic turn-signal example115ms onset, zero scheduling lateness. Estimated beat/bar/key values are not ground truth.
|
||
-61 tests pass including real-log future-mutation causality regression. Initial test invocations had environment import failures; corrected test environment uses native runtime plus analysis libraries.
|
||
|
||
Physical output remains muted throughout. Alternative models are isolated under experiments; resident Chestnut worker stays exclusive and unchanged. No MusicGen edits.
|
||
|
||
## Model screen and porting decision
|
||
|
||
The musical winner is **not selected**. Listen to the new [single review page](results/composition_20260916/index.html). Runtime/role labels cannot establish that a chorus is convincing or that transitions sound natural. No candidate has earned Tier3 on human musical evidence; no speculative Chestnut port was started.
|
||
|
||
| Candidate | Interface and evidence | Execution / decision |
|
||
|---|---|---|
|
||
| SA3 Small-Music control | Native tinygrad DiT + decoder, cached CPU text conditions, audio-prefix inpainting. One fresh identity, five extreme roles and three explicitly generated transitions. | Demonstrated Chestnut runtime retained. Control roles:28.05s unique audio in about19.3s. Current live continuations:26.006s unique in about19.9–20.3s,RTF0.77–0.78. Clear macro-form remains a listening question. |
|
||
| ACE-Step1.5 turbo | Official hybrid planner + flow/diffusion model; instrumental tags, tempo/key metadata, reference audio; upstream repaint/cover interfaces. We tested full-form generation and reference-conditioned roles, not live repaint. | Seven actual local Mac outputs.90s/108.03s,RTF1.20;30s roles take63.15–91.40s,RTF2.11–3.05, including reference encoding;40s transition73.33s,RTF1.83. Preserve for human evaluation as a possible longer-horizon composer. |
|
||
| YuE2-3B | Native symbolic ABC melody/chord planning, autoregressive semantic audio tokens, non-autoregressive acoustic flow synthesis, VAE. Instrumental request through style/structure text, not a verified dedicated no-vocal mode. | Actual MPS end-to-end output81.279s/374.55s,RTF4.61. Neither score nor semantic stream hit our token cap. However37.7s of100ms windows are below−50dBFS, much of the latter half; the symbolic plan choseF minor despiteD minor conditioning. Keep original evidence; no further investment this pass. |
|
||
| MAGNeT | Documented masked-token generation,300M/1.5B family,10/30s clips, EnCodec decode. No documented full-song structural interface in the inspected guide. | Cheap screen only. Does not present an obvious solution to this pass's missing macro-form; no model downloaded/executed. This is an interface-based screen, not a listening rejection. |
|
||
| LeVo2 / SongGeneration | Full-song model claims surfaced in search, but the official repository/raw source returned unavailable during this run. | Cheap screen blocked on source availability; exact failure saved in`levo_access.txt`. Do not infer model quality or claim an executed benchmark. |
|
||
|
||
Primary sources: [ACE-Step1.5 official implementation](https://github.com/ACE-Step/ACE-Step-1.5), [YuE official implementation](https://github.com/multimodal-art-projection/YuE), [MAGNeT official guide](https://github.com/facebookresearch/audiocraft/blob/main/docs/MAGNET.md), [SongGeneration official repository checked](https://github.com/tencent-ailab/SongGeneration). Local model revisions, sizes and YuE integrity hashes are retained in the experiment/results. No vendor source was patched; only isolated research runners were added.
|
||
|
||
### Why not port the alternatives yet?
|
||
|
||
ACE's inspected turbo configuration has24 decoder layers,2048 hidden width,16 query/8 KV heads, alternating sliding/full attention, rotary embeddings, RMS normalization, gated feedforwards, and extra lyric/timbre conditioning. Its local packages occupy4.463GiB for the core,3.498GiB for the1.7B planner,1.125GiB for the text embedding model and0.314GiB for the VAE. These are on-disk footprints, not peak allocations. The narrow matrix/attention/convolution primitives are familiar, but this is a much larger conditioning/decoder stack than the SA3 port. Complete simultaneous residency is not established on the8GB Chestnut; host3.5GB makes staging/offload difficult. Quantization, serial model lifetimes or a separate local planner would need engineering and measurement. A2B core does not automatically imply an impossible port, nor does familiar attention imply real-time performance.
|
||
|
||
YuE's inspected backbone has28 layers,2048 hidden width,16 query/8 KV heads,6144 feedforward width and184704 vocabulary. Audio uses64-channel latents at25Hz and a convolutional VAE. Its official MPS path actually ran; a general PyTorch compatibility layer was unnecessary. Nevertheless it combines sequential planning/semantic generation with32 acoustic flow steps. The unquantized backbone file is7,261,441,640 bytes plus530,512,720-byte VAE, before activations/KV cache. Mac MPS driver allocation at the end was18,031,886,336 bytes (~16.79GiB), **not a measured peak**. The configured16GiB budget does not enforce a hard MPS cap in this implementation. It is not a demonstrated viable8GB live backend.
|
||
|
||
ACE peak process RSS was8,291.9MiB, while unified allocations plus other system use caused substantial swap (17.38GiB observed). YuE RSS was1,275.6MiB despite the large MPS driver allocation: RSS alone is particularly misleading here. Alternative runs were serialized, never run concurrently with each other. Their model operations executed on Mac GPU via MLX/MPS; Python/tokenization, file handling and some conditioning/offload work execute on CPU. ACE uses its upstream mixed MLX/PyTorch path. YuE's operator fallback environment was enabled, so this is not proof every operator ran on GPU. Neither is a Chestnut benchmark.
|
||
|
||
SA3 worker process peak RSS ~175–178MiB and tracked post-decode allocation1,111.85MiB are not whole-bench RAM or verified peak VRAM. Its DiT and decoder execute on Chestnut; CPU handles already-cached text conditioning, noise, replay, resampling, mixing and gestures. Text encoding was performed separately on CPU before the runs. Existing process-cold127–185s and resident15–24s launcher evidence remains historical; this pass's native curve resident launch14.62s is recorded. Alternative initialization timings are not apples-to-apples: ACE248.42s includes downloads; YuE7.42s resolves/checks local files, with lazy model loading included in the generation wall timer. No warm YuE repeat was run merely to improve a number.
|
||
|
||
## Minimum viable architecture recommendation, conditional on listening
|
||
|
||
Keep the model-independent route/replay → causal input → final PCM → host/archive pipeline. Split musical control by time horizon:
|
||
|
||
- **Immediate / next beat:** source-derived deterministic percussion gestures. Turn signal uses an eighth-note motif whose first entrance is quantized to the next sixteenth. Curve preparation uses a one-bar fill and a predicted-peak accent. Navigation uses a small phrase, with a cooldown. No new guessed chord/key is introduced.
|
||
- **Tens of seconds ahead:** cached text conditions and prefix-conditioned generated transitions. Current navigation can select a chorus/bridge transition or concluding material far enough ahead to survive generation and existing buffered audio. We do not pretend the neural model can create a new chorus six seconds before a curve.
|
||
- **Longer-form composition:** if ACE's actual examples win human listening, investigate its longer-form material as a local prebuffered composer or arrangement guide. SA3/Chestnut would still generate meaningful continuing material, and deterministic gestures would own precise timing. This hybrid is a candidate architecture, not a demonstrated coherent multi-model composition system.
|
||
|
||
The live prototype now uses continuous generated windows, not the previously rejected independent section bank. It still overlaps retained context at decoder boundaries; this cannot guarantee identical tempo, harmony or groove across windows. Three context-included transition outputs expose the underlying transition for listening rather than hiding it with seam tuning. The alternative runners own model-specific preparation and decoding;`backends.py` records their actual capabilities/results. They are not selectable live workers. The existing transport/score format remains final PCM, avoiding alternative-model imports in the replay process.
|
||
|
||
The earlier rock example's human impact cannot be re-proved while muted. The earlier dense-source report does document stable107–109BPM pulse, denser transient vocabulary and no premature silence in its continuous run. A plausible lesson is stronger attack contrast and headroom for event punctuation, not simply more distortion. The present source-derived hats/fills/crashes test that hypothesis; this is explicitly an inference awaiting listening.
|
||
|
||
### What the new controls do and do not establish
|
||
|
||
- A persistent identity is a source prefix and repeated musical target, not a verified recurring motif.
|
||
- A source-derived drum bank is stylistically related by material and approximate tempo. Filtered source attacks can still contain pitched leakage. We deliberately avoid claiming a new in-key bass/synth part from unreliable key estimates.
|
||
- Gesture start frames are exact in the captured sample stream. The source's beat/downbeat inference is uncertain and can drift across subsequent generated windows. Input polling and host output buffering add latency (Mac target buffer250ms); synthetic115ms scheduling is not an acoustic115ms measurement.
|
||
- Curve payoff now follows revisions to the currently delivered model forecast until it is within one beat. Already-started audio is never moved. Curvature forecast and physical steering onset are different quantities; no future recorded steering drives runtime.
|
||
- Outro conditioning begins well before the final cadence in the demonstrated arrival run. This is necessary evidence, not proof of musically conclusive phrasing. Original counter failure and the corrected history/intent separation are documented rather than rewriting old records.
|
||
- Physical speakers, subjective form, instrumental adherence, groove/key continuity and a convincing ending all remain human verification items. No fine-tuning, model port, or subjective acceptance was claimed.
|
||
|
||
## Replay validation and known capture failures
|
||
|
||
Native arrival253.9s:8 fresh jobs,0 fallbacks/renderer underflows/output flags;81 repeated turn-signal phrase starts,7 curve preparation/payoff pairs,3 navigation cues and1 arrival preparation. Native curve176.3s:6 fresh jobs,0 fallbacks/underflows/output flags;27 repeated signal phrase starts and4 curve pairs. These counts are phrases, **not distinct signal activations**. Both had zero sample scheduling lateness. Repeated signal phrases are intentionally queued up to one beat ahead; their queue wait is not the first-signal response latency.
|
||
|
||
A native video-recording attempt failed the existing1x replay-clock guard after20.8s audio. Retrying without recording passed. A Mac live recording attempt had a persistent~94ms reported DAC offset beginning at42.5s, despite zero PortAudio flags/starvation/recovery. It was stopped at159.4s rather than presented as a synchronized pass. Recording load is a plausible contributor, not a proven root cause. Both failures and original logs remain in[failures.json](results/composition_20260916/failures.json); no timing tolerance was weakened. A full no-recording Mac rerun is being validated, followed by separate stored-score capture.
|
||
|
||
Stored overlay now derives road phase/lead only from decisions whose archived sample time has been reached; cancelled future gestures stay cancelled. The normal camera/path/lane UI remains the existing replay UI. Capture screenshots now wait for a ready onroad display instead of preserving the preparation screen. No host PCM scheduling or jitter logic changed; its friendly style label was the only host-audio edit.
|
||
|
||
Final regression: full community route`normal_1789576104` completed571.6s,21 accepted fresh jobs, median unique-audioRTF0.7795, zero host starvation/late samples/output flags and max host alignment0.324ms. Live outro counter correctly recorded21.653s before cadence. Normal camera/path/lane drawing remained active. Mac stored capture`normal_1789576767` verified contiguous samples with zero output flags and0.583ms maximum clock error; VFR review uses10,978 recorded UI timestamps over186.46s including startup, with149ms worst frame gap. Native stored arrival`normal_1789576781` verified contiguous samples with zero flags and9.90ms maximum clock error. Neither stored run invoked generation. The review video is the archived fresh score, not a new generation presented as live.
|
||
|
||
Shutdown: owned resident worker stopped; power snapshots match exactly. Normal bench UI running, realIsOnroad=false. Mac system mute and both session locks remain; every test stream was muted. No Bluetooth output or speaker output was enabled. ALSA mixer queries were unavailable even through the permitted read attempt, so no hardware-mixer-register verification is claimed; native muted PCM and session-policy evidence are retained. Public StarPilot tree was not edited; MusicGen remains untouched.
|