mirror of
https://github.com/firestar5683/StarPilot.git
synced 2026-10-04 21:33:58 +08:00
Specify coherent Prism hook development and isolated host planning
This commit is contained in:
@@ -0,0 +1,88 @@
|
||||
# Prism hook candidate — proposal, not deployed
|
||||
|
||||
Base: `06c8e7ffe8` (isolated GOLD60 diagnostic), itself based on `a12d1b7854`.
|
||||
Protected GOLD and user-approved planned60 assets remain unchanged.
|
||||
|
||||
## Musical intent
|
||||
|
||||
`prism_hook_spec.py` defines one shared full Prism identity and hook contract for
|
||||
initial, verse, build/prechorus, chorus, bridge and outro. Every role retains
|
||||
128 BPM, D minor, crystal pluck/glass lead timbre, syncopated bass and tight drums.
|
||||
The hook has a recognizable four-note rising contour, signature syncopated rhythm
|
||||
and short answering phrase. Verses expose fragments; builds intensify them;
|
||||
choruses state the complete original; the bridge transforms rhythm/register while
|
||||
keeping identity; the return restores the original phrase and outro resolves it.
|
||||
This replaces the old generic restrained-verse emphasis and keeps specific hook
|
||||
and instrumental instructions present across all role prompts. It cannot guarantee
|
||||
memorability or uniqueness without listening. Fixed seeds intentionally reproduce
|
||||
the same composition; route-derived seeds distinguish comparable route runs.
|
||||
|
||||
## Proposed first demo composition
|
||||
|
||||
The first candidate is one coherent LM-planned60-second composition, not repeated
|
||||
8-second-context verses. At the requested128BPM,32bars occupy60seconds:
|
||||
|
||||
| Target time | Role | Musical development |
|
||||
|---|---|---|
|
||||
|0–3.75s|Introduction,2bars|State recognizable hook|
|
||||
|3.75–18.75s|Verse,8bars|Lighter fragments, breathing space and answer|
|
||||
|18.75–26.25s|Build,4bars|Rising register/subdivisions of same motif|
|
||||
|26.25–41.25s|Chorus,8bars|Complete melody, fuller bass and drums|
|
||||
|41.25–48.75s|Bridge,4bars|Spacious contrasting transformation|
|
||||
|48.75–56.25s|Chorus reprise,4bars|Return the unmistakable original|
|
||||
|56.25–60s|Outro,2bars|Answer and resolution|
|
||||
|
||||
These are prompt targets, **not verified generated timestamps**. Do not force
|
||||
audio cuts or treat these times as detected beats/section boundaries. An actual
|
||||
generated performance may not follow them. The fixed narrative is independent
|
||||
of recorded future route events; operator excerpt selection does not grant
|
||||
runtime access to future telemetry. Current delivered conditions may later inform
|
||||
causal adaptation, but this candidate does not implement it.
|
||||
|
||||
Prefer a60-second route excerpt for this first musical test. A90–120second route
|
||||
requires a newly prepared longer semantic plan, larger native shape validation,
|
||||
or a tested continuation strategy. Do not loop this60s piece, stretch its timing,
|
||||
or restart its initial plan and call that coherent longer-form composition.
|
||||
|
||||
Presentation is subordinate: proposed cue/engagement lanes must not overwrite
|
||||
the hook or rhythmic phase. If a cue conflicts, omit/simplify it. This composition
|
||||
lane does not add signal, curve, engagement or other sonification. The master
|
||||
will combine the route, composition, one signal motif, one curve treatment and
|
||||
one engagement transition in the operator proposal before new generation.
|
||||
|
||||
## Actual preparation status and path
|
||||
|
||||
**Prompt specification only. No new conditioning tensors, semantic codes, native
|
||||
audio, or deployed behavior yet.** Three CPU specification tests pass. Local Mac
|
||||
ACE packages and existing9.4GB model assets are available; no model downloads are
|
||||
needed. New preparation is paused for the combined demo-plan review.
|
||||
|
||||
`prepare_prism_hook.py` is an opt-in Mac preparation tool. Default mode saves the
|
||||
specification only. `--prepare` initializes the existing official ACE/1.7BLM,
|
||||
uses `thinking=True` and fixed seed33602 (or deterministic route-derived seed),
|
||||
records actual returned semantic codes/seed/LM costs, then captures full60s
|
||||
encoder/context tensors immediately before DiT diffusion. It refuses repaint,
|
||||
missing semantic planning, nonfinite/wrong-duration tensors and an unavailable
|
||||
MLX interception path. It never executes native hardware or intentionally
|
||||
generates audio. Candidate output uses a new private directory; protected files
|
||||
and active profiles are never replaced. Exact prompt and tensor hashes are saved.
|
||||
|
||||
The full composition caption/section text are what this first preparation feeds
|
||||
into the model. Per-role continuation contracts are saved for future use but
|
||||
**are not prepared role tensors** and are not wired into current worker startup.
|
||||
Editing these strings cannot alter deployed prepared embeddings. All role-based
|
||||
continuation restoration remains future work after the full-plan musical test.
|
||||
|
||||
Historical preparation recorded12.4s for the old60s LM plan, plus text/conditioning
|
||||
and cold model load; budget minutes rather than promise that runtime. Longer
|
||||
hook text may change encoder length/cost. Tensor output should be a few MiB
|
||||
(1500×128 float32 context≈0.73MiB plus variable-length encoder≈1–severalMiB), not
|
||||
new multi-GB weights. Model loading can consume substantial RAM; current observed
|
||||
diskfree was≈7GB. The approved past run is the playback fallback while preparing.
|
||||
|
||||
After review: run one captured plan, check its actual codes/tensors and provenance;
|
||||
have hardware owner review a candidate-capable isolated native harness (the GOLD
|
||||
probe intentionally pins original asset hashes and must reject this new case);
|
||||
then generate one fixed-seed sample with rawgain processing, inspect/listen,
|
||||
and only afterward integrate subordinate presentation. No seed auditions or
|
||||
per-route cherry-picking. Master/user musical acceptance remains the gate.
|
||||
@@ -0,0 +1,145 @@
|
||||
"""Mac-only semantic planning/tensor capture. Stops BEFORE audio diffusion or decoding."""
|
||||
import argparse
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
from pathlib import Path
|
||||
import random
|
||||
import sys
|
||||
import time
|
||||
import traceback
|
||||
|
||||
from prism_hook_spec import spec
|
||||
|
||||
|
||||
class BoundaryCaptured(BaseException):
|
||||
pass
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument('--assets-root', type=Path, required=True, help='Existing RoadScore assets; read only')
|
||||
parser.add_argument('--output-root', type=Path, required=True, help='New private demo artifacts only')
|
||||
parser.add_argument('--route-seed', type=int)
|
||||
parser.add_argument('--prepare', action='store_true', help='Run one host semantic plan; no audio diffusion')
|
||||
args = parser.parse_args()
|
||||
recipe = spec(args.route_seed)
|
||||
out = args.output_root.resolve() / f'prism_hook_v1_{time.time_ns()}'
|
||||
out.mkdir(parents=True, exist_ok=False)
|
||||
(out / 'hook_spec.json').write_text(json.dumps(recipe, indent=2))
|
||||
state = {'phase': 'spec_only', 'output': str(out), 'spec_sha256': recipe['spec_sha256'],
|
||||
'audio_generated': False, 'native_deployed': False, 'continuation_tensors_prepared': False}
|
||||
|
||||
def record():
|
||||
(out / 'preparation.json').write_text(json.dumps(state, indent=2))
|
||||
|
||||
record()
|
||||
if not args.prepare:
|
||||
print(json.dumps(state))
|
||||
return
|
||||
if sys.platform != 'darwin':
|
||||
raise RuntimeError('This host-preparation entry point is Mac-only; no device execution')
|
||||
base = args.assets_root.resolve() / 'experiments/composition_20260916'
|
||||
if not (base / 'models/ace/checkpoints').is_dir():
|
||||
raise FileNotFoundError('Existing local ACE checkpoints missing; no download attempted')
|
||||
os.environ.update(HF_HOME=str(base / 'cache/hf'), HF_HUB_OFFLINE='1', HF_HUB_DISABLE_TELEMETRY='1',
|
||||
TOKENIZERS_PARALLELISM='false', PYTORCH_ENABLE_MPS_FALLBACK='1')
|
||||
started = time.monotonic()
|
||||
state.update(phase='loading', started_wall=time.time())
|
||||
record()
|
||||
try:
|
||||
import numpy as np
|
||||
import torch
|
||||
import mlx.core as mx
|
||||
from acestep.handler import AceStepHandler
|
||||
from acestep.llm_inference import LLMHandler
|
||||
from acestep.inference import GenerationParams, GenerationConfig, generate_music
|
||||
seed = recipe['seed']
|
||||
random.seed(seed)
|
||||
np.random.seed(seed)
|
||||
torch.manual_seed(seed)
|
||||
mx.random.seed(seed)
|
||||
handler, lm = AceStepHandler(), LLMHandler()
|
||||
status, ok = handler.initialize_service(str(base / 'models/ace'), config_path='acestep-v15-turbo',
|
||||
device='mps', use_mlx_dit=True, offload_to_cpu=True, offload_dit_to_cpu=False)
|
||||
if not ok:
|
||||
raise RuntimeError(status)
|
||||
if not handler.use_mlx_dit or handler.mlx_decoder is None:
|
||||
raise RuntimeError('MLX boundary unavailable; refuse fallback audio generation')
|
||||
status, ok = lm.initialize(str(base / 'models/ace/checkpoints'), 'acestep-5Hz-lm-1.7B', backend='mlx', device='mps')
|
||||
if not ok:
|
||||
raise RuntimeError(status)
|
||||
original_plan = lm.generate_with_stop_condition
|
||||
|
||||
def record_plan(*values, **kwargs):
|
||||
result = original_plan(*values, **kwargs)
|
||||
codes = result.get('audio_codes', '')
|
||||
if not result.get('success') or not codes:
|
||||
raise RuntimeError('Semantic planner did not return audio codes')
|
||||
(out / 'semantic_plan.json').write_text(json.dumps({
|
||||
'caption': kwargs.get('caption'), 'lyrics': kwargs.get('lyrics'),
|
||||
'seeds': kwargs.get('seeds'), 'infer_type': kwargs.get('infer_type'),
|
||||
'temperature': kwargs.get('temperature'), 'cfg_scale': kwargs.get('cfg_scale'),
|
||||
'audio_codes': codes, 'metadata': result.get('metadata'),
|
||||
'time_costs': result.get('extra_outputs', {}).get('time_costs'),
|
||||
}, indent=2))
|
||||
state['semantic_plan_present'] = True
|
||||
return result
|
||||
|
||||
lm.generate_with_stop_condition = record_plan
|
||||
state.update(phase='planning', load_seconds=time.monotonic() - started)
|
||||
record()
|
||||
|
||||
def capture(*unused, **kwargs):
|
||||
required = ('encoder_hidden_states', 'encoder_attention_mask', 'context_latents')
|
||||
arrays = {}
|
||||
for key in required:
|
||||
tensor = kwargs.get(key)
|
||||
if not isinstance(tensor, torch.Tensor):
|
||||
raise ValueError(f'Missing prepared boundary tensor: {key}')
|
||||
array = tensor.detach().float().cpu().numpy()
|
||||
if not np.isfinite(array).all():
|
||||
raise ValueError(f'Non-finite {key}')
|
||||
arrays[key] = array
|
||||
if arrays['context_latents'].shape != (1, 1500, 128):
|
||||
raise ValueError('Preparation did not produce a full60s context')
|
||||
if kwargs.get('repaint_mask') is not None:
|
||||
raise ValueError('Full planned composition must not silently become repaint')
|
||||
if not state.get('semantic_plan_present'):
|
||||
raise ValueError('Refuse unplanned conditioning')
|
||||
case = dict(recipe, name='prism_hook_v1_60', reference_audio=None,
|
||||
sampler={k: v for k, v in kwargs.items() if isinstance(v, (str, int, float, bool)) or v is None},
|
||||
tensor_shapes={k: list(v.shape) for k, v in arrays.items()})
|
||||
for key, array in arrays.items():
|
||||
np.save(out / (key + '.npy'), array)
|
||||
(out / 'case.json').write_text(json.dumps(case, indent=2))
|
||||
state.update(phase='prepared_boundary', audio_diffusion_called=False,
|
||||
boundary_sha256={p.name: hashlib.sha256(p.read_bytes()).hexdigest()
|
||||
for p in out.glob('*.npy')},
|
||||
tensor_shapes=case['tensor_shapes'], elapsed_seconds=time.monotonic() - started)
|
||||
record()
|
||||
raise BoundaryCaptured()
|
||||
|
||||
handler._mlx_run_diffusion = capture
|
||||
params = GenerationParams(caption=recipe['caption'], lyrics=recipe['lyrics'], instrumental=True,
|
||||
bpm=128, keyscale='D minor', timesignature='4', duration=60, inference_steps=8, seed=seed,
|
||||
thinking=True, dcw_enabled=False, use_cot_caption=False, use_cot_metas=False, use_cot_language=False)
|
||||
try:
|
||||
result = generate_music(handler, lm, params,
|
||||
GenerationConfig(batch_size=1, allow_lm_batch=False, use_random_seed=False, seeds=[seed], audio_format='wav'),
|
||||
save_dir=str(out / 'unused_audio'))
|
||||
except BoundaryCaptured:
|
||||
pass
|
||||
else:
|
||||
raise RuntimeError(f'Expected boundary capture, got success={result.success}: {result.error}')
|
||||
if state['phase'] != 'prepared_boundary':
|
||||
raise RuntimeError('No native-compatible tensor capture')
|
||||
print(json.dumps(state))
|
||||
except BaseException as error:
|
||||
state.update(phase='failed', error=repr(error), traceback=traceback.format_exc(), elapsed_seconds=time.monotonic() - started)
|
||||
record()
|
||||
raise
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,75 @@
|
||||
"""Shared musical intent for a candidate; this alone does not change prepared tensors."""
|
||||
import hashlib
|
||||
import json
|
||||
|
||||
VERSION = 'prism-hook-v1'
|
||||
IDENTITY = (
|
||||
'Instrumental polished K-pop and modern electronic game score. No vocals, singing or speech. '
|
||||
'128 BPM, D minor, 4/4. Tight electronic drums, punchy rubbery syncopated bass, '
|
||||
'crystal pluck arpeggios and bright glass synth leads. '
|
||||
)
|
||||
HOOK = (
|
||||
'Give this composition one distinctive, immediately memorable four-note rising synth hook '
|
||||
'with a catchy syncopated rhythm and a short answering phrase. Establish its recognizable '
|
||||
'melodic contour and rhythm early. Keep that SAME hook identity throughout this song: '
|
||||
'develop its rhythm, register, accompaniment and dynamics, then bring back the clear original '
|
||||
'phrase at each chorus. Make the motif specific to this composition, not a generic running '
|
||||
'arpeggio. Leave breathing space between hook statements. Do not replace it with unrelated '
|
||||
'lead melodies or mechanically repeat an unchanged loop. '
|
||||
)
|
||||
ROLES = {
|
||||
'initial': 'Introduce the hook clearly, develop a lighter verse, rise into a build and deliver a confident first chorus. ',
|
||||
'verse': 'Develop a lighter verse using recognizable fragments and a quieter call-and-response of the established hook over an active bass groove. Vary the accompaniment and leave space for the hook return. ',
|
||||
'prechorus': 'Build tension with shorter recognizable hook fragments, rising register and denser drum subdivisions. Aim the phrase toward the chorus; do not introduce a new main melody. ',
|
||||
'chorus': 'State the complete established hook prominently with its original contour and signature rhythm; answer it with the same phrase in a wider register and fuller bass/drums. Make this a clear melodic payoff, not merely louder texture. ',
|
||||
'bridge': 'Create contrast by reducing the arrangement and transforming the same hook into a spacious half-time rhythmic answer in compatible glass/pluck timbre. Retain its melodic identity and D minor harmony, then rebuild toward the original chorus hook. ',
|
||||
'outro': 'Return to the recognizable complete hook, answer and resolve its last phrase, then thin the arrangement into a gentle intentional conclusion. ',
|
||||
}
|
||||
TIMELINE = [
|
||||
('initial', 2, 'Clear hook introduction'),
|
||||
('verse', 8, 'Hook fragments and lighter call-and-response'),
|
||||
('prechorus', 4, 'Rising fragment development'),
|
||||
('chorus', 8, 'Complete hook and melodic payoff'),
|
||||
('bridge', 4, 'Contrasting transformation of same motif'),
|
||||
('chorus', 4, 'Recognizable original hook reprise'),
|
||||
('outro', 2, 'Motif answer and resolution'),
|
||||
]
|
||||
|
||||
|
||||
def spec(route_seed=None):
|
||||
if route_seed is not None and (type(route_seed) is not int or not 0 <= route_seed < 2**32):
|
||||
raise ValueError('route_seed must be an unsigned 32-bit integer')
|
||||
seed = 33602 if route_seed is None else int.from_bytes(
|
||||
hashlib.sha256(f'roadscore-sample-v1:{route_seed}:prepare:0'.encode()).digest()[:4], 'big')
|
||||
interior = 'This is a continuing interior section: preserve the pulse and hand off naturally without a terminal fade or silence. '
|
||||
prompts = {role: IDENTITY + HOOK + text + (interior if role not in ('initial', 'outro') else '')
|
||||
for role, text in ROLES.items()}
|
||||
cursor = 0
|
||||
timeline = []
|
||||
for role, bars, intent in TIMELINE:
|
||||
duration = bars * 4 * 60 / 128
|
||||
timeline.append({'role': role, 'bars': bars, 'start_seconds': cursor,
|
||||
'end_seconds': cursor + duration, 'intent': intent})
|
||||
cursor += duration
|
||||
result = {
|
||||
'version': VERSION, 'profile': 'prism', 'bpm': 128, 'keyscale': 'D minor', 'timesignature': '4',
|
||||
'duration': 60, 'seed': seed, 'seed_mode': 'fixed-reference' if route_seed is None else 'route-derived',
|
||||
'route_seed': route_seed, 'thinking': True, 'role_captions': prompts,
|
||||
'role_lyrics': {r: '[Instrumental]\n[' + ('Pre-Chorus' if r == 'prechorus' else r.title()) + ']'
|
||||
for r in ROLES if r != 'initial'},
|
||||
'caption': IDENTITY + HOOK + (
|
||||
'Compose a coherent one-minute arc: establish the hook, develop a lighter verse, '
|
||||
'build anticipation, reveal a full chorus, give a brief contrasting bridge variation, '
|
||||
'then reprise the original hook and resolve naturally. Each section should change '
|
||||
'the musical arrangement while retaining the song\'s identity. '),
|
||||
'lyrics': '[Instrumental]\n[Intro]\n[Verse]\n[Pre-Chorus]\n[Chorus]\n[Bridge]\n[Chorus]\n[Outro]',
|
||||
'desired_timeline': timeline,
|
||||
'timeline_status': '32-bar intent at 128 BPM, not verified generated section boundaries; never force cuts to these timestamps',
|
||||
'continuation_policy': 'role contracts only; no prepared continuation tensors or runtime deployment claimed',
|
||||
'uniqueness_limit': 'Distinct hook is musical intent, not a metric guarantee; fixed seeds intentionally reproduce a composition.',
|
||||
}
|
||||
result['caption'] += ('Target a 32-bar arc: 2-bar introduction, 8-bar verse, 4-bar build, '
|
||||
'8-bar chorus, 4-bar bridge variation, 4-bar chorus reprise and 2-bar resolution. ')
|
||||
result['role_lyrics']['initial'] = result['lyrics']
|
||||
result['spec_sha256'] = hashlib.sha256(json.dumps(result, sort_keys=True).encode()).hexdigest()
|
||||
return result
|
||||
@@ -0,0 +1,39 @@
|
||||
import unittest
|
||||
from prism_hook_spec import spec, IDENTITY, HOOK, ROLES
|
||||
|
||||
|
||||
class HookSpecTests(unittest.TestCase):
|
||||
def test_identity_and_hook_survive_every_role(self):
|
||||
candidate = spec()
|
||||
self.assertEqual(set(candidate['role_captions']), set(ROLES))
|
||||
for caption in candidate['role_captions'].values():
|
||||
self.assertTrue(caption.startswith(IDENTITY + HOOK))
|
||||
self.assertIn('128 BPM, D minor', caption)
|
||||
self.assertEqual(len(set(candidate['role_captions'].values())), 6)
|
||||
self.assertIn('complete established hook', candidate['role_captions']['chorus'])
|
||||
self.assertIn('transforming the same hook', candidate['role_captions']['bridge'])
|
||||
self.assertIn('resolve', candidate['role_captions']['outro'])
|
||||
|
||||
def test_real_arc_not_repeated_verse(self):
|
||||
candidate = spec()
|
||||
self.assertTrue(candidate['thinking'])
|
||||
self.assertEqual(candidate['lyrics'].count('[Chorus]'), 2)
|
||||
self.assertIn('[Bridge]', candidate['lyrics'])
|
||||
self.assertIn('[Pre-Chorus]', candidate['lyrics'])
|
||||
self.assertIn('no prepared continuation tensors', candidate['continuation_policy'])
|
||||
self.assertEqual(sum(x['bars'] for x in candidate['desired_timeline']), 32)
|
||||
self.assertEqual(candidate['desired_timeline'][-1]['end_seconds'], 60)
|
||||
self.assertIn('not verified', candidate['timeline_status'])
|
||||
|
||||
def test_seed_policy_reproducible_and_not_seed_shopping(self):
|
||||
self.assertEqual(spec()['seed'], 33602)
|
||||
self.assertEqual(spec(123), spec(123))
|
||||
self.assertNotEqual(spec(123)['seed'], spec(124)['seed'])
|
||||
self.assertEqual(spec(123)['caption'], spec(124)['caption'])
|
||||
for bad in (-1, 2**32, True, 1.5):
|
||||
with self.assertRaises(ValueError):
|
||||
spec(bad)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
unittest.main()
|
||||
Reference in New Issue
Block a user