Specify coherent Prism hook development and isolated host planning

This commit is contained in:
firestar5683
2026-09-19 13:22:33 -07:00
parent 1132379613
commit 299ed93ae3
4 changed files with 347 additions and 0 deletions
@@ -0,0 +1,88 @@
# Prism hook candidate — proposal, not deployed
Base: `06c8e7ffe8` (isolated GOLD60 diagnostic), itself based on `a12d1b7854`.
Protected GOLD and user-approved planned60 assets remain unchanged.
## Musical intent
`prism_hook_spec.py` defines one shared full Prism identity and hook contract for
initial, verse, build/prechorus, chorus, bridge and outro. Every role retains
128 BPM, D minor, crystal pluck/glass lead timbre, syncopated bass and tight drums.
The hook has a recognizable four-note rising contour, signature syncopated rhythm
and short answering phrase. Verses expose fragments; builds intensify them;
choruses state the complete original; the bridge transforms rhythm/register while
keeping identity; the return restores the original phrase and outro resolves it.
This replaces the old generic restrained-verse emphasis and keeps specific hook
and instrumental instructions present across all role prompts. It cannot guarantee
memorability or uniqueness without listening. Fixed seeds intentionally reproduce
the same composition; route-derived seeds distinguish comparable route runs.
## Proposed first demo composition
The first candidate is one coherent LM-planned60-second composition, not repeated
8-second-context verses. At the requested128BPM,32bars occupy60seconds:
| Target time | Role | Musical development |
|---|---|---|
|0–3.75s|Introduction,2bars|State recognizable hook|
|3.75–18.75s|Verse,8bars|Lighter fragments, breathing space and answer|
|18.75–26.25s|Build,4bars|Rising register/subdivisions of same motif|
|26.25–41.25s|Chorus,8bars|Complete melody, fuller bass and drums|
|41.25–48.75s|Bridge,4bars|Spacious contrasting transformation|
|48.75–56.25s|Chorus reprise,4bars|Return the unmistakable original|
|56.25–60s|Outro,2bars|Answer and resolution|
These are prompt targets, **not verified generated timestamps**. Do not force
audio cuts or treat these times as detected beats/section boundaries. An actual
generated performance may not follow them. The fixed narrative is independent
of recorded future route events; operator excerpt selection does not grant
runtime access to future telemetry. Current delivered conditions may later inform
causal adaptation, but this candidate does not implement it.
Prefer a60-second route excerpt for this first musical test. A90–120second route
requires a newly prepared longer semantic plan, larger native shape validation,
or a tested continuation strategy. Do not loop this60s piece, stretch its timing,
or restart its initial plan and call that coherent longer-form composition.
Presentation is subordinate: proposed cue/engagement lanes must not overwrite
the hook or rhythmic phase. If a cue conflicts, omit/simplify it. This composition
lane does not add signal, curve, engagement or other sonification. The master
will combine the route, composition, one signal motif, one curve treatment and
one engagement transition in the operator proposal before new generation.
## Actual preparation status and path
**Prompt specification only. No new conditioning tensors, semantic codes, native
audio, or deployed behavior yet.** Three CPU specification tests pass. Local Mac
ACE packages and existing9.4GB model assets are available; no model downloads are
needed. New preparation is paused for the combined demo-plan review.
`prepare_prism_hook.py` is an opt-in Mac preparation tool. Default mode saves the
specification only. `--prepare` initializes the existing official ACE/1.7BLM,
uses `thinking=True` and fixed seed33602 (or deterministic route-derived seed),
records actual returned semantic codes/seed/LM costs, then captures full60s
encoder/context tensors immediately before DiT diffusion. It refuses repaint,
missing semantic planning, nonfinite/wrong-duration tensors and an unavailable
MLX interception path. It never executes native hardware or intentionally
generates audio. Candidate output uses a new private directory; protected files
and active profiles are never replaced. Exact prompt and tensor hashes are saved.
The full composition caption/section text are what this first preparation feeds
into the model. Per-role continuation contracts are saved for future use but
**are not prepared role tensors** and are not wired into current worker startup.
Editing these strings cannot alter deployed prepared embeddings. All role-based
continuation restoration remains future work after the full-plan musical test.
Historical preparation recorded12.4s for the old60s LM plan, plus text/conditioning
and cold model load; budget minutes rather than promise that runtime. Longer
hook text may change encoder length/cost. Tensor output should be a few MiB
(1500×128 float32 context≈0.73MiB plus variable-length encoder≈1–severalMiB), not
new multi-GB weights. Model loading can consume substantial RAM; current observed
diskfree was≈7GB. The approved past run is the playback fallback while preparing.
After review: run one captured plan, check its actual codes/tensors and provenance;
have hardware owner review a candidate-capable isolated native harness (the GOLD
probe intentionally pins original asset hashes and must reject this new case);
then generate one fixed-seed sample with rawgain processing, inspect/listen,
and only afterward integrate subordinate presentation. No seed auditions or
per-route cherry-picking. Master/user musical acceptance remains the gate.
@@ -0,0 +1,145 @@
"""Mac-only semantic planning/tensor capture. Stops BEFORE audio diffusion or decoding."""
import argparse
import hashlib
import json
import os
from pathlib import Path
import random
import sys
import time
import traceback
from prism_hook_spec import spec
class BoundaryCaptured(BaseException):
pass
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('--assets-root', type=Path, required=True, help='Existing RoadScore assets; read only')
parser.add_argument('--output-root', type=Path, required=True, help='New private demo artifacts only')
parser.add_argument('--route-seed', type=int)
parser.add_argument('--prepare', action='store_true', help='Run one host semantic plan; no audio diffusion')
args = parser.parse_args()
recipe = spec(args.route_seed)
out = args.output_root.resolve() / f'prism_hook_v1_{time.time_ns()}'
out.mkdir(parents=True, exist_ok=False)
(out / 'hook_spec.json').write_text(json.dumps(recipe, indent=2))
state = {'phase': 'spec_only', 'output': str(out), 'spec_sha256': recipe['spec_sha256'],
'audio_generated': False, 'native_deployed': False, 'continuation_tensors_prepared': False}
def record():
(out / 'preparation.json').write_text(json.dumps(state, indent=2))
record()
if not args.prepare:
print(json.dumps(state))
return
if sys.platform != 'darwin':
raise RuntimeError('This host-preparation entry point is Mac-only; no device execution')
base = args.assets_root.resolve() / 'experiments/composition_20260916'
if not (base / 'models/ace/checkpoints').is_dir():
raise FileNotFoundError('Existing local ACE checkpoints missing; no download attempted')
os.environ.update(HF_HOME=str(base / 'cache/hf'), HF_HUB_OFFLINE='1', HF_HUB_DISABLE_TELEMETRY='1',
TOKENIZERS_PARALLELISM='false', PYTORCH_ENABLE_MPS_FALLBACK='1')
started = time.monotonic()
state.update(phase='loading', started_wall=time.time())
record()
try:
import numpy as np
import torch
import mlx.core as mx
from acestep.handler import AceStepHandler
from acestep.llm_inference import LLMHandler
from acestep.inference import GenerationParams, GenerationConfig, generate_music
seed = recipe['seed']
random.seed(seed)
np.random.seed(seed)
torch.manual_seed(seed)
mx.random.seed(seed)
handler, lm = AceStepHandler(), LLMHandler()
status, ok = handler.initialize_service(str(base / 'models/ace'), config_path='acestep-v15-turbo',
device='mps', use_mlx_dit=True, offload_to_cpu=True, offload_dit_to_cpu=False)
if not ok:
raise RuntimeError(status)
if not handler.use_mlx_dit or handler.mlx_decoder is None:
raise RuntimeError('MLX boundary unavailable; refuse fallback audio generation')
status, ok = lm.initialize(str(base / 'models/ace/checkpoints'), 'acestep-5Hz-lm-1.7B', backend='mlx', device='mps')
if not ok:
raise RuntimeError(status)
original_plan = lm.generate_with_stop_condition
def record_plan(*values, **kwargs):
result = original_plan(*values, **kwargs)
codes = result.get('audio_codes', '')
if not result.get('success') or not codes:
raise RuntimeError('Semantic planner did not return audio codes')
(out / 'semantic_plan.json').write_text(json.dumps({
'caption': kwargs.get('caption'), 'lyrics': kwargs.get('lyrics'),
'seeds': kwargs.get('seeds'), 'infer_type': kwargs.get('infer_type'),
'temperature': kwargs.get('temperature'), 'cfg_scale': kwargs.get('cfg_scale'),
'audio_codes': codes, 'metadata': result.get('metadata'),
'time_costs': result.get('extra_outputs', {}).get('time_costs'),
}, indent=2))
state['semantic_plan_present'] = True
return result
lm.generate_with_stop_condition = record_plan
state.update(phase='planning', load_seconds=time.monotonic() - started)
record()
def capture(*unused, **kwargs):
required = ('encoder_hidden_states', 'encoder_attention_mask', 'context_latents')
arrays = {}
for key in required:
tensor = kwargs.get(key)
if not isinstance(tensor, torch.Tensor):
raise ValueError(f'Missing prepared boundary tensor: {key}')
array = tensor.detach().float().cpu().numpy()
if not np.isfinite(array).all():
raise ValueError(f'Non-finite {key}')
arrays[key] = array
if arrays['context_latents'].shape != (1, 1500, 128):
raise ValueError('Preparation did not produce a full60s context')
if kwargs.get('repaint_mask') is not None:
raise ValueError('Full planned composition must not silently become repaint')
if not state.get('semantic_plan_present'):
raise ValueError('Refuse unplanned conditioning')
case = dict(recipe, name='prism_hook_v1_60', reference_audio=None,
sampler={k: v for k, v in kwargs.items() if isinstance(v, (str, int, float, bool)) or v is None},
tensor_shapes={k: list(v.shape) for k, v in arrays.items()})
for key, array in arrays.items():
np.save(out / (key + '.npy'), array)
(out / 'case.json').write_text(json.dumps(case, indent=2))
state.update(phase='prepared_boundary', audio_diffusion_called=False,
boundary_sha256={p.name: hashlib.sha256(p.read_bytes()).hexdigest()
for p in out.glob('*.npy')},
tensor_shapes=case['tensor_shapes'], elapsed_seconds=time.monotonic() - started)
record()
raise BoundaryCaptured()
handler._mlx_run_diffusion = capture
params = GenerationParams(caption=recipe['caption'], lyrics=recipe['lyrics'], instrumental=True,
bpm=128, keyscale='D minor', timesignature='4', duration=60, inference_steps=8, seed=seed,
thinking=True, dcw_enabled=False, use_cot_caption=False, use_cot_metas=False, use_cot_language=False)
try:
result = generate_music(handler, lm, params,
GenerationConfig(batch_size=1, allow_lm_batch=False, use_random_seed=False, seeds=[seed], audio_format='wav'),
save_dir=str(out / 'unused_audio'))
except BoundaryCaptured:
pass
else:
raise RuntimeError(f'Expected boundary capture, got success={result.success}: {result.error}')
if state['phase'] != 'prepared_boundary':
raise RuntimeError('No native-compatible tensor capture')
print(json.dumps(state))
except BaseException as error:
state.update(phase='failed', error=repr(error), traceback=traceback.format_exc(), elapsed_seconds=time.monotonic() - started)
record()
raise
if __name__ == '__main__':
main()
@@ -0,0 +1,75 @@
"""Shared musical intent for a candidate; this alone does not change prepared tensors."""
import hashlib
import json
VERSION = 'prism-hook-v1'
IDENTITY = (
'Instrumental polished K-pop and modern electronic game score. No vocals, singing or speech. '
'128 BPM, D minor, 4/4. Tight electronic drums, punchy rubbery syncopated bass, '
'crystal pluck arpeggios and bright glass synth leads. '
)
HOOK = (
'Give this composition one distinctive, immediately memorable four-note rising synth hook '
'with a catchy syncopated rhythm and a short answering phrase. Establish its recognizable '
'melodic contour and rhythm early. Keep that SAME hook identity throughout this song: '
'develop its rhythm, register, accompaniment and dynamics, then bring back the clear original '
'phrase at each chorus. Make the motif specific to this composition, not a generic running '
'arpeggio. Leave breathing space between hook statements. Do not replace it with unrelated '
'lead melodies or mechanically repeat an unchanged loop. '
)
ROLES = {
'initial': 'Introduce the hook clearly, develop a lighter verse, rise into a build and deliver a confident first chorus. ',
'verse': 'Develop a lighter verse using recognizable fragments and a quieter call-and-response of the established hook over an active bass groove. Vary the accompaniment and leave space for the hook return. ',
'prechorus': 'Build tension with shorter recognizable hook fragments, rising register and denser drum subdivisions. Aim the phrase toward the chorus; do not introduce a new main melody. ',
'chorus': 'State the complete established hook prominently with its original contour and signature rhythm; answer it with the same phrase in a wider register and fuller bass/drums. Make this a clear melodic payoff, not merely louder texture. ',
'bridge': 'Create contrast by reducing the arrangement and transforming the same hook into a spacious half-time rhythmic answer in compatible glass/pluck timbre. Retain its melodic identity and D minor harmony, then rebuild toward the original chorus hook. ',
'outro': 'Return to the recognizable complete hook, answer and resolve its last phrase, then thin the arrangement into a gentle intentional conclusion. ',
}
TIMELINE = [
('initial', 2, 'Clear hook introduction'),
('verse', 8, 'Hook fragments and lighter call-and-response'),
('prechorus', 4, 'Rising fragment development'),
('chorus', 8, 'Complete hook and melodic payoff'),
('bridge', 4, 'Contrasting transformation of same motif'),
('chorus', 4, 'Recognizable original hook reprise'),
('outro', 2, 'Motif answer and resolution'),
]
def spec(route_seed=None):
if route_seed is not None and (type(route_seed) is not int or not 0 <= route_seed < 2**32):
raise ValueError('route_seed must be an unsigned 32-bit integer')
seed = 33602 if route_seed is None else int.from_bytes(
hashlib.sha256(f'roadscore-sample-v1:{route_seed}:prepare:0'.encode()).digest()[:4], 'big')
interior = 'This is a continuing interior section: preserve the pulse and hand off naturally without a terminal fade or silence. '
prompts = {role: IDENTITY + HOOK + text + (interior if role not in ('initial', 'outro') else '')
for role, text in ROLES.items()}
cursor = 0
timeline = []
for role, bars, intent in TIMELINE:
duration = bars * 4 * 60 / 128
timeline.append({'role': role, 'bars': bars, 'start_seconds': cursor,
'end_seconds': cursor + duration, 'intent': intent})
cursor += duration
result = {
'version': VERSION, 'profile': 'prism', 'bpm': 128, 'keyscale': 'D minor', 'timesignature': '4',
'duration': 60, 'seed': seed, 'seed_mode': 'fixed-reference' if route_seed is None else 'route-derived',
'route_seed': route_seed, 'thinking': True, 'role_captions': prompts,
'role_lyrics': {r: '[Instrumental]\n[' + ('Pre-Chorus' if r == 'prechorus' else r.title()) + ']'
for r in ROLES if r != 'initial'},
'caption': IDENTITY + HOOK + (
'Compose a coherent one-minute arc: establish the hook, develop a lighter verse, '
'build anticipation, reveal a full chorus, give a brief contrasting bridge variation, '
'then reprise the original hook and resolve naturally. Each section should change '
'the musical arrangement while retaining the song\'s identity. '),
'lyrics': '[Instrumental]\n[Intro]\n[Verse]\n[Pre-Chorus]\n[Chorus]\n[Bridge]\n[Chorus]\n[Outro]',
'desired_timeline': timeline,
'timeline_status': '32-bar intent at 128 BPM, not verified generated section boundaries; never force cuts to these timestamps',
'continuation_policy': 'role contracts only; no prepared continuation tensors or runtime deployment claimed',
'uniqueness_limit': 'Distinct hook is musical intent, not a metric guarantee; fixed seeds intentionally reproduce a composition.',
}
result['caption'] += ('Target a 32-bar arc: 2-bar introduction, 8-bar verse, 4-bar build, '
'8-bar chorus, 4-bar bridge variation, 4-bar chorus reprise and 2-bar resolution. ')
result['role_lyrics']['initial'] = result['lyrics']
result['spec_sha256'] = hashlib.sha256(json.dumps(result, sort_keys=True).encode()).hexdigest()
return result
@@ -0,0 +1,39 @@
import unittest
from prism_hook_spec import spec, IDENTITY, HOOK, ROLES
class HookSpecTests(unittest.TestCase):
def test_identity_and_hook_survive_every_role(self):
candidate = spec()
self.assertEqual(set(candidate['role_captions']), set(ROLES))
for caption in candidate['role_captions'].values():
self.assertTrue(caption.startswith(IDENTITY + HOOK))
self.assertIn('128 BPM, D minor', caption)
self.assertEqual(len(set(candidate['role_captions'].values())), 6)
self.assertIn('complete established hook', candidate['role_captions']['chorus'])
self.assertIn('transforming the same hook', candidate['role_captions']['bridge'])
self.assertIn('resolve', candidate['role_captions']['outro'])
def test_real_arc_not_repeated_verse(self):
candidate = spec()
self.assertTrue(candidate['thinking'])
self.assertEqual(candidate['lyrics'].count('[Chorus]'), 2)
self.assertIn('[Bridge]', candidate['lyrics'])
self.assertIn('[Pre-Chorus]', candidate['lyrics'])
self.assertIn('no prepared continuation tensors', candidate['continuation_policy'])
self.assertEqual(sum(x['bars'] for x in candidate['desired_timeline']), 32)
self.assertEqual(candidate['desired_timeline'][-1]['end_seconds'], 60)
self.assertIn('not verified', candidate['timeline_status'])
def test_seed_policy_reproducible_and_not_seed_shopping(self):
self.assertEqual(spec()['seed'], 33602)
self.assertEqual(spec(123), spec(123))
self.assertNotEqual(spec(123)['seed'], spec(124)['seed'])
self.assertEqual(spec(123)['caption'], spec(124)['caption'])
for bad in (-1, 2**32, True, 1.5):
with self.assertRaises(ValueError):
spec(bad)
if __name__ == '__main__':
unittest.main()