mirror of
https://github.com/firestar5683/StarPilot.git
synced 2026-10-04 21:33:58 +08:00
Record first clean event Chestnut replay and add acceptance audit
This commit is contained in:
@@ -95,3 +95,15 @@ ACE is now the event default; explicit SA3 selection is preserved. Composer test
|
||||
Actual native Params report `Model=DrivingModel=rdf43`, version v15, real `IsOnroad=false`. The built-in model path uses the QCOM backend and does not select an external-GPU artifact merely because Chestnut is connected. Model Lab configuration must still be checked before any test.
|
||||
|
||||
The existing `selfdrive/test/process_replay/process_replay.py` supports isolated `modeld` execution in an `OpenpilotPrefix`, feeding road/wide camera frames and device/calibration/car state; it publishes modelV2, drivingModelData and cameraOdometry. The cached known route contains `fcamera.hevc`, `ecamera.hevc` and rlog. A bounded test should call that local harness directly and record execution times first alone, then during ACE generation. Do not use the CI model-replay report/upload entrypoint for private routes. No coexistence or live-driving test has been run.
|
||||
|
||||
### 45 W full preparation
|
||||
|
||||
The explicitly capped 45 W resident worker completed 112 seconds of accepted Prism audio in 530.441 seconds, including 38.190 seconds model load. Peak host RSS was 758,308 KiB; tracked allocation reached 5,789,487,104 bytes. Initial 28-second music required 13.651 seconds generation plus 6.206 seconds decode (0.709 RTF excluding cold compile). Warm continuation required 19.51–19.59 seconds generation plus 10.30–10.32 seconds decode per 28 new seconds (1.065–1.068 compute RTF; 30.12–30.42 seconds wall). This does not establish sustained faster-than-playback continuation.
|
||||
|
||||
Native muted replay `normal_1789795416` is running without continuous UI video encoding. Camera/path auditing and the overlay image remain enabled; the causal timing guard is unchanged. No event native pass is claimed before its final audit. Exact preparation metadata is preserved privately in `results/event_night_one/prism_power45/`.
|
||||
|
||||
### First clean event native replay — 45 W
|
||||
|
||||
`normal_1789795416` passed the instrumented native gate through final-segment EOF: 254.1 seconds captured, six accepted fresh generation jobs, zero accepted-music holds, zero underflows, zero emergency fallbacks, no worker failure, zero output flags and all blocks muted. Camera accepted 5,088 frames; path/lane drawing, navigation and ten curve activations were observed. All request timestamps were causal. Maximum source-clock drift was 24.671 ms. Continuous UI video encoding was disabled; the overlay snapshot and UI audit remain available. This single pass does not prove that encoding caused the earlier timing failure.
|
||||
|
||||
The local private evidence is `results/event_night_one/prism_power45/`: preparation metadata, native audit, overlay and captured score. The reusable post-run audit correctly rejects the earlier full-route failure and has three focused negative-evidence tests. It does not claim human musical approval, physical speaker/Bluetooth validation or modeld coexistence. Phase 2 has not yet begun; the native prerequisite is now met. No upload or push was performed.
|
||||
|
||||
@@ -1,3 +1,7 @@
|
||||
# Event checkpoint — clean 45 W native replay
|
||||
|
||||
Prism completed a full 254.1-second muted route on the event comma + Chestnut with six accepted fresh music jobs, no holds, no underflows, no worker failure, and clean causal timing. Camera, path, lanes, navigation and ten curve activations were recorded. See [event evidence](EVENT.md). This is an instrumented replay pass; physical listening, Bluetooth and modeld coexistence remain unverified. Community judging has not begun.
|
||||
|
||||
# Event checkpoint — RoadScore branch
|
||||
|
||||
Current event work and exact failures are recorded in [EVENT.md](EVENT.md). Source is migrated. Official Chestnut validation and 30 W Prism preparation pass; native replay still has unresolved timing/ending-context failures. A 45 W fixed-decoder comparison passed 40/40 and full-model testing is starting. No event-baseline or judging-batch pass is claimed. All automated audio remains muted; no files have been pushed by this work.
|
||||
|
||||
@@ -0,0 +1,55 @@
|
||||
"""Read-only post-run acceptance checks. Never used to drive the score."""
|
||||
import argparse
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def audit(run):
|
||||
def read(name):
|
||||
p = run / name
|
||||
return json.loads(p.read_text()) if p.exists() else {}
|
||||
|
||||
def rows(name):
|
||||
p = run / name
|
||||
return [json.loads(line) for line in p.read_text().splitlines() if line] if p.exists() else []
|
||||
|
||||
launch, summary, bridge = (read(n) for n in ('launch.json', 'summary.json', 'bridge.json'))
|
||||
trace, jobs, blocks, ui = (rows(n) for n in ('trace.jsonl', 'jobs.jsonl', 'host_audio.jsonl', 'ui_audit.jsonl'))
|
||||
accepted = [g for g in summary.get('generation', []) if g.get('quality_accepted') is True]
|
||||
violations = [j for j in jobs if any(t > j.get('cutoff_ns', -1) for t in j.get('input_times', {}).values())]
|
||||
last_ui = ui[-1] if ui else {}
|
||||
checks = {
|
||||
'native_eof': launch.get('end_reason') == 'native final segment exhausted',
|
||||
'bridge_clean': bool(bridge) and bridge.get('failure') is None,
|
||||
'fresh_generation': bool(accepted),
|
||||
'captured_audio': summary.get('audio_seconds', 0) > 0,
|
||||
'no_underflow': summary.get('underflows') == 0,
|
||||
'no_emergency_fallback': summary.get('emergency_fallbacks') == 0,
|
||||
'worker_healthy_throughout': bool(trace) and not any(t.get('worker_failed') for t in trace),
|
||||
'causal_requests': bool(jobs) and not violations,
|
||||
'muted_launch': launch.get('muted') is True,
|
||||
'muted_blocks': bool(blocks) and all(b.get('muted') is True for b in blocks),
|
||||
'no_output_flags': bool(blocks) and all(not b.get('portaudio_status') for b in blocks),
|
||||
'camera': last_ui.get('accepted_camera_frames', 0) > 0,
|
||||
'path': last_ui.get('nonempty_path_draws', 0) > 0,
|
||||
'lanes': last_ui.get('nonempty_lane_draws', 0) > 0,
|
||||
}
|
||||
return {
|
||||
'checks': checks, 'instrumented_pass': all(checks.values()),
|
||||
'audio_seconds': summary.get('audio_seconds'), 'accepted_jobs': len(accepted),
|
||||
'accepted_music_holds': summary.get('accepted_music_holds'),
|
||||
'navigation_present': any(t.get('nav', {}).get('valid') for t in trace),
|
||||
'curve_activations': len({t['activation'] for t in trace if t.get('activation') is not None and t.get('kind') == 'curve'}),
|
||||
'input_time_violations': len(violations), 'ui_last': last_ui,
|
||||
'bridge': bridge,
|
||||
'limitations': 'Instrumented evidence only; musical quality, speakers, Bluetooth and live driving are not verified.',
|
||||
}
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument('run', type=Path)
|
||||
args = parser.parse_args()
|
||||
result = audit(args.run)
|
||||
(args.run / 'event_replay_audit.json').write_text(json.dumps(result, indent=2))
|
||||
print(json.dumps(result, indent=2))
|
||||
@@ -0,0 +1,34 @@
|
||||
import json
|
||||
from pathlib import Path
|
||||
import tempfile
|
||||
import unittest
|
||||
from event_replay_audit import audit
|
||||
|
||||
|
||||
class AuditTests(unittest.TestCase):
|
||||
def test_missing_evidence_cannot_pass(self):
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
self.assertFalse(audit(Path(directory))['instrumented_pass'])
|
||||
|
||||
def test_failed_generation_is_not_fresh_music(self):
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
(root / 'summary.json').write_text(json.dumps({'generation': [
|
||||
{'quality_rejected': True, 'quality_attempts': []}]}))
|
||||
result = audit(root)
|
||||
self.assertFalse(result['checks']['fresh_generation'])
|
||||
self.assertEqual(result['accepted_jobs'], 0)
|
||||
|
||||
def test_future_request_and_single_unmuted_block_fail(self):
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
(root / 'jobs.jsonl').write_text(json.dumps({'cutoff_ns': 10, 'input_times': {'modelV2': 11}})+'\n')
|
||||
(root / 'host_audio.jsonl').write_text('{"muted": true}\n{"muted": false}\n')
|
||||
result = audit(root)
|
||||
self.assertEqual(result['input_time_violations'], 1)
|
||||
self.assertFalse(result['checks']['causal_requests'])
|
||||
self.assertFalse(result['checks']['muted_blocks'])
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
unittest.main()
|
||||
Reference in New Issue
Block a user