Ford: verify unwind timing and reject incompatible offline replays

Compare v7 and recorded v6 unwind instructions on a5 using both common request levels and own-peak thresholds. Preserve the later C0 release and earlier C1 release in the report without claiming a physical improvement.

Exercise saturated release through the actual sender for both signs and every send phase. Require original model clocks for current replay extracts and reject current candidates in the historical unchanged-C1 damping replay. Record the complete 585-test suite and refreshed stress evidence.

Assisted-by: OpenAI Codex
This commit is contained in:
Isaac Barham
2026-09-08 13:53:15 -04:00
parent 6f83b17457
commit 4bd841eccd
7 changed files with 1544 additions and 8 deletions
+54 -1
View File
@@ -36,7 +36,7 @@ maneuver test reference is explicitly unsupported and disengages this candidate.
## Offline evidence
The broad run passed **572 tests and 9,146 subtests**, with 178 inherited safety
The broad run passed **585 tests and 9,146 subtests**, with 178 inherited safety
cases skipped as inapplicable. Tests of the removed yaw forecast were retired;
new tests cover actual model clocks, a nonconstant-speed trajectory, shared-point
sampling, the distance floor, endpoint holding, heading unwrap, source selection,
@@ -73,5 +73,58 @@ No root cause or physical fix is proven by frozen inputs. The latest a5 road
logs used 100 Hz sends, whereas the immediately preceding code revision already
changed to 20 Hz; a comparison against that drive also includes the cadence change.
## Unwind comparison
The follow-up a5 comparison includes all seven clearly separated tight turns and
four completed ordinary bends identified in the full-route command plot. A fifth
bend runs into the next turn and is excluded from the summary. All scored windows
remain continuously paired-active, without control gaps over 30 ms. The tight
turns contain driver input; this compares instructions on frozen inputs, not
unassisted tracking or hypothetical truck motion.
Measure the first time each instruction falls below the same fixed level on
exit and stays below for 100 ms. All seven tight turns exceed these levels:
| Instruction | Candidate minus recorded clearance time |
| --- | --- |
| C0 below 0.5 m into the turn | 0.10 s later median; five later by 0.010.46 s, two unchanged |
| C1 below 0.1 rad into the turn | 0.10 s earlier median; all seven 0.040.17 s earlier |
These are control-publication times. Actual packet timing also includes the
20 Hz send phase; shifts of only a few tens of milliseconds should not be
interpreted as equally precise changes at the PSCM.
The report also compares half of each command's own peak: C1 reaches that level
0.25 s earlier in the median tight turn and C0 0.12 s later. Those normalized
crossings have different absolute thresholds when amplitudes differ. Comparing
half the smaller peak at an identical level instead gives C0 earlier in one
turn and later in six. No single threshold captures the whole release waveform.
For the ordinary bend around 108 s, C0 reaches half of its own peak 0.20 s later;
C1 reaches half-peak 0.24 s earlier. The larger C0 takes 0.95 s longer to fall
below the same 0.05 m threshold. On the last tight turn, C0 clears 0.05 m 0.23 s
later. Near-zero thresholds are sensitive to small residual model offsets; the
complete report retains 90%, 50%, 10%, common-level and near-zero crossings
rather than treating any one threshold as a physical success criterion.
The plotted candidate target and command largely coincide during these exits:
the longer C0 tail comes from the selected model point continuing to ask for
lateral offset, rather than a retained yaw correction. C1 can release sooner
because it now follows model orientation directly. These results do not show
that every command unwinds earlier, or that the vehicle will unwind earlier.
Ten additional end-to-end sender cases cover both signs and all five CAN send
phases. After a full-cap turn, a zero model target starts reducing both states on
the first control update and appears on the next scheduled CAN message, 040 ms
later at nominal 100 Hz calculation. There is no additional release hold. The
unchanged slew itself takes 1.28 s of updates to clear 5.11 m C0 and 1.00 s to
clear 0.5 rad C1, plus scheduling to transmit zero. These are software timing
checks, not measured PSCM dynamics. The [unwind report](ford_model_points_unwind.json)
records the episode timings, method and source hashes.
Older offline utilities now fail explicitly when their input lacks original
model clocks or their historical unchanged-C1 comparison cannot support v7.
This prevents missing model inputs from producing an all-invalid apparent match.
`calibration_approved=false` remains. Installation instructions and restore
behavior are in the [drive-test guide](ford_model_action_drive_test.md).
File diff suppressed because it is too large Load Diff
+243 -2
View File
@@ -4,7 +4,7 @@
"scope": "Model-point construction only; physical tracking not validated.",
"calibration_approved": false,
"tests": {
"passed": 572,
"passed": 585,
"subtests_passed": 9146,
"inapplicable_skips": 178
},
@@ -13,7 +13,11 @@
"openpilot/selfdrive/controls/controlsd.py": "bea720b8ec68d6b8a6ad4376ae37324e2404833fb13097a8ab7710a1cedf37a5",
"tools/ford_pscm_lab/stress_model_action.py": "0221f85ea2f5757d3b54d077ed8be1faaa6a5d2ed8371de1137f3b6f3d70bf47",
"openpilot/sunnypilot/sunnylink/settings_ui_src/pages/vehicle.yaml": "ec5e0d9023f7260ae619a4cc6ec538d39053ac13e4515bd0de7d854b79d8abf6",
"openpilot/sunnypilot/sunnylink/settings_ui.json": "62e5178a8deee32ef5a88026094cedde359130741688e8a4bcfa82155ef4cca3"
"openpilot/sunnypilot/sunnylink/settings_ui.json": "62e5178a8deee32ef5a88026094cedde359130741688e8a4bcfa82155ef4cca3",
"tools/ford_pscm_lab/model_action_replay.py": "95d546cc87c065fa7d581b41382e1ab78bacc4030931b81388cb69c61adf32f7",
"tools/ford_pscm_lab/damping_replay.py": "1c253b99342d8113739b0a0551ca66ec6ced9e0aa96338ce6cb42e09c04178a4",
"tools/ford_pscm_lab/test_model_action_replay.py": "7aa01cbf197155175215aa4494cfc909029526fbacad0a1e1a2a214583064c4b",
"openpilot/selfdrive/controls/tests/test_ford_model_action_cadence.py": "d647e33c83581257f2662c3fec725bb9cd245fb48d579c2e38388318592a2690"
},
"routes": {
"a5": {
@@ -387,5 +391,242 @@
"/Users/ibpersonal/.codex/worktrees/1a1c/sunnypilot/.cache/ford_model_points/9e_models.npz": "acacf8f0d16fd986b6d74e4a02d704fa5b5b3327fe66dcbe149806f2a5bc3d65"
}
}
},
"stress": {
"seed": 20260908,
"random_cycles": 200000,
"mirrored_core_updates": 200000,
"invalid_or_inactive_resets": 3537,
"field_boundary_cases": 18138,
"float32_can_round_trips": 218138,
"analytic_targets_scalar_slew_and_mirror_checks_pass": true,
"valid_yaw_does_not_affect_targets_checked": true,
"shared_model_point_checked": true,
"direct_raw_float32_packing_matches_host_output": true,
"max_continuous_step_c0_c1": [
0.40000000000000147,
0.05000000000000002
],
"calibration_approved": false,
"scope": "Numerical construction only; no PSCM response or closed-loop performance claims.",
"opendbc_import_head": "87ca78e6e641eefb2d654f260a6ab08df3058bd5",
"source_sha256": {
"/Users/ibpersonal/.codex/worktrees/1a1c/sunnypilot/tools/ford_pscm_lab/stress_model_action.py": "0221f85ea2f5757d3b54d077ed8be1faaa6a5d2ed8371de1137f3b6f3d70bf47",
"/Users/ibpersonal/.codex/worktrees/1a1c/sunnypilot/openpilot/selfdrive/controls/lib/ford_model_action.py": "e7502ab68d04aef52edcffcef4b1d42c04e9008ab6892789aac3a865c100cb6d",
"/Users/ibpersonal/.codex/worktrees/1a1c/sunnypilot/tools/ford_pscm_lab/model_action_replay.py": "95d546cc87c065fa7d581b41382e1ab78bacc4030931b81388cb69c61adf32f7"
}
},
"unwind": {
"report": "docs/ford_model_points_unwind.json",
"sha256": "7395a4e66a275fd72d744f5a8244f1b9c8840bb7efff2eec225f15b9c033cd6e",
"full_cap_release_send_phase_cases": 10,
"summary": {
"tight": {
"episodes": 7,
"c0": {
"below_90pct_s": {
"n": 7,
"median_s": 0.0,
"min_s": -0.12269046900000546,
"max_s": 0.04266988000000538,
"earlier": 2,
"later": 2
},
"below_50pct_s": {
"n": 7,
"median_s": 0.12129962300002717,
"min_s": 0.0848213179999675,
"max_s": 0.8246327600000001,
"earlier": 0,
"later": 7
},
"below_10pct_s": {
"n": 7,
"median_s": 0.1376267710000434,
"min_s": 0.0,
"max_s": 0.45046578099999124,
"earlier": 0,
"later": 6
},
"near_zero_s": {
"n": 5,
"median_s": 0.1399018539999588,
"min_s": -0.5224231490000193,
"max_s": 1.9510819790000085,
"earlier": 1,
"later": 4
},
"common_half_s": {
"n": 7,
"median_s": 0.07685077599995793,
"min_s": -0.2591174110000338,
"max_s": 0.5661551039999893,
"earlier": 1,
"later": 6
},
"fixed_clearance_s": {
"n": 7,
"median_s": 0.10080577899998389,
"min_s": 0.0,
"max_s": 0.4609723959999883,
"earlier": 0,
"later": 5
}
},
"c1": {
"below_90pct_s": {
"n": 7,
"median_s": -0.3856273459999784,
"min_s": -0.5246007750000103,
"max_s": 0.5545458859999997,
"earlier": 6,
"later": 1
},
"below_50pct_s": {
"n": 7,
"median_s": -0.2540558009999927,
"min_s": -0.46442294700000275,
"max_s": -0.21144918100003451,
"earlier": 7,
"later": 0
},
"below_10pct_s": {
"n": 6,
"median_s": -0.15774414349999688,
"min_s": -0.20159088100001554,
"max_s": 0.0,
"earlier": 5,
"later": 0
},
"near_zero_s": {
"n": 5,
"median_s": 0.0,
"min_s": -0.148710938000022,
"max_s": 1.5534367700000047,
"earlier": 2,
"later": 2
},
"common_half_s": {
"n": 7,
"median_s": -0.2528352759999848,
"min_s": -0.45502116000000115,
"max_s": -0.15168304499999863,
"earlier": 7,
"later": 0
},
"fixed_clearance_s": {
"n": 7,
"median_s": -0.10267424500000288,
"min_s": -0.1688626670000417,
"max_s": -0.040530361999969955,
"earlier": 7,
"later": 0
}
}
},
"ordinary": {
"episodes": 4,
"c0": {
"below_90pct_s": {
"n": 4,
"median_s": 0.49771416100000465,
"min_s": 0.2530698499999744,
"max_s": 3.195503137000003,
"earlier": 0,
"later": 4
},
"below_50pct_s": {
"n": 4,
"median_s": 0.3417062340000143,
"min_s": 0.19507780300000377,
"max_s": 0.4595077099999685,
"earlier": 0,
"later": 4
},
"below_10pct_s": {
"n": 4,
"median_s": 0.11352378200000146,
"min_s": 0.09310488600004874,
"max_s": 1.1537201169999776,
"earlier": 0,
"later": 4
},
"near_zero_s": {
"n": 4,
"median_s": 0.6988682280000234,
"min_s": 0.40513631699997177,
"max_s": 3.1885321799999815,
"earlier": 0,
"later": 4
},
"common_half_s": {
"n": 4,
"median_s": 1.0263607459999946,
"min_s": 0.7940629349999995,
"max_s": 2.846801904000017,
"earlier": 0,
"later": 4
},
"fixed_clearance_s": {
"n": 2,
"median_s": 1.97568246000003,
"min_s": 1.5486082510000188,
"max_s": 2.4027566690000413,
"earlier": 0,
"later": 2
}
},
"c1": {
"below_90pct_s": {
"n": 4,
"median_s": -0.0037014085000066643,
"min_s": -0.25255103900002496,
"max_s": 0.3999567310000316,
"earlier": 2,
"later": 1
},
"below_50pct_s": {
"n": 4,
"median_s": -0.17102558000001977,
"min_s": -0.2489269939999872,
"max_s": -0.046780566999984785,
"earlier": 4,
"later": 0
},
"below_10pct_s": {
"n": 4,
"median_s": -0.10130154349999998,
"min_s": -0.29764589699999533,
"max_s": 0.05121605200002932,
"earlier": 2,
"later": 1
},
"near_zero_s": {
"n": 4,
"median_s": -0.022888650000027155,
"min_s": -0.10325871399999187,
"max_s": 0.24679704999999785,
"earlier": 2,
"later": 1
},
"common_half_s": {
"n": 4,
"median_s": -0.22263683349999042,
"min_s": -0.2573995380000156,
"max_s": -0.09939024700003074,
"earlier": 4,
"later": 0
},
"fixed_clearance_s": {
"n": 2,
"median_s": -0.17009618150001415,
"min_s": -0.24998850800000127,
"max_s": -0.09020385500002703,
"earlier": 2,
"later": 0
}
}
}
}
}
}
@@ -106,3 +106,41 @@ def test_core_slew_per_second_and_actual_panda_acceptance():
assert wire['LatCtlPath_An_Actl'] == pytest.approx(-command.path_angle)
assert wire['LatCtlCurv_No_Actl'] == wire['LatCtlCrv_NoRate2_Actl'] == 0.
assert frames == list(range(0, 100, 5))
@pytest.mark.parametrize('sign', [-1., 1.])
@pytest.mark.parametrize('phase', range(5))
def test_saturated_turn_unwinds_on_first_update_and_next_scheduled_can_frame(sign, phase):
controller, cs = sender()
core = ModelActionController()
cc, sp = structs.CarControl(latActive=True), structs.CarControlSP()
parser = CANParser('ford_lincoln_base_pt', [('LateralMotionControl2', 0)], controller.CAN.main)
change_frame = 200 + phase
first_unwind = first_zero_c0 = first_zero_c1 = None
turn, released = straight(sign*10., sign), straight()
for frame in range(change_frame+135):
command = core.update(turn if frame < change_frame else released, sign*.1, speed=20., dt=.01)
if frame >= change_frame:
elapsed = (frame-change_frame+1)*.01
assert sign*core.c0 == pytest.approx(max(0., 5.11-4.*elapsed), abs=1e-10)
assert sign*core.c1 == pytest.approx(max(0., .5-.5*elapsed), abs=1e-10)
sp.fordLateralPath.valid = command.valid
sp.fordLateralPath.pathOffset, sp.fordLateralPath.pathAngle = command.path_offset, command.path_angle
_, packets = controller.update(cc.as_reader(), sp, cs, frame*10_000_000)
lateral = [p for p in packets if p[0] == 0x3d6]
if not lateral or frame < change_frame:
continue
parser.update([frame*10_000_000, lateral])
wire = parser.vl['LateralMotionControl2']
c0, c1 = -sign*wire['LatCtlPathOffst_L_Actl'], -sign*wire['LatCtlPath_An_Actl']
assert 0. <= c0 < 5.11 and 0. <= c1 < .5
assert wire['LatCtlCurv_No_Actl'] == wire['LatCtlCrv_NoRate2_Actl'] == 0.
if first_unwind is None:
first_unwind = frame
if c0 == 0. and first_zero_c0 is None:
first_zero_c0 = frame
if c1 == 0. and first_zero_c1 is None:
first_zero_c1 = frame
assert first_unwind == ((change_frame+4)//5)*5
assert first_zero_c0 == ((change_frame+127+4)//5)*5
assert first_zero_c1 == ((change_frame+99+4)//5)*5
+6 -4
View File
@@ -1,4 +1,4 @@
"""Compare pinned selected-action controllers or current code on complete rlogs.
"""Compare pinned historical selected-action controllers on complete rlogs.
Original controls publication times proxy computation time. Consumed model
timestamps are exact; carState is causal and carControl is matched within 5 ms.
@@ -74,10 +74,13 @@ def run(directory, output, baseline_version='v1', candidate_version='v2', window
directory, output = directory.resolve(), output.resolve()
if output == directory or directory in output.parents:
raise ValueError('Output must be outside the source route directory')
if candidate_version == 'current':
raise ValueError('This historical damping replay requires unchanged C1 and does not preserve model clocks; ' +
'select a pinned v2/v3/v4 candidate, or use a model-point replay for current code')
verify_dependency(DEPLOYMENT_OPENDBC)
revisions = {'v1': V1_REVISION, 'v2': V2_REVISION, 'v3': V3_REVISION, 'v4': V4_REVISION}
baseline_source = load_controller(revisions[baseline_version])
candidate_source = ford_model_action if candidate_version == 'current' else load_controller(revisions[candidate_version])
candidate_source = load_controller(revisions[candidate_version])
streams, models, sources, t0 = extract(directory)
controls, model = streams['controls'], streams['model']
t = controls['t']
@@ -149,8 +152,7 @@ def run(directory, output, baseline_version='v1', candidate_version='v2', window
'baseline_version': baseline_version, 'baseline_revision': revisions[baseline_version],
'candidate_version': candidate_version, 'candidate_revision': revisions.get(candidate_version, 'working_tree'),
'focus_windows': windows, 'baseline_source_sha256': baseline_source.source_sha256,
'candidate_source_sha256': (hashlib.sha256(Path(ford_model_action.__file__).read_bytes()).hexdigest()
if candidate_version == 'current' else candidate_source.source_sha256),
'candidate_source_sha256': candidate_source.source_sha256,
'source_rlog_sha256': sources, 'opendbc_head': DEPLOYMENT_OPENDBC,
'source_sha256': {str(p): hashlib.sha256(p.read_bytes()).hexdigest() for p in (Path(__file__), Path(ford_model_action.__file__))}}
output.mkdir(parents=True, exist_ok=True)
+19 -1
View File
@@ -5,6 +5,10 @@ command compatibility. The separate current adapter pass reconstructs input elig
from service records, never from candidate/baseline output validity. Neither
pass scores counterfactual motion. Source extracts and archived reports are
read-only; --output selects a separate destination.
Current comparisons require model_position_t and model_orientation_t arrays
preserved from the rlogs. Older extracts lack these clocks and must be replayed
using their archived tool revision; do not synthesize timestamps for them.
"""
import argparse
from collections import Counter
@@ -62,6 +66,20 @@ def table(raw, name):
return dict(zip(raw[name+'_names'], raw[name].T, strict=True))
def extract_models(raw):
clock_fields = ('model_position_t', 'model_orientation_t')
if any(name not in raw for name in clock_fields):
raise ValueError('Current replay requires original model_position_t and model_orientation_t; ' +
'use the archived replay revision for older extracts, or re-extract the original rlogs with both clocks')
paths = raw['model_paths']
if paths.ndim != 3 or paths.shape[1] != 4 or len(paths) == 0:
raise ValueError('Expected nonempty model_paths with four arrays per model')
if any(raw[name].shape != (len(paths), paths.shape[2]) for name in clock_fields):
raise ValueError('Model clocks must match the extracted model and point counts')
return [SimpleNamespace(position=SimpleNamespace(x=p[1], y=p[2], t=pt), orientation=SimpleNamespace(z=p[3], t=ht))
for p, pt, ht in zip(paths, raw[clock_fields[0]], raw[clock_fields[1]], strict=True)]
def sample(stream, query, *, nearest=False):
if len(stream['t']) == 0 or np.any(np.diff(stream['t']) < 0):
raise ValueError('Replay requires nonempty streams in original timestamp order')
@@ -122,10 +140,10 @@ def run(directory, output):
raise ValueError('Output must be outside the source route directory')
dependency = verify_dependency()
with np.load(directory/'route.npz', allow_pickle=False) as raw:
models = extract_models(raw)
streams = {name: table(raw, name) for name in ('controls', 'cs', 'cc', 'model', 'params', 'path')}
if len(raw['maneuver']):
raise ValueError('This extract cannot identify the selected maneuver service per cycle; use the integration tests for that source')
models = [SimpleNamespace(position=SimpleNamespace(x=p[1], y=p[2]), orientation=SimpleNamespace(z=p[3])) for p in raw['model_paths']]
with np.load(directory/'encoder_comparison.npz', allow_pickle=False) as archive:
baseline = {key: archive[key] for key in ('t', 'valid', 'action_heading')}
with np.load(directory/'pose_candidate/pose_replay.npz', allow_pickle=False) as pose:
@@ -5,6 +5,7 @@ import pytest
from openpilot.selfdrive.controls.lib.ford_path import FordPath
from tools.ford_pscm_lab import model_action_replay as replay
from tools.ford_pscm_lab import damping_replay
def test_service_sampling_keeps_original_gaps_and_never_pulls_future_inputs():
@@ -70,6 +71,32 @@ def test_source_route_directory_cannot_be_overwritten(tmp_path):
replay.run(tmp_path, tmp_path/'selected_controller')
def test_current_replay_rejects_missing_model_clocks_before_comparing_outputs(tmp_path, monkeypatch):
source = tmp_path/'source'
source.mkdir()
np.savez(source/'route.npz', model_paths=np.zeros((1, 4, 33)))
monkeypatch.setattr(replay, 'verify_dependency', lambda: tmp_path)
with pytest.raises(ValueError, match='requires original model_position_t and model_orientation_t'):
replay.run(source, tmp_path/'output')
assert not (tmp_path/'output').exists()
def test_extract_models_preserves_both_actual_clocks_and_rejects_count_mismatch():
clocks = np.array([[0., .3, 1.7]])
raw = {'model_paths': np.zeros((1, 4, 3)), 'model_position_t': clocks, 'model_orientation_t': clocks+.01}
model, = replay.extract_models(raw)
np.testing.assert_array_equal(model.position.t, clocks[0])
np.testing.assert_array_equal(model.orientation.t, clocks[0]+.01)
raw['model_orientation_t'] = clocks[:, :2]
with pytest.raises(ValueError, match='must match'):
replay.extract_models(raw)
def test_historical_damping_replay_rejects_current_controller(tmp_path):
with pytest.raises(ValueError, match='historical damping replay requires unchanged C1'):
damping_replay.run(tmp_path/'source', tmp_path/'output', candidate_version='current')
@pytest.mark.parametrize('field,value', [(0, 5.12), (1, .501), (2, .00002), (3, .000001), (0, np.nan)])
def test_field_validation_catches_range_and_zero_c2_c3_violations(field, value):
command = np.zeros((2, 4))