mirror of
https://github.com/sunnypilot/sunnypilot.git
synced 2026-10-01 06:03:43 +08:00
Ford: reduce v24 Python overhead and trace health-loop stalls
This commit is contained in:
@@ -68,3 +68,19 @@ Route 174 ran v24 on `528ed3615` and recorded System Lagging during turns. Full
|
||||
The search now rejects candidates that fail its existing immediate-error constraint before evaluating their full return. Within each selection it also reuses return costs for identical predicted states and prefix costs. No cache survives the selection, and the candidate set, full-return policy, scoring, tie breaks and command cadence are unchanged.
|
||||
|
||||
Compared with the original native library built at the same optimization level, all nine outputs matched exactly in 600 deterministic stress cases and 720 comparisons using the 360 active logged states from route 174 at both firmware tick counts. On the development Mac, the recorded-state benchmark's 95th-percentile selection time fell from 0.317 ms to 0.090 ms; aggregate speedup was 1.47×. The broader stress benchmark improved 5.41× in aggregate. These are local encoder timings, not measured post-fix comma CPU utilization or a guarantee of system scheduling latency. Frozen pre-optimization commands and costs are also covered by the regression suite.
|
||||
|
||||
## Route 175 follow-up
|
||||
|
||||
Route `00000175--9529c33e36` ran `99b3fb03e` with v24 enabled. All ten full rlogs contain no System Lagging alert. Remaining communication warnings identify deviceState and, sometimes, managerState. DeviceState publication timestamps contain gaps up to 8.06 seconds; manager polls that service and falls back to its one-second timeout. The recorded CPU samples still show card/controlsd/selfdrived close to one core in sustained sections, so further optimization is useful independently of the health-process stalls.
|
||||
|
||||
Two wheel pauses near 100 and 115 degrees follow short steeringPressed detections. At about 175.35 and 176.94 seconds, the existing override handling transmits mode 0 / zero C0/C1 for approximately 87 and 117 ms, then rebuilds the request. Those pauses do not show limitReached. None of the 1,710 active v24 diagnostic samples reports the inverse acceleration allowance clipping the target. The user reported small left corrections against a rightward tendency, so these detections cannot be classified as false driver input. Override behavior is unchanged.
|
||||
|
||||
The bookmark marks a moderately good but slightly wide left turn. It contains brief limitReached states with continued mode-2 commands, rather than a frozen command. Wheel-tracking error remains; this route is not evidence of closed-loop equivalence to another steering platform. An 8.71-second carState recording gap around 290.55–299.26 seconds, with other card streams continuing and buffered diagnostics arriving later, is excluded from tracking statistics.
|
||||
|
||||
The additional encoder optimization reuses the already-computed C0/C1 calibration gains, constructs the same sorted candidate grid with fewer NumPy temporaries, and decodes fixed calibration entries once. Relative to `99b3fb03e`, all nine search outputs matched exactly in 600 deterministic stress cases and 3,420 comparisons using active route-175 states at both tick counts. Another 20,000 grid-boundary comparisons matched. A 6,000-cycle adapter/CAN/observer exercise, including stop, override, reengagement, steps and reversals, produced exactly the same packed commands and observer states.
|
||||
|
||||
On the development Mac, recorded-state selection p95 decreased from 0.174 to 0.102 ms (1.63× aggregate speedup). Full pipeline p95 decreased from 0.199 to 0.162 ms (1.24× aggregate speedup). These local measurements do not establish comma scheduling latency, and there are scheduler outliers in both runs. No target, gain, bound, override rule, or steering command was intentionally changed.
|
||||
|
||||
The hardwared patch avoids redundant synchronous Params removal for hidden alerts whose unused extra text changes, and routes the Tici-support alert through the same change-only helper. Initial cleanup, visible text updates, and clearing remain synchronous; failed writes are retried. An unchanged hidden-temperature alert reproduces the unnecessary blocking operation on the previous implementation. The logs do not identify which blocking call caused every on-device stall. Slow-stage timing now records operations exceeding 100 ms to distinguish thermal/system reads, Chestnut status, startup/engagement parameters, power reads, publication, and persistence. Watchdog thresholds and health checks remain unchanged.
|
||||
|
||||
Validation: 524 focused controller/hardware tests passed; Ruff and `git diff --check` passed. Diagnostic artifacts, route data, and benchmarks are local under `.cache/ford_v24_followup`; the visual report is `route-175.html` in the existing report directory. No truck validation of this follow-up optimization is claimed.
|
||||
|
||||
@@ -6,6 +6,7 @@ The next-update accuracy constraint is enabled for the trial.
|
||||
|
||||
import copy
|
||||
import ctypes
|
||||
import math
|
||||
import sys
|
||||
from pathlib import Path
|
||||
import numpy as np
|
||||
@@ -68,13 +69,14 @@ LIB.paired_select.restype = None
|
||||
|
||||
def levels(held, pref, rate, count, lsb, bound, clear):
|
||||
reach = rate * 0.008 * count + lsb / 2
|
||||
lo = max(round(-bound / lsb), int(np.floor((held - reach) / lsb)))
|
||||
hi = min(round(bound / lsb), int(np.ceil((held + reach) / lsb)))
|
||||
lo = max(round(-bound / lsb), math.floor((held - reach) / lsb))
|
||||
hi = min(round(bound / lsb), math.ceil((held + reach) / lsb))
|
||||
# Include both sides of the fast-latch clearing thresholds even when outside
|
||||
# the reachable range: equal slew endpoints can otherwise have different flags.
|
||||
near = np.floor(np.array([-clear, clear]) / lsb).astype(int)
|
||||
thresholds = np.r_[near - 1, near, near + 1] * lsb
|
||||
return np.unique(np.clip(np.r_[np.arange(lo, hi + 1) * lsb, pref, round(held / lsb) * lsb, thresholds, -bound, bound], -bound, bound))
|
||||
candidates = [i * lsb for i in range(lo, hi + 1)]
|
||||
candidates.extend((math.floor(sign * clear / lsb) + offset) * lsb for sign in (-1, 1) for offset in (-1, 0, 1))
|
||||
candidates.extend((pref, round(held / lsb) * lsb, -bound, bound))
|
||||
return np.array(sorted({min(max(x, -bound), bound) for x in candidates}), dtype=np.float64)
|
||||
|
||||
|
||||
class PairedRelease:
|
||||
@@ -88,11 +90,12 @@ class PairedRelease:
|
||||
m = self.request
|
||||
if not self.freeze_i or m.c0_i != 0:
|
||||
raise ValueError('Joint encoder requires its nominal zero-I estimate')
|
||||
if phase not in range(10) or speed_kmh <= 0 or not np.isfinite([speed_kmh, curvature, *state(m)]).all():
|
||||
s = state(m)
|
||||
if phase not in range(10) or speed_kmh <= 0 or not all(math.isfinite(x) for x in (speed_kmh, curvature, *s)):
|
||||
raise ValueError('Finite moving-vehicle inputs and scheduler phase required')
|
||||
p = parameters(m, speed_kmh, True, self.interaction)
|
||||
s = state(m)
|
||||
pref = np.array(quantize(static_pair(m, curvature, speed_kmh)))
|
||||
# parameters() already probed this same state/speed for the channel gains.
|
||||
pref = np.array(quantize(static_pair(m, curvature, speed_kmh, gains=(float(p[0]), float(p[1])))))
|
||||
target = float(p[0] * pref[0] + p[1] * pref[1])
|
||||
count = next(k for k in range(1, 3) if (phase + 8 * k) // 10 > 0)
|
||||
c0s = levels(m.c0, pref[0], max(p[10:12]), count, 0.01, 5.11, p[14])
|
||||
|
||||
@@ -65,12 +65,13 @@ def arc_pair(curvature, speed_kmh):
|
||||
return clip(c0, -C0_BOUND, C0_BOUND), clipped_c1
|
||||
|
||||
|
||||
def static_pair(request, curvature, speed_kmh):
|
||||
def static_pair(request, curvature, speed_kmh, *, gains=None):
|
||||
"""Preserve the base's channel proportion, solve its static curvature sum."""
|
||||
probe = copy.copy(request)
|
||||
gains = probe.step(speed_kmh, request.c0, request.c1, freeze_i=True)
|
||||
if gains is None:
|
||||
probe = copy.copy(request).step(speed_kmh, request.c0, request.c1, freeze_i=True)
|
||||
gains = probe['g0'], probe['g1']
|
||||
p0, p1 = arc_pair(curvature, speed_kmh)
|
||||
total = gains['g0'] * p0 + gains['g1'] * p1
|
||||
total = gains[0] * p0 + gains[1] * p1
|
||||
scale = curvature / total if total else 1.0
|
||||
return clip(p0 * scale, -C0_BOUND, C0_BOUND), clip(p1 * scale, -C1_BOUND, C1_BOUND)
|
||||
|
||||
|
||||
@@ -14,13 +14,14 @@ DT = 0.008
|
||||
|
||||
class Calibration:
|
||||
def __init__(self):
|
||||
self.entries = json.loads(Path(__file__).with_name('calibration.json').read_text())['entries']
|
||||
entries = json.loads(Path(__file__).with_name('calibration.json').read_text())['entries']
|
||||
self.entries = {int(address, 16): (item['format'], tuple(item['values'])) for address, item in entries.items()}
|
||||
|
||||
def read(self, address, fmt='f'):
|
||||
item = self.entries[hex(address)]
|
||||
if item['format'] != fmt:
|
||||
stored_format, values = self.entries[address]
|
||||
if stored_format != fmt:
|
||||
raise ValueError(f'Unexpected calibration format: {address:x}')
|
||||
return tuple(item['values'])
|
||||
return values
|
||||
|
||||
def f(self, address):
|
||||
return self.read(address)[0]
|
||||
|
||||
@@ -109,10 +109,21 @@ prev_offroad_states: dict[str, tuple[bool, str | None]] = {}
|
||||
|
||||
|
||||
def set_offroad_alert_if_changed(offroad_alert: str, show_alert: bool, extra_text: str | None=None):
|
||||
# Hidden alerts have no displayed text. Avoid taking the Params file lock
|
||||
# again whenever a temperature changes while its alert remains hidden.
|
||||
if not show_alert:
|
||||
extra_text = None
|
||||
if prev_offroad_states.get(offroad_alert, None) == (show_alert, extra_text):
|
||||
return
|
||||
prev_offroad_states[offroad_alert] = (show_alert, extra_text)
|
||||
set_offroad_alert(offroad_alert, show_alert, extra_text)
|
||||
prev_offroad_states[offroad_alert] = (show_alert, extra_text)
|
||||
|
||||
|
||||
def log_slow_hardware_stage(stage: str, start: float) -> float:
|
||||
now = time.monotonic()
|
||||
if now - start > 0.1:
|
||||
cloudlog.event("Hardware loop slow stage", stage=stage, elapsed=now - start)
|
||||
return now
|
||||
|
||||
def touch_thread(end_event):
|
||||
count = 0
|
||||
@@ -278,9 +289,11 @@ def hardware_thread(end_event, hw_queue) -> None:
|
||||
if (sm.frame % round(SERVICE_LIST['pandaStates'].frequency * DT_HW) != 0) and not ign_edge:
|
||||
continue
|
||||
|
||||
stage_start = time.monotonic()
|
||||
msg = messaging.new_message('deviceState', valid=True)
|
||||
msg.deviceState = thermal_config.get_msg()
|
||||
msg.deviceState.deviceType = HARDWARE.get_device_type()
|
||||
stage_start = log_slow_hardware_stage("thermal_read", stage_start)
|
||||
|
||||
try:
|
||||
last_hw_state = hw_queue.get_nowait()
|
||||
@@ -304,6 +317,7 @@ def hardware_thread(end_event, hw_queue) -> None:
|
||||
msg.deviceState.modemTempC = last_hw_state.modem_temps
|
||||
|
||||
msg.deviceState.screenBrightnessPercent = HARDWARE.get_screen_brightness()
|
||||
stage_start = log_slow_hardware_stage("system_stats", stage_start)
|
||||
|
||||
set_usb_state(msg.deviceState, last_hw_state.usb_state)
|
||||
chestnut.update(started_ts is None, last_hw_state.usb_state)
|
||||
@@ -312,6 +326,7 @@ def hardware_thread(end_event, hw_queue) -> None:
|
||||
chestnut_status.update(started_ts is None, branch, last_hw_state.usb_state, chestnut.failed,
|
||||
params.get_bool("ChestnutLoading"), params.get("ChestnutActive"),
|
||||
chestnut_state if chestnut_valid else None, set_offroad_alert_if_changed)
|
||||
stage_start = log_slow_hardware_stage("chestnut_status", stage_start)
|
||||
# this subset is only used for offroad
|
||||
temp_sources = [
|
||||
msg.deviceState.memoryTempC,
|
||||
@@ -372,13 +387,14 @@ def hardware_thread(end_event, hw_queue) -> None:
|
||||
is_unsupported_combo = COMMA_HARDWARE and HARDWARE.get_device_type() == "tici" and build_metadata.channel_type != "tici"
|
||||
startup_conditions["not_tici"] = not is_unsupported_combo
|
||||
onroad_conditions["not_tici"] = not is_unsupported_combo
|
||||
set_offroad_alert("Offroad_TiciSupport", is_unsupported_combo, extra_text=build_metadata.channel)
|
||||
set_offroad_alert_if_changed("Offroad_TiciSupport", is_unsupported_combo, extra_text=build_metadata.channel)
|
||||
|
||||
# if the temperature enters the danger zone, go offroad to cool down
|
||||
onroad_conditions["device_temp_good"] = thermal_status < ThermalStatus.critical
|
||||
extra_text = f"{offroad_comp_temp:.1f}C"
|
||||
show_alert = (not onroad_conditions["device_temp_good"] or not startup_conditions["device_temp_engageable"]) and onroad_conditions["ignition"]
|
||||
set_offroad_alert_if_changed("Offroad_TemperatureTooHigh", show_alert, extra_text=extra_text)
|
||||
stage_start = log_slow_hardware_stage("startup_conditions", stage_start)
|
||||
|
||||
if show_alert:
|
||||
msg.deviceState.fanSpeedPercentDesired = 100
|
||||
@@ -404,6 +420,7 @@ def hardware_thread(end_event, hw_queue) -> None:
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
stage_start = log_slow_hardware_stage("engagement_params", stage_start)
|
||||
should_pwrsave = not onroad_conditions["ignition"] and msg.deviceState.screenBrightnessPercent < 1e-3
|
||||
if should_pwrsave != pwrsave or (count == 0):
|
||||
HARDWARE.set_power_save(should_pwrsave)
|
||||
@@ -439,13 +456,16 @@ def hardware_thread(end_event, hw_queue) -> None:
|
||||
power_monitor.calculate(voltage, onroad_conditions["ignition"])
|
||||
msg.deviceState.offroadPowerUsageUwh = power_monitor.get_power_used()
|
||||
msg.deviceState.carBatteryCapacityUwh = max(0, power_monitor.get_car_battery_capacity())
|
||||
stage_start = log_slow_hardware_stage("power_monitor", stage_start)
|
||||
current_power_draw = HARDWARE.get_current_power_draw()
|
||||
statlog.sample("power_draw", current_power_draw)
|
||||
msg.deviceState.powerDrawW = current_power_draw
|
||||
stage_start = log_slow_hardware_stage("power_draw_read", stage_start)
|
||||
|
||||
som_power_draw = HARDWARE.get_som_power_draw()
|
||||
statlog.sample("som_power_draw", som_power_draw)
|
||||
msg.deviceState.somPowerDrawW = som_power_draw
|
||||
stage_start = log_slow_hardware_stage("som_power_read", stage_start)
|
||||
|
||||
# Check if we need to shut down
|
||||
if power_monitor.should_shutdown(onroad_conditions["ignition"], in_car, off_ts, started_seen):
|
||||
@@ -461,6 +481,7 @@ def hardware_thread(end_event, hw_queue) -> None:
|
||||
|
||||
msg.deviceState.thermalStatus = thermal_status
|
||||
pm.send("deviceState", msg)
|
||||
stage_start = log_slow_hardware_stage("device_state_publish", stage_start)
|
||||
|
||||
statlog.gauge("free_space_percent", msg.deviceState.freeSpacePercent)
|
||||
statlog.gauge("gpu_usage_percent", msg.deviceState.gpuUsagePercent)
|
||||
@@ -514,6 +535,7 @@ def hardware_thread(end_event, hw_queue) -> None:
|
||||
|
||||
count += 1
|
||||
should_start_prev = should_start
|
||||
log_slow_hardware_stage("stats_and_persistence", stage_start)
|
||||
|
||||
|
||||
def main():
|
||||
|
||||
@@ -0,0 +1,55 @@
|
||||
from unittest.mock import Mock
|
||||
|
||||
import pytest
|
||||
|
||||
from openpilot.system.hardware import hardwared
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def clear_alert_cache(monkeypatch):
|
||||
monkeypatch.setattr(hardwared, 'prev_offroad_states', {})
|
||||
|
||||
|
||||
def test_hidden_alert_does_not_repeat_blocking_param_operation(monkeypatch):
|
||||
persist = Mock()
|
||||
monkeypatch.setattr(hardwared, 'set_offroad_alert', persist)
|
||||
# The first call must remove an alert left by a previous process.
|
||||
hardwared.set_offroad_alert_if_changed('Offroad_TemperatureTooHigh', False, '42.0C')
|
||||
persist.assert_called_once_with('Offroad_TemperatureTooHigh', False, None)
|
||||
# An unrelated writer may now hold the shared Params lock. Hidden temperature
|
||||
# changes must not touch it again, nor delay the deviceState heartbeat.
|
||||
persist.side_effect = AssertionError('Unexpected blocking Params operation')
|
||||
for temp in ('42.1C', '42.2C', '45.0C'):
|
||||
hardwared.set_offroad_alert_if_changed('Offroad_TemperatureTooHigh', False, temp)
|
||||
|
||||
|
||||
def test_visible_alert_text_and_clear_still_persist(monkeypatch):
|
||||
persist = Mock()
|
||||
monkeypatch.setattr(hardwared, 'set_offroad_alert', persist)
|
||||
for show, text in [(False, '42C'), (True, '107C'), (True, '107C'), (True, '108C'), (False, '90C')]:
|
||||
hardwared.set_offroad_alert_if_changed('Offroad_TemperatureTooHigh', show, text)
|
||||
assert [call.args for call in persist.call_args_list] == [
|
||||
('Offroad_TemperatureTooHigh', False, None),
|
||||
('Offroad_TemperatureTooHigh', True, '107C'),
|
||||
('Offroad_TemperatureTooHigh', True, '108C'),
|
||||
('Offroad_TemperatureTooHigh', False, None),
|
||||
]
|
||||
|
||||
|
||||
def test_failed_alert_write_is_retried(monkeypatch):
|
||||
persist = Mock(side_effect=[OSError('write failed'), None])
|
||||
monkeypatch.setattr(hardwared, 'set_offroad_alert', persist)
|
||||
with pytest.raises(OSError):
|
||||
hardwared.set_offroad_alert_if_changed('Offroad_TiciSupport', False)
|
||||
hardwared.set_offroad_alert_if_changed('Offroad_TiciSupport', False)
|
||||
assert persist.call_count == 2
|
||||
|
||||
|
||||
def test_stage_timing_reports_slow_stage_only(monkeypatch):
|
||||
report = Mock()
|
||||
monkeypatch.setattr(hardwared.cloudlog, 'event', report)
|
||||
monkeypatch.setattr(hardwared.time, 'monotonic', Mock(side_effect=[1.05, 6.05, 6.07]))
|
||||
start = hardwared.log_slow_hardware_stage('thermal_read', 1.)
|
||||
start = hardwared.log_slow_hardware_stage('system_stats', start)
|
||||
hardwared.log_slow_hardware_stage('chestnut_status', start)
|
||||
report.assert_called_once_with('Hardware loop slow stage', stage='system_stats', elapsed=5.)
|
||||
Reference in New Issue
Block a user