Ford: reduce v24 search work during turns

Reject candidates failing the existing immediate-error bound before simulating their full return. Reuse identical state and prefix-cost evaluations within each selection, preserving candidate order, scoring, and tie breaks.

Route 174 shows card, controlsd, and selfdrived sharing a saturated core during System Lagging alerts. This reduces encoder work without changing steering targets, commands, rates, or alert thresholds.

Validation: all nine native outputs exactly match the original library in 600 seeded cases and 720 recorded-state comparisons. Recorded-state p95 chooser runtime improves from 0.317 ms to 0.090 ms on the development Mac. 492 regression tests, Ruff, strict C++ compilation, and diff checks pass. Post-fix device timing remains to be measured.
This commit is contained in:
Isaac Barham
2026-09-18 00:21:32 -04:00
parent 528ed36155
commit 99b3fb03e9
3 changed files with 68 additions and 8 deletions
+8
View File
@@ -60,3 +60,11 @@ python -m pytest openpilot/selfdrive/car/tests/test_ford_joint_control.py \
Build the native kernel through SCons. On the development checkout, its real SConscript and native parameter library were compiled successfully. A complete root build could not run because the checkout lacks the msgq/rednose SCons tool submodules. No on-device build or road validation is claimed.
The trial's question is whether coordinating both channels preserves entry while reducing unnecessary correction during release. Offline results justify the opt-in comparison; they cannot establish smoothness or closed-loop stability.
## Runtime optimization after route 174
Route 174 ran v24 on `528ed3615` and recorded System Lagging during turns. Full rlogs show card using about 63% of one CPU core during the first alert; card, controlsd and selfdrived share that core and priority, and their combined measured load was about 101%. This supports CPU contention, not a PSCM limit, as the explanation for this alert. Comma timing must still be checked after the optimization.
The search now rejects candidates that fail its existing immediate-error constraint before evaluating their full return. Within each selection it also reuses return costs for identical predicted states and prefix costs. No cache survives the selection, and the candidate set, full-return policy, scoring, tie breaks and command cadence are unchanged.
Compared with the original native library built at the same optimization level, all nine outputs matched exactly in 600 deterministic stress cases and 720 comparisons using the 360 active logged states from route 174 at both firmware tick counts. On the development Mac, the recorded-state benchmark's 95th-percentile selection time fell from 0.317 ms to 0.090 ms; aggregate speedup was 1.47×. The broader stress benchmark improved 5.41× in aggregate. These are local encoder timings, not measured post-fix comma CPU utilization or a guarantee of system scheduling latency. Frozen pre-optimization commands and costs are also covered by the regression suite.
@@ -200,6 +200,24 @@ def test_candidate_does_not_mutate_state_and_native_step_matches_python():
assert abs(info['first_state'][3]-info['planned_target']) <= info['immediate_error_bound']+1e-12
@pytest.mark.parametrize('speed,c0,c1,filtered,fast,target,phase,pair,cost', [
(20., 0., 0., 0., False, .03, 0, (.03, .002), 5.6554469893989605),
(20., 3., .4, .03, False, -.03, 2, (2.98, .399), 151.49476621872228),
(40., 3., .4, .03, True, 0., 0, (2.95, .395), 12.167867914611671),
(6., 5.11, .5, .1, False, -.1, 0, (5.08, .498), 1391.5462859829215),
(80., -5.11, -.5, -.1, True, .03, 2, (-5.08, -.4975), 34.51646821335084),
(40., 3., -.4, 0., False, 0., 2, (3.02, -.399), 6.184073685045232),
])
def test_optimized_selection_preserves_v24_commands(speed, c0, c1, filtered, fast, target, phase, pair, cost):
# Frozen outputs from 528ed3615 before pruning/caching the full-return search.
# Cover entry, reversal, release, saturation, cancellation and both tick counts.
m = MainRequest(native_lookup=True)
m.c0, m.c1, m.filtered, m.fast = c0, c1, filtered, fast
command, info = PairedRelease(m).choose(speed, target, phase)
assert command == pytest.approx(pair, abs=1e-15, rel=0)
assert info['cost'] == pytest.approx(cost, abs=1e-12, rel=1e-12)
@pytest.mark.parametrize('active,override', list(itertools.product((False, True), repeat=2)))
def test_card_transmit_hook_and_fault_alert_use_actual_source(active, override):
import ast
@@ -1,7 +1,9 @@
// Opt-in Ford joint encoder. Numerical request state is estimated, not ECU RAM.
#include <algorithm>
#include <array>
#include <cmath>
#include <cstring>
#include <map>
static double clip(double x, double lo, double hi) {
return std::min(std::max(x, lo), hi);
@@ -31,19 +33,29 @@ static void step(double *s, const double *p, double c0, double c1) {
s[3] += alpha * (raw - s[3]);
}
extern "C" double paired_cost(const double *initial, const double *p, double target,
const double *pref, int count, double c0, double c1, double *first) {
double s[5];
std::memcpy(s, initial, sizeof(s));
static double command_cost(const double *initial, const double *p, double target,
const double *pref, int count, double c0, double c1, double *s) {
std::memcpy(s, initial, 5 * sizeof(double));
double steady_raw = p[0] * pref[0] + p[1] * pref[1];
double steady_error = (steady_raw - target) / .01, cost = 0;
for (int j = 0; j < count; j++) {
step(s, p, c0, c1);
double error = (s[3] - target) / .01;
cost += .008 * (error * error - steady_error * steady_error);
}
return cost;
}
static double return_cost(const double *first, const double *p, double target, const double *pref, double cost) {
double s[5];
std::memcpy(s, first, sizeof(s));
double steady_raw = p[0] * pref[0] + p[1] * pref[1];
double steady_error = (steady_raw - target) / .01;
auto tick = [&](double a, double b) {
step(s, p, a, b);
double error = (s[3] - target) / .01;
cost += .008 * (error * error - steady_error * steady_error);
};
for (int j = 0; j < count; j++) tick(c0, c1);
if (first) std::memcpy(first, s, sizeof(s));
// Include both channels' entire return, followed by the exact filter tail.
// There is no adjustable planning horizon or retained future command plan.
int n = 2 + std::ceil(std::max(std::abs(s[0] - pref[0]) / std::min(p[10], p[11]),
@@ -55,17 +67,37 @@ extern "C" double paired_cost(const double *initial, const double *p, double tar
return cost + .008 * (2 * steady_error * d * a / alpha + d * d * a * a / (1 - a * a));
}
extern "C" double paired_cost(const double *initial, const double *p, double target,
const double *pref, int count, double c0, double c1, double *first) {
double s[5];
double cost = command_cost(initial, p, target, pref, count, c0, c1, s);
if (first) std::memcpy(first, s, sizeof(s));
return return_cost(s, p, target, pref, cost);
}
extern "C" void paired_select(const double *initial, const double *p, double target,
const double *pref, int count, const double *c0s, int n0,
const double *c1s, int n1, int preserve_now, double *result) {
double best = 1e300, best_move = 1e300, best_remaining = 1e300;
double first[5];
// Slew clipping makes different command fields reach identical states. Reuse
// their exact return cost within this selection only. Include the prefix cost
// so the original floating-point accumulation and tie breaks are preserved.
std::map<std::array<double, 6>, double> costs;
auto finish = [&](double cost) {
if (!std::isfinite(cost)) return return_cost(first, p, target, pref, cost);
std::array<double, 6> key = {first[0], first[1], first[2], first[3], first[4], cost};
auto entry = costs.emplace(key, 0.0);
if (entry.second) entry.first->second = return_cost(first, p, target, pref, cost);
return entry.first->second;
};
double max_error = 1e300, anchor_score = 1e300, anchor_move = 1e300, anchor_remaining = 1e300;
if (preserve_now) {
// Preserve the C1-anchored policy's immediate target accuracy.
// An inequality against a feasible reference, not a new gain/deadband.
for (int i = 0; i < n0; i++) {
double cost = paired_cost(initial, p, target, pref, count, c0s[i], pref[1], first);
double cost = command_cost(initial, p, target, pref, count, c0s[i], pref[1], first);
cost = finish(cost);
double score = std::nearbyint(cost * 1e12) / 1e12;
double move = std::abs(c0s[i] - initial[0]), remaining = std::abs(c0s[i] - pref[0]);
if (score < anchor_score || (score == anchor_score && (move < anchor_move || (move == anchor_move && remaining < anchor_remaining)))) {
@@ -78,8 +110,10 @@ extern "C" void paired_select(const double *initial, const double *p, double tar
}
for (int i = 0; i < n0; i++) {
for (int j = 0; j < n1; j++) {
double cost = paired_cost(initial, p, target, pref, count, c0s[i], c1s[j], first);
double cost = command_cost(initial, p, target, pref, count, c0s[i], c1s[j], first);
if (std::abs(first[3] - target) > max_error + 1e-12) continue;
// Reject infeasible candidates before simulating their full return.
cost = finish(cost);
// Numerical equality only. Tie breaks cannot trade worse tracking for
// less channel motion; normalize them using the existing field spans.
double score = std::nearbyint(cost * 1e12) / 1e12;