feat(vapi): recover CAN layouts from the MCU decoder, and default every tool to the generated DBC #16

Merged
outlandnish merged 54 commits from feat/vapi-layout into main 2026-09-09 19:19:35 -05:00
Owner

Recovers CAN signal layout from the MCU's own decoder and makes the generated DBC
the default signal database everywhere.

vapi layout recovery — vapi_layout.py walks libQtCarVAPI's dispatch tree to
recover bit layout for ~98% of 26227 signals; vapi_emu.py runs the firmware's
decoder under Unicorn rather than modelling it. Validated against compact.json at
100% bit agreement on 2020, with mux pages cross-checked against alert names on
every extraction.

The catalog is the database — layouts are an optional overlay, and every tool
now defaults to the generated DBC instead of the partial compact.json (which
carries only a fraction of the catalog). alertlog decodes its fields from the .so
by default, with packer JSON as fallback.

ODIN runner — flashes through dfu for real instead of faking success, picks the
platform tree and lists a car's procedures, confirms a flash with pickers before it
starts, and lets the operator decline bootloader updates or declare on the bench
what the bus cannot report.

Also: CP PLC subcomponents route to their own flash regions, the signed
bootloader's handoff window is held open with TesterPresent, and a batch of
config/test/UI fixes.

53 commits, 52 files, +12207/-309.

Note: the branch is 9 commits behind main at time of opening.

🤖 Generated with Claude Code

Recovers CAN signal layout from the MCU's own decoder and makes the generated DBC the default signal database everywhere. **vapi layout recovery** — `vapi_layout.py` walks libQtCarVAPI's dispatch tree to recover bit layout for ~98% of 26227 signals; `vapi_emu.py` runs the firmware's decoder under Unicorn rather than modelling it. Validated against compact.json at 100% bit agreement on 2020, with mux pages cross-checked against alert names on every extraction. **The catalog is the database** — layouts are an optional overlay, and every tool now defaults to the generated DBC instead of the partial compact.json (which carries only a fraction of the catalog). alertlog decodes its fields from the .so by default, with packer JSON as fallback. **ODIN runner** — flashes through dfu for real instead of faking success, picks the platform tree and lists a car's procedures, confirms a flash with pickers before it starts, and lets the operator decline bootloader updates or declare on the bench what the bus cannot report. Also: CP PLC subcomponents route to their own flash regions, the signed bootloader's handoff window is held open with TesterPresent, and a batch of config/test/UI fixes. 53 commits, 52 files, +12207/-309. Note: the branch is 9 commits behind `main` at time of opening. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
libQtCarCANData.so carries the complete signal catalog but no bit-layout, and
Model3_ETH.compact.json -- which does -- is stripped down every release (2022:
347 signals vs the catalog's 26227). candata_to_dbc was therefore borrowing
layout from OLDER revisions, which is unsound: layouts move between releases.
0x118 DI_immobilizerState is 27|3 in 2020 but 13|3 in 2022, DI_driveBlocked
12|2 -> 24|2, and bits 27-28 became DI_hvilSystemStatus -- so the 2020 table
read hvilSystemStatus AS immobilizerState, silently, since both are 0 on a
dormant bench unit.

The layout does exist same-revision, as code, in libQtCarVAPI.so:

    CAN frame -> GUICanCracker::crackMessage(bus, msgId, payload)
              -> CANDataManager::storeSignalValue(key, double, valid, bus)

crackMessage is one generated function (2022: ~1.1 MB) whose cases load payload
bytes, shift/mask the field, scale it, and call storeSignalValue with a 32-bit
KEY -- and that KEY is the u32 at +0x08 in the catalog's signal struct, which
recovers the name. The call count equals the catalog signal count exactly, so
every catalogued signal has exactly one decode site.

vapi_layout.py disassembles it with capstone and models the field arithmetic,
emitting nothing rather than a guess wherever the model does not fully follow
the sequence. Multiplexed messages are most of the catalog, and GCC compiles
their page switch three ways -- jump table, if-else compare chain, and a chain
that walks the page number down with `sub` -- all three are recognised, as is
`ja default` being a binary split rather than the end of the chain.

Measured against the same-revision 2020 DBC (real ground truth): 98.4% of
start|width exact, 99.99% of signedness, 99.5% of scale+offset, 22 genuine
Intel disagreements in 8767 signals. Coverage 98.0% of the 2022 catalog, 86.0%
of 2020; the generated 2022 DBC goes from 21368 to 25768 signals.

candata_to_dbc auto-detects the sibling VAPI library and prepends it as the
highest-priority donor (--no-vapi opts out), and di.py's 0x118 overlay becomes
revision-aware, selecting floor-with-clamp on the DIR's firmware.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PROC_DI_X_RESOLVER-LEARN failed with "Not in Dyno Mode" even though the status
dump showed DI_tractionControlMode = TC_DYNO_MODE. The gate does not read that
signal: it reads the CID data-value GUI_tractionControlModeRequest, which is the
TOUCHSCREEN's request published by the MCU, not the inverter's CAN echo. On a
bench there is no MCU, so it read None and the graph took the failure branch.
The same trap applies to VAPI_shiftState vs DI_gear.

BenchBackend now derives those CID names from the bus when no MCU is present,
via CidStore's derive hook, so the graph's gates evaluate as they would in a
car. An unopenable bus reads as absent rather than raising.

Separately, ODIN could only see signals can_live happened to be decoding for its
own HUD. can_live now retains the last raw frame per (bus, CAN id) alongside the
arrival stamp, and the bench backend takes that as a frame_source plus the
signal database -- so ANY message seen on the bus is readable by ODIN, decoded
on demand through the catalog's signal->message index rather than a fixed list.

can_live also threads the per-revision 0x118 overlay table (see the VAPI layout
commit) into _overlay_0x118, resolved once per flush.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A mux page body can open on a payload byte loaded BEFORE the dispatch --
IBST_a181_DIchassisControlDLC is just `shr al,4; and eax,1`, where `al` holds
the selector byte the compiler kept live across the branch. extract_stores
sweeps crackMessage linearly, and a page body sits far from its own case, so it
arrived there with no register state and dropped the signal.

find_compare_chain now tracks which payload loads are still pristine and
snapshots them PER PAGE, and extract_stores installs that state on reaching the
page's entry address. Only untouched loads carry over: a register the chain did
arithmetic on is not what the page body expects, and a page that loads its own
byte overwrites the seed harmlessly -- so this can restore signals, not invent
them.

Per page is the whole point. Snapshotting once at the case's first branch and
reusing it for every page regressed 2020 from 86.0% to 82.4% (overlap drops
14 -> 84), because the chain mutates registers between one page's branch and the
next. Table dispatches get no seed at all: nothing walks them instruction by
instruction, so the live state at each branch is unknown.

Coverage 2022 98.0% -> 98.2% (unmodelled sites 195 -> 140), 2020 86.0% -> 86.2%
(272 -> 249), with ground-truth accuracy unchanged at 22 wrong in 8767.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every one of the 117 Motorola signals the 2020 extraction got wrong was a single
unmodelled instruction:

    mov   eax, dword ptr [rbx + 4]
    bswap eax                        ; ESP_infoApplicationCRC -> 39|32 @0

All 117 already had the right WIDTH and were emitted LITTLE with an LSB start --
the byte reversal states the byte order outright, rather than it having to be
inferred from which byte supplies the high half of a split field. The DBC start
of a Motorola field is its MSB, which after the swap is bit 7 of the first
payload byte. A swap of anything already shifted or masked is refused, since
then the field is not the whole load.

Motorola disagreements 117 -> 46, and all 46 that remain cover exactly the same
BITS as compact.json -- a field inside a single byte has an equivalent Motorola
spelling (start + width - 1, @0), so they decode identically and there is
nothing in the binary that says which notation to prefer.

Validation no longer goes through a generated DBC. It now compares against the
raw same-revision Model3_ETH.compact.json: Tesla's own layout DATA, shipped in
the same image as the CODE being decoded, so the two are independent. Of the
8781 signals in both: 8706 exact (99.15%), 8759 same bits (99.75%), 22
genuinely different. The docstring now also records what this cannot check --
compact.json covers 8781 of 10193 catalogued signals, so the ~1400 only this
recovers are exactly the ones no DBC could verify.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two sign-extension idioms were misread, and between them they accounted for
every remaining disagreement with Tesla's own layout data.

`shl eax,K` then `sar ax,M` is NOT a right shift of M. The shl first pushes the
load's top bits out of the 16-bit sub-register, so it is a net shift of (M-K)
and the width is however much of the load still sits above bit M. Reading it as
M put the start (M-K) bits too high: BMS_cacMinUpdateAhError 39|9 instead of
37|9, DAS_TE_vlC1 23|9 instead of 18|9 -- 19 signals. The full-register form,
where the shifts genuinely cancel, is a different rule and is now checked
second: it is the less specific case, and letting it win first shrank
BMS_loadReg_output's high half from 1 bit to 0.

`cwde` was treated as a no-op. It is a statement that the assembled field is 16
bits and signed, which is what caps VCFRONT_PCSCurrent at 16 rather than the 17
its `or` join implies. `cdqe` likewise at 32.

Against the raw same-revision compact.json, all 8820 signals present in both now
decode to exactly the same BITS -- 0 genuinely different, down from 22. 8767 are
also an exact start|width|endianness match; the other 53 differ only in
notation, since a field inside one byte has an equivalent Motorola spelling and
nothing in the binary says which the DBC author would pick.

Coverage 2022 98.2% -> 98.4%, 2020 86.2% -> 86.5%, and overlap drops are now
zero in both revisions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The `or` handler required the low side to sit at bit 0, which only expresses a
TWO-way join. PMR_bootGitHash takes bytes 1, 2-3 and 4-7, and at every
intermediate step of a chain like that BOTH sides carry a shift:

    mov   eax, dword ptr [rbx+4] ; shl rax,0x18
    or    rdx, rax               ; <- rdx already holds bytes 2-3 at bit 8
    movzx eax, byte ptr [rbx+1]
    or    rax, rdx

so the whole field was refused. Each side occupies result bits
[shl, shl+width); they join iff those runs are adjacent, and the result keeps
the lower side's payload origin and offset. The two-way case is the same rule
with lo.shl == 0, so it is unchanged.

Unmodelled decode sites 2022 140 -> 66, 2020 249 -> 165. Coverage 2022 98.4% ->
98.5%, 2020 86.5% -> 87.3%. Agreement with compact.json holds at 100% of the
bits over 80 more comparable signals (8900).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Model3_ETH.compact.json is only the subset Tesla ships to the diagnostic tool,
and it shrinks every release: 2022.45.15 carries 140 messages / 347 signals
against the catalog's 446 / 26227. Every tool defaulting to it could therefore
see a fraction of the bus, which on a bench looks like a bus fault rather than a
missing database. The DBC candata_to_dbc now builds covers the whole catalog
with layout recovered from that same firmware's own decoder.

config.ETH_DBC resolves TM3_ETH_DBC, else <PRODUCT>_ETH.<rev>.dbc in the project
root -- the path candata_to_dbc already writes. can_decoder.default_db() prefers
it and warns ONCE when falling back, rather than quietly degrading; vehicle_sim,
can_live, ecu_bench and the ODIN bench backend all route through it.

node_config gains _message_ids(), which reads either source. It only ever wanted
the name -> CAN id map for a node's UDS request/response pair, so the DBC is
strictly better there: a node whose UDS messages were stripped from compact.json
resolves off the DBC but raises KeyError off the JSON.

Nothing is lost by the switch. CanDatabase.from_dbc already reproduces the same
dict shape including receivers/min/max, and nothing in the tree consumes the
compact-only fields.

Loading the 2022 DBC gives 526 messages / 26010 signals, and 0x118 now carries
21 signals natively -- including DI_accelPedalPos and DI_driveBlocked, the ones
di.py has to reconstruct with a hand-maintained bit overlay when running off
compact.json.

Also adds progress output to candata_to_dbc: recovering layout disassembles
~1.1 MB of generated code several times, which was half a minute of silence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two defects that only showed up on 2026.8.3, whose catalog is 40455 signals
against 2022.45.15's 26227.

apply_guards checked "a message using mux_id must name its multiplexor" BEFORE
pruning overlaps, then pruned. In SDCR_info the signal it pruned WAS the
multiplexor (SDCR_infoIndex), so three m<id> pages were emitted with nothing to
select on. cantools reads such a message flat, and the pages collide -- it then
refuses the ENTIRE file, so one bad message cost all 673. Re-check after
pruning and hand the message to a donor. 2022 regenerates byte-identical.

config.ETH_DBC only resolved off TM3_FW, which is the revision vehicle_sim
TRANSMITS and is normally unset -- so generating a DBC did not make anything
use it, and every tool quietly fell back to compact.json. Resolve from TM3_ROOT
first (the root that already supplies nodes.json, the ODJ tree and
compact.json), keeping TM3_FW as a fallback and TM3_ETH_DBC as the override.
config.rev_from_root is now the one place the extraction suffix is stripped;
candata_to_dbc calls it, so the name it writes and the name config looks for
cannot drift.

2026.8.3: 673 messages / 39284 signals, 0 stranded pages, loads clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The overlap guard had one response to a message it could not make consistent:
discard it. On 2026.8.3 that cost 1243 signals across 22 messages -- including
all 458 of VCFRONT2_alertLog, of which 444 had been recovered perfectly. The
condemning signals were 14 stragglers whose mux page we failed to attribute.

An unpaged signal in a multiplexed message is read on EVERY page, so one whose
page we missed collides with all of them and trips the pervasive-overlap rule
on its own. But a genuinely mux-independent signal -- a counter, a checksum --
sits in bits no page uses and conflicts with nothing. Conflict is what tells
those two apart, not the absence of a page, so prune on that.

Guards now prune in order of how well corroborated a signal is:

  unpaged-and-conflicting  <  ordinary signal  <  the multiplexor

The multiplexor comes last because it is the best-evidenced signal in the
message: the dispatch we read the pages from is a branch on exactly its bits.
It also cannot be replaced -- without it the pages cannot be expressed at all
and the message has to go, which is what happened to SDCR_info when
SDCR_infoAppCrc landed on the selector's byte and both were dropped.

Only when conflicts survive all of that is the message handed to a donor.

  2026.8.3   gaps 1346 -> 647    VCFRONT2_alertLog 0 -> 444 signals
                                 VCFRONT1_alertLog 0 -> 162
                                 VCBATT2_alertLog  0 ->  85
                                 SDCR_info         0 ->   8
  2022.45.15 gaps  366 -> 334    UI_alertLog recovers 32 signals

Both revisions load clean; no message is left with stranded pages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A message whose pages hang on ONE payload bit gets neither of the shapes we
looked for. There is nothing to build a jump table from and nothing to compare,
so GCC emits a bare branch:

    test byte ptr [rbx + 6], 1      ; VC_pcsInterfaceMuxIndex, bit 48
    je   page0                      ; ... and page 1 is the fall-through

With no dispatch found, every signal landed on one flat page, the pages
collided, and the whole message was discarded. VCFRONT_vehicleStatus loads the
byte first and tests the register (`test al,1`), so both operand forms count;
the existing `test` rule only matched `test dl,dl` on an already-identified
selector, and worse, dropped the selector register on anything else.

Only a single-bit mask qualifies. `test` separates zero from nonzero, which is
an unambiguous two-way split exactly when the field is one bit wide.

Two details the real code forces:

* The bit test REDEFINES the selector, so a preceding `and` must not bound it.
  0x441 masks byte 6 with 0xf to read an unrelated signal immediately before
  testing bit 0 of that same byte; inheriting that 0..15 bound leaves 15 pages
  unaccounted for and the fall-through page is never attributed.
* The walk continues past the branch into the page body, where further `and`s
  are ordinary field extraction. The selector is pinned when it branches so
  they cannot be folded into the reported mask.

name_multiplexors now locates the selector signal from any contiguous mask
rather than assuming it is low-aligned in its byte.

  TAS_axleData    0 -> 13 signals, and reads as it should: page 0 front
                  (rawHeightFL/FR), page 1 rear (RL/RR), on TAS_axleIndex
  VC_pcsInterface 0 -> 13, VCBATT_pcsInterface 0 -> 13
  VCFRONT_vehicleStatus 0 -> 34, VCLEFT_switchStatus 9 -> 75,
  VCLEFT_windowStatus 5 -> 31, UI_systemMonitor 0 -> 12

  2026.8.3   gaps  647 -> 420
  2022.45.15 gaps  334 -> 149

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An alertLog's mux id IS its alert number, and the signal names carry it
(VCBATT2_a192_hvState is alert 192), so the catalog is its own oracle. It said
13247 of 18592 alertLog signals had the wrong page -- every message off by its
own constant. Not a missing page: a page numbered wrong, which silently decodes
one alert's payload as another alert's. Confirmed against the mux_id compact.json
does ship (21 signals across five revisions, all agreeing with the name).

The constant is the table's low bound, and GCC subtracts it three ways. Only
one was modelled:

    add ax, 0x3f1 / and ax, 0x3ff    TRCM_alertLog   -15 mod 1024
    add eax, 0x6d / cmp al, 9        EPAS3S_alertLog -147 mod 256, the
                                     truncation to 8 bits IS the modulus
    sub eax, 0x79                    OCS1P_alertLog  base alert 121

The modular forms were not read at all, and the plain `sub` was guarded by
str.isdigit(), which capstone's hex operands never satisfy -- so a bound of
0x79 parsed as 0 while a bound of 9 parsed fine.

Arithmetic on rbx/rsp/rbp is excluded: `sub rsp,0x28` in a case body is a frame
adjustment, not the selector being rebased. (_canon returns family names -- "b",
"sp" -- not register names, which is what makes that guard actually fire.)

The top-level message switch is untouched: its own bound stays 17 and all 446 /
580 cases still resolve to real catalog message ids in both revisions.

  2026.8.3    13247 wrong -> 0    (18351 correct, 241 unpaged)
  2022.45.15                 0    (11946 correct,  22 unpaged)

Every remaining unpaged alert lies OUTSIDE the id range of the pages we found,
with none inside: those are the sibling subtrees of a binary search tree whose
other branches we do not yet follow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A big switch is not one table, it is a TREE of range splits with tables at the
leaves. VCBATT2_alertLog tests alert 227 on its own, hands everything above it
to one subtree and everything below 0x3b to another, and tables only the run in
between:

    cmp ax, 0xe3 / je 0x6f16ad      ; alert 227
                   ja 0x6c5e4e      ; 228..    -> subtree
    cmp ax, 0x3a / jbe 0x70c277     ; ..58     -> subtree
    add ax, 0x3c5 ... jmp rax       ; 59..192  -> table

Reading only the root left 128 of that message's 209 signals with no page at
all -- and an unpaged signal does not simply go missing, it inherits whichever
mark precedes it in the address space, so VCBATT2 alerts were landing in
regions owned by LS_lightShow and BMS_kwhCounter.

The walker gains what the tree needs: `jbe`/`jb` split downward as `ja`/`jae`
already split upward, each run now carries its whole [lo, hi] range so a
sibling inherits the right bound, selectors may be read at any WIDTH (an alert
id is a 10-bit field read as a word), and a table met mid-walk is expanded in
place rather than ending the walk.

Three things this forced, each found by regression rather than reasoning:

* Comparisons are reduced modulo the selector mask only PAST a rebase, i.e.
  when the bias is negative. Reducing unconditionally folds ordinary
  comparisons in page bodies onto real pages -- with a one-bit selector
  `cmp al,5` becomes page 1 -- which cost TAS_axleData and both PARK_psc*Slot.
* A second register holding the same payload bytes is added to the selector
  set, not swapped in. PARK_pscEnvSlot reads byte 0 into eax and then word 0
  into ecx before comparing `al`; replacing the set loses the register the
  chain actually branches on.
* The selector is pinned when the first page resolves, and pages above its mask
  are refused. The running selector drifts as the walk continues into the page
  bodies, and a bogus page is worse than a missing one because it claims a code
  region belonging to something else.

Also: a rebase's modulus is the mask wherever the mask is visible, even when
the selector load is not -- a subtree is entered below it, and falling back on
the register width read `add ax,0x2d4` as base 64812 rather than 300.

  2026.8.3   40035 -> 40287 of 40455 (99.6%)   gaps 420 -> 168
             alertLog pages 18351 -> 18568 correct, 241 -> 24 unpaged
  2022.45.15 26078 -> 26127 of 26227 (99.6%)   gaps 149 -> 100
             DAS_alertLog, FC_alertLog and SCS_alertLog recover whole
  2020.8.1                    94.9% (was 87.3%)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The failure this catches is silent. A dispatch whose low bound we misread still
yields a full set of pages, all shifted by the same constant -- coverage looks
perfect, the DBC loads, and every alert decodes as a different alert. 13247
signals were wrong that way and nothing in the tool said so; it took comparing
output against the catalog to notice.

So make that comparison part of the run. An <NODE>_alertLog selects on the alert
code and the signal names carry it (VCBATT2_a192_hvState belongs to page 192),
which gives ~18500 assertions per extraction from a source independent of the
code the pages were recovered from. Mismatches print as a WARNING from both
vapi_layout and candata_to_dbc rather than shipping quietly.

Unpaged is counted separately from wrong: missing a page is a coverage gap,
being on the wrong one is a defect.

  2020.8.1       19 correct, 0 unpaged, 0 wrong
  2022.45.15  11966 correct, 2 unpaged, 0 wrong
  2026.8.3    18566 correct, 2 unpaged, 0 wrong

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2022.45.15 reaches 100.0% of its catalog (26220/26227) and 2026.8.3 99.8%.
Three fixes, all found by looking at what the code actually does rather than at
what the model assumed.

**A left shift bounds the width.** Bits shifted off the top of the register are
gone: `shl eax,0x1f` on a byte load keeps ONE bit, not eight. Without that the
piece is read at full width and the join runs past the end of the field --
GTW_nmDebugWakeUp came out 39 bits instead of 32 and swallowed the three
signals above it, which cost the whole of GTW_vehNm. The 56-bit bootGitHash
fields are unaffected: GCC shifts those through RAX, and knows why.

**`sar` smaller than the preceding `shl`.** With M < K the pair does not move
the field down, it places the high half at sub-register bit (K-M) and
sign-extends from the top of the byte. Read as a right shift of M it kept a
stale `shl` and modelled nothing at all, which was most of the unmodelled
sites. The width comes from the ORIGINAL shift -- `shl eax,6` leaves 2 bits of
the byte inside al and `sar` moves those 2 back down -- so sizing it by the net
shift over-reads: DAS_TE_accMinJ 8 bits where compact.json says 5.

**Prune the worst offender, not the conflicted set.** One over-wide field
collides with everything beneath it, so dropping every signal involved
discards its victims too. Highest conflict degree, then widest, is the one
least likely to be right; bound the DROPS rather than the conflicts, since a
single bad field can conflict with many signals and still be one deletion from
consistent. GTW_vehNm went to a donor over one bad signal, taking ten good ones.

Validated against the same-revision 2020 compact.json, now over a bigger
overlap than before because more signals are recovered at all:

    9735 signals in both (was 8900)
    same decoded bits : 9735 (100.00%)
    genuinely different : 0

  2022.45.15  26170 -> 26220 of 26227 (100.0%), 7 unmodelled sites and nothing
              else; GTW_vehNm, GTW_info and DAS_gpsStatus recover whole
  2026.8.3    40327 -> 40384 of 40455 (99.8%)
  2020.8.1     9711 ->  9736 of 10193 (95.5%)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2020.8.1 goes 95.5% -> 99.2%, and the check against Tesla's own layout data now
covers 10111 of its 10193 catalogued signals -- all of them decoding to exactly
the same bits.

Every <NODE>_alertMatrix reaches page 0 through a branch we were searching
instead of reading:

    cmp dl, 1
    je  page1
    jb  page0        ; can only be taken when dl == 0

`jb` over a single value is not a subtree, it IS the page. Queued as a range to
walk it found nothing, and page 0 is the biggest page in the message -- 287 of
2020's signals had no page for want of this one branch. Recording it also stops
the walk wandering into a page body, where an ordinary `cmp al,5` on the
selector register can invent a page.

Two smaller ones:

* `movbe` is a load that byte-reverses on the way in -- the same Motorola
  statement `bswap` after a plain load makes, folded into one instruction. It
  was not modelled at all, so SCCM_infoAppCrc, ESP_infoApplicationCRC and
  GTW_uptimeSeconds emitted nothing rather than a big-endian 32-bit field.
* On a tie, prefer the dispatch model that WALKED the code. It resolved the
  same pages and also knows which payload loads were still live at each branch;
  a page body that opens on a register loaded before the dispatch (`shr al,4`
  with no load in the block) models as nothing without those seeds.

  2020.8.1     9736 -> 10112 of 10193 (99.2%), and both the page-attribution
               and overlap gaps close entirely: 81 unmodelled sites, nothing else
  2026.8.3    40384 -> 40397 of 40455 (99.9%)
  2022.45.15  unchanged at 100.0%

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2020.8.1 now recovers all 10193 catalogued signals with no gaps of any kind,
and 10192 of them decode to exactly the same bits as Tesla's own compact.json.

A dispatch preamble often narrows the payload byte it loaded and leaves the
RESULT live for the page bodies. DAS_telemetryEvent computes its jump table in
rdx/rcx precisely so that `eax = byte0 >> 5` survives the branch, and all ten
of its pages open by using it:

    movzx eax, byte ptr [rbx]
    ...
    shr   al, 5              ; <- the value the pages read
    lea   rdx, [rip + table] ; table in rdx, NOT rax
    jmp   rdx

The tracker treated that `shr` as destroying the load, so every page was seeded
with nothing and its first signal modelled as nothing. Carry shifts, masks and
self-truncations through instead, and drop the register only on something the
model genuinely cannot follow (a shift by a register, say).

That exposed a second bug: the per-page snapshot was `dict(pristine)`, which
copies the dict but SHARES the Field objects. It never showed while the tracker
only ever deleted fields; now that it mutates them, a later `shr` reached
backwards into pages already recorded, and page 1 was seeded with page 2's
state. Snapshots now copy the fields.

  2020.8.1     10112 -> 10193 of 10193 (100.0%), nothing dropped at all
  2026.8.3     40397 of 40455 (99.9%)
  2022.45.15   26220 of 26227 (100.0%)

Validation against the same-revision compact.json grew with the coverage:
10192 signals compared (was 8900 when this started), 100.00% same bits, 0
different.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
BMS_kwhCounter is the last case in address order, so no following case bounds
it and find_dispatch scanned the 1.2 MB to the end of the function, latching
onto VCBATT2_alertLog's table. It claimed nine pages at VCBATT2's own
addresses; VCBATT2's stores then landed in a region owned by the wrong message
and 22 of them were left with no page at all.

Two guards, either of which would have caught it:

* Bound the case-body scan. An in-line case body is small -- the largest in
  2026.8.3 is under 1.3 KB -- so 4 KB is generous, and only the LAST case was
  ever unbounded.
* A message cannot have more pages than it has signals to put on them.
  BMS_kwhCounter has two signals and was claiming nine. The catalog knows the
  count, so pass it in and use it.

  2026.8.3  40397 -> 40419 of 40455 (99.9%), pages unattributed 24 -> 2
            alertLog pages 18584 -> 18606 correct, still none wrong

2020 and 2022 are unchanged, and both still agree with compact.json on every
signal they share (10192 and 192, 100.00% same bits).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every message case ends by telling the manager the frame arrived, and a mux
whose selector has a single live value is written as one branch over exactly
that block:

    movzx eax, byte ptr [rbx]
    test  al, 0xf
    je    page0
    mov   edx, 0x3e6        ; the frame arrived, nothing was decoded
    mov   esi, 3
    call  CANDataManager::messageArrived

A lone comparison is refused, and rightly -- a plausibility check on a payload
byte looks identical. What tells them apart is where the OTHER side goes: over
the sign-off block, the case has given up, which no ordinary check does. So
find_compare_chain now takes the storeSignalValue PLT and accepts a single
target when the branch's other side reaches a call having stored nothing.

29 messages in 2026.8.3 were affected. Their signals were emitted as if
unconditional when they are only valid on one selector value, and the twelve
whose page body opens on a register loaded BEFORE the branch (`shr al,4` with
no load in the block) modelled nothing at all for want of a seed.

Also read `test al,0xf` as the selector mask it is -- the `and` form without
the write-back, which is all GCC needs when there is one page -- and stop
inferring a fall-through page when the fall-through is the sign-off: with a
one-bit selector that handed page 1 to a block that decodes nothing.

  2022.45.15  26220 -> 26227 of 26227 (100.0%), no gaps of any kind
  2026.8.3    40419 -> 40432 of 40455 (99.9%), unmodelled sites 27 -> 14
              alertLog pages 18606 -> 18608 correct, unpaged 2 -> 0, none wrong

2020.8.1 stays at 10193/10193. Checked against Tesla's own compact.json in the
same image: 10192 + 192 + 281 signals compared across the three revisions, all
still 100.00% on the decoded bits, and the mux ids now agree on 10609 of them
with none disagreeing. compact.json is what settles this shape independently --
it gives GTW_hrl mux ids 1 and 2, and the branch says 2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
TCU_log and TCU_alertLog convert straight out of memory:

    cvtsi2sd xmm1, dword ptr [rbx + 4]
    xorpd    xmm0, xmm0
    addsd    xmm0, xmm1              ; a MOVE, written as an add

Only the register form of cvtsi2sd was followed, so all six of those signals
modelled as nothing. cvtsi2sd reads a SIGNED integer, which is what the load is.

Two smaller things the same shape needed. `xorpd`/`xorps` on a register with
itself is `pxor` by another name -- GCC picks between them by which unit is
free -- and adding INTO a zeroed register is a move, not an unknown constant,
so the scale survives it. And an xmm register stops being zero once something
is written to it, or a later `mulsd` would be read against a value that has
already moved on.

  2026.8.3  40432 -> 40438 of 40455 (99.9%), unmodelled sites 14 -> 8

2020.8.1 and 2022.45.15 stay at 100%, and all three still agree with
compact.json on every signal they share.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signal keys are unique, so two messages can never share a store. RCM_collision's
compare chain walked out of its own case body and into GTW_info's dispatch, and
resolved four perfectly good pages -- all four of them GTW_info's. GTW_info was
then left flat, and with no pages to separate them every one of its seven
signals overlapped every other: two were dropped and the whole message was
handed to the donor.

So read the first store at each claimed page and check whose signal it is. This
is the same class as the BMS_kwhCounter fix, caught by a different invariant --
that one bounded the SCAN, this one checks the RESULT, and either would have
caught the other's case.

  2026.8.3  40438 -> 40445 of 40455 (99.9%), GTW_info fully recovered
            and no message deferred to the donor at all

2020.8.1 and 2022.45.15 stay at 100%. All three still agree with compact.json on
every signal they share, and on the mux id of all 10609 that carry one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`add ax,0x3ff` / `and ax,0x3ff` is ONE modular rebase, not a fresh field, so
the bias has to carry through the `and`. VCFRONT1_alertLog bounds its table
with the REBASED value:

    cmp ax, 0x10a / ja subtree      ; alerts above 266
    test ax, ax   / je default      ; alert 0
    add ax, 0x3ff / and ax, 0x3ff   ; -1 mod 1024
    cmp ax, 0x109 / ja default      ; ...so this is alert 266, not 265

Zeroing the bias put that bound one short. The single-value split then had
exactly one alert unaccounted for and handed 266 to the sign-off block, so when
the jump table was expanded a moment later 266 was already claimed and its two
signals were left with no page at all.

find_dispatch has read the pair correctly since the table-low-bound work; this
is the same rule, in the walker that follows the chain rather than the one that
reads the table.

  2026.8.3  40445 -> 40447 of 40455 (99.9%)
            alertLog pages 18608 -> 18610 correct, still none wrong

2020.8.1 and 2022.45.15 stay at 100%, and all three still agree with
compact.json on every signal and every mux id they share.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The bit layouts were right and the SCALES were quietly wrong: an unreadable
scale falls back to 1, so 12% of 2026.8.3's signals decoded as raw counts. The
bit comparison against compact.json says nothing about this, which is why it
went unnoticed; comparing scale and offset too is what found it.

Three shapes:

* The 0.0 GCC adds to a scaled signal comes from an integer register --
  `xor r12d,r12d` ONCE at the top of crackMessage, then `movq xmm6,r12` in each
  block that needs it. r12 is callee-saved, which is exactly why it is parked
  there, so its zero survives the 40455 calls in between.
* The scale need not be an operand of the multiply: GCC loads it into its own
  register and multiplies the VALUE into that -- `movsd xmm1,[rip+X]` /
  `mulsd xmm1,xmm0`. Looking only at the source operand found nothing.
* Nothing reset the running scale at a block boundary. GCC's unsigned-64-to-
  double idiom ends `addsd xmm0,xmm0` / `movq rax,xmm0` / `jmp`, and that
  doubling leaked forward into every block until the next store: GTW_hrlPagesCount
  came out scaled by 2^24, DIR_Vsx by 2^26. Nothing falls through an
  unconditional transfer, so the accumulated state is cleared there.

And _snap_float now leaves a constant alone when its exact decimal is already
short. 0.0732421875 is 75/1024 written out in full; float32 also accepts
0.07324219, and snapping to it threw away digits compact.json still carries.

  2026.8.3  unreadable scales 4922 -> 4
  2020.8.1  unreadable scales   37 -> 1

Measured against Tesla's compact.json in the same image, 2020.8.1 now agrees on
10166 of 10192 scales exactly and 25 more to float32 precision. The one left is
ESP_pForceBlendingMC, where the decoder multiplies by a register it has just
zeroed -- the code says 0 and compact.json says 0.25, so Tesla's own two
artifacts disagree and we keep reporting the scale as unknown.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
GCC merges page bodies that decode the same field under different keys, leaving
stubs that do nothing but name the signal:

    movzx eax, byte ptr [rbx+1]
    cvtsi2sd xmm0, eax
    test  byte ptr [rbx], 1     ; the mux bit
    mulsd xmm0, [rip+X]
    je    left_stub             ; mov esi,<key> / call storeSignalValue

The linear sweep arrives at the stub long after the value has been reset, so
UI_ventPanelLeftPositionX modelled as nothing. The branch that reached it is the
only place the value can have come from, and a stub is recognisable outright:
it stores without loading, converting or scaling anything.

Also separate the last class from the gaps. The generator emits one store per
CATALOGUED signal, so a field this revision's frame does not carry is still
stored -- from a constant, with no payload read at all. Six of DI_alertLog's
track voltages are that, and reporting them as an unmodelled decode said we had
failed to follow something the code plainly does not do.

  2026.8.3  40447 -> 40448 of 40455 (99.98%)

What is left in 2026.8.3 is 6 signals the decoder itself stores as constants and
ONE genuinely unmodelled: VCFRONT_isAdaptiveHighBeamAvailable, whose page is
chosen by a 3-bit selector split across byte 4 bit 7 and byte 5 bits 0-1. That
field is not contiguous in either byte order, so it cannot be written as a DBC
multiplexor even once it is read.

2020.8.1 and 2022.45.15 remain at 100% with no gaps of any kind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
vapi_layout recovers bit positions by reading GUICanCracker::crackMessage out
of libQtCarVAPI. That is a model, and a wrong layout is silent -- it decodes,
it just decodes the wrong thing. So run the real function instead.

The whole thing hinges on storeSignalValue(key, double, valid, bus) being an
UND import: it leaves through the PLT, so the decode can be watched by mapping
libQtCarVAPI's own sections into Unicorn and stubbing every PLT slot, with none
of the 103 libraries it links (Qt, gRPC, protobuf, abseil, ICU) present. What
comes back is key -> value, and the catalog names the keys, so no bit position
is involved at runtime at all.

Measured on 2026.8.3: 0 faults across all 580 catalog ids, 0.65 ms a frame,
and against the generated DBC over 1740 random payloads -- 19784 the same, 4
different, 755 only-emu, 212 only-dbc. The 4 are real DBC bugs (0x247
DAS_controlDistance decodes at 2x its true scale); the 212 are DBC signals the
firmware does NOT store for that payload, false positives on a mux page that no
static check can see. `vapi_emu.py parity` is that harness.

Threads scale NEGATIVELY here -- 1341 -> 735 -> 165 decodes/s at 1/2/4 -- because
the emulator calls back into Python on every PLT hit. So parallelism is
processes: 6162 decodes/s at 4 workers, started lazily on the first decode so a
tool that only ENCODES never pays the warmup, and every worker warmed behind a
barrier so the second one does not pay it inside the first batch.

VapiDatabase subclasses CanDatabase and overrides decode alone; encode_frame
and everything vehicle_sim, ecu_bench and tesla_frames stand on is inherited
untouched, because crackMessage only goes one way. can_live's flusher collects
a tick and awaits ONE batch rather than decoding inline. A faulted decode falls
back to the recovered layouts, TM3_VAPI=0 and --no-vapi force them, and the
firmware's own per-signal valid flag -- which the layout decoder cannot produce
-- now stops a bad store latching a fault.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
VCFRONT_lightStatus picks its page with a 3-bit index that straddles a byte
boundary:

    movzx eax, byte ptr [rbx + 5]
    mov   edx, eax
    and   edx, 3                  ; byte 5 bits 0-1
    lea   ecx, [rdx + rdx]        ; ...shifted up one
    movzx edx, byte ptr [rbx + 4]
    shr   dl, 7                   ; byte 4 bit 7 -- the LSB
    or    edx, ecx
    sub   dl, 1                   ; ZF iff the index is 1
    je    page1                   ; ...and the fall-through gives up

In Intel numbering that selector is contiguous -- bits 39..41, which the
extraction already recovers as VCFRONT_lightStatusMuxIndex -- so the message is
perfectly expressible. The compare chain simply tracks a selector as a byte plus
a mask, and this one is neither. Rather than teach that walker a second model
(it is the most delicate code here and has cost regressions before), read this
shape on its own terms and offer it as another candidate.

Deliberately narrow: the assembled field must SPAN A BYTE BOUNDARY, which is
exactly what the chain cannot express, and the branch's other side must be the
case's bare sign-off. A single-byte selector is left to the chain, so the two
models never claim the same dispatch.

The `or` join is now one helper shared with extract_stores rather than two
copies of the same rule.

  2026.8.3  40448 -> 40449 of 40455 (99.985%)

VCFRONT_isAdaptiveHighBeamAvailable was the gap, but five more signals were
worse off than missing: isAdaptiveHighBeamFaulted, isBendLightingAvailable,
advancedDrivingBeamStatus, headlampBendAngle and headlampRoadClass were being
emitted as unconditional when they are only valid on page 1.

All that is left in 2026.8.3 is six signals the decoder itself stores from a
constant. 2020.8.1 and 2022.45.15 stay at 100%, all three still agree with
compact.json on every signal and every mux id they share, and the alert-name
oracle still reports none wrong and none unpaged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The bench host names every extraction <rev>.ice with nothing appended, so
rev_from_root handed back "2022.4.15.ice" and the DBC it names could never be
found -- config.ETH_DBC stayed None and every tool silently dropped to the
stripped compact.json.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
can_read(bus="ETH") maps the bus token through the developer .env, but the
frames were filed under a hardcoded "can0". On the bench host, where the
vehicle bus IS can1, all six reads looked up an empty channel and every
assertion failed on None. Nothing to do with the code under test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
spawn re-imports the parent __main__ in every worker, so any caller whose
script does its work at module level runs it again, once per worker. Caught on
the bench: a capture script without an if __name__ guard opened the DU CAN bus
four times and printed its report four times. ecu_bench.py has no guard either.

forkserver hands the child only what it is given, and like spawn (unlike plain
fork) gives it a clean interpreter -- which matters because the pool is built
lazily, quite possibly after can_live has started its reader threads. Measured
head to head it is a wash on throughput and slightly faster to warm: 6227 vs
6375 decodes/s, 0.84s vs 1.00s warmup.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
3e1c49a claimed forkserver stops a worker re-importing the caller's __main__.
It does not. Measured, unguarded and guarded, both ways:

    spawn       guarded=False -> body ran 4x
    spawn       guarded=True  -> body ran 4x
    forkserver  guarded=False -> body ran 4x
    forkserver  guarded=True  -> body ran 4x

Both methods run spawn.prepare() in the child, which re-imports __main__ as
__mp_main__ regardless of start method or preload, so module-level code runs
once per worker either way -- and an `if __name__ == "__main__"` guard spares
only the guarded block, not the module body. What actually fixed the bench
capture script was moving the CAN-bus open inside a function; the guard alone
would not have. ecu_bench.py needs no guard either: its only top-level
statement is its docstring.

forkserver stays, on its real merit -- it forks each worker from a warm server
rather than exec'ing a fresh interpreter, so it warms faster (0.84s vs 1.00s
for four) with throughput a wash, and still gives a clean interpreter, which
plain fork would not. The preload is narrowed to this module so workers inherit
it with capstone and unicorn instead of each importing them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
compact.json spells every node lowercase (vcfront) where the .so catalog uses
VCFRONT. enrich() takes the node from the catalog when it has the message and
from the donor otherwise -- so the *_udsRequest/*_udsResponse pairs, which
Tesla ships ONLY in compact.json, landed under the donor spelling and gave 25
ECUs a near-empty twin node holding just those two.

That is not cosmetic: can_live lists nodes from the database, so picking
vcfront showed one message where VCFRONT has 23. 2022.45.15 goes from 85 nodes
with 33 case-duplicate pairs to 58 with none.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The rich per-field alertLog decode (the alert panel) took its geometry from
our hand-recovered packer layouts (alertlog_layouts_<rev>.json). Take it from
the firmware instead: the active CanDatabase already decodes the whole
multiplexed alertLog page with libQtCarVAPI's own crackMessage (VapiDatabase
under emulation, or the geometry recovered statically out of it), so the alert
decoder now consumes those <NODE>_aNNN_<field> rows. In lockstep with the
loaded build, nothing modelled here.

AlertLogDecoder.decode() gains a `signals` param and maps the db rows via the
new decoded_from_signals(); can_live passes in the rows it already computed
(alertlog ids added to the flusher decode set, VapiDatabase cache covers
repeats). CAN-rationality alerts keep their firmware-independent name-resolving
path. The packer JSON (_LAYOUTS) is now an OFFLINE fallback only -- used when
no firmware decode is available or it yields nothing -- not a cross-check; the
extractor/build_layouts stay as the standalone JSON exporter.

Proof on the a162 bench frame: both paths agree except timeSinceCruiseCancel
(.so 60 vs JSON 120 -- a real off-by-one-bit in the hand-recovered layout the
firmware's decoder gets right), exactly the lockstep win this is for.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
VapiDatabase was built on from_dbc, so the shim only engaged where someone had
already run candata_to_dbc -- a 90-second build per host for something decoding
never needed. It doesn't: libQtCarCANData carries every name, unit, enum table,
DLC, cycle time and originNode the decode path renders, and the emulator
supplies the values. No bit position is involved anywhere in it.

So build from the catalog and layer a layout source on top only for the two
things running the decoder cannot do: encode_frame, and decoding a frame whose
emulation faulted. The generated DBC when there is one, else compact.json, and
an unreadable one is now a warning rather than the loss of the shim. A message
only the layout source has -- the *_udsRequest/*_udsResponse pairs -- is kept
whole, so coverage never shrinks by preferring the catalog.

Measured against a DBC-backed database on 2022.45.15, 446 shared messages:

    same 4875   different 5   only-catalog 0   only-dbc 0

The five are *Crc / *BootGitHash signals whose VALUES agree exactly; only the
hex label is absent, because rendering bytes needs a width and the catalog has
none. That also fixes a latent crash: _is_hex_signal reads a missing width as
0 % 8 == 0, so those signals took the hex path and then KeyError'd on it.

The catalog's node names are the canonical spelling, so a catalog-first
database cannot reproduce the case-duplicate nodes fixed in a232168 at all:
50 nodes, 0 duplicate pairs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
renderFaults and renderLog wiped their container and rebuilt every card, on a
1 Hz poll -- so the node the pointer was over was destroyed and recreated once
a second, dropping and re-applying :hover. The content rarely changes between
polls, so compare a signature of it first and leave the DOM alone when it has
not. Shared by the Alerts tab and the Drive HUD, so both stop flickering.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
_overlay_0x118 re-read DI_systemState, DI_immobilizerState, accel pedal and
traction mode straight out of the raw bytes using a per-revision table in
di.py, and it ran AFTER the decode -- so it overwrote whatever the database had
produced. The table is picked by TM3_FW, which is unset on every bench, and an
unset rev clamps to the 2020 baseline. On the 2022 unit here that put
DI_immobilizerState at bit 27, where 2022 keeps DI_hvilSystemStatus: the HUD
was showing one signal labelled as another. It only looked right because both
read 0 on a dormant unit, exactly as di.py's own comment warned.

There is nothing left for it to recover. The database decodes 0x118 with the
loaded firmware's OWN decoder, which cannot drift from the firmware the way a
table here can -- all six signals, with their enum labels:

    DI_systemState 2.0 STANDBY      DI_immobilizerState 0.0 INIT_SNA
    DI_driveBlocked 3.0 FALCON      DI_hvilSystemStatus 0.0 DISABLED
    DI_accelPedalPos 1.6            DI_tractionControlMode 5.0 DYNO_MODE

So the overlay goes, and with it _DI_0X118_LAYOUTS, di_0x118_layout and
_DI_0X118_RECOVERED, which had no other caller. A layout borrowed from another
revision is worse than a gap: the gap is visible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
TM3_ROOT is the extraction being read; TM3_FW is the revision vehicle_sim
TRANSMITS, and may deliberately be older -- an older car emulated for an older
ECU under test. config._resolve_eth_dbc already had that precedence, and says
so, but the two places that WRITE the artifacts it reads back had it inverted:

    candata_to_dbc._resolve_rev  args.rev or FW_VERSION or _rev_from_lib(...)
    so_alerts                    a.rev   or FW_VERSION or _rev_from_lib(...)

So TM3_FW=2020.8.1 made a DBC built from a 2026 root come out named
Model3_ETH.2020.8.1.dbc, which config would then load as though it were the
2020 database -- every layout borrowed from the wrong release, silently. Now:

    $ TM3_FW=2020.8.1 TM3_ROOT=.../2022.45.15.ice.extracted candata_to_dbc extract
    libQtCarCANData.so: rev=2022.45.15 ...

_rev_from_lib returned the string "unknown" when it could derive nothing, which
is truthy, so simply reordering would have made TM3_FW unreachable rather than
a fallback. It now returns None and each caller supplies its own last resort.

alert_log._layouts_path keeps FW_VERSION deliberately: those layouts come from
DRIVE-UNIT firmware, so the bench ECU's revision is the right one to match, not
the MCU extraction's.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The bundle is split by PLATFORM, not just by car, and only Model3/tasks was
ever scanned -- so on 2022.45.15 the 625 entries under Gen3 were invisible,
including every DIR/DIF resolver-learn procedure. A Model 3 with a Gen3 drive
unit runs those.

list_platforms reports the trees a bundle carries (a tree counts if it has
tasks/; Common/ is a library, pulled in by reference). The picker sits in the
ODIN sidebar and is remembered in localStorage; the server chooses the first
one (TM3_PRODUCT when the bundle has it, else Model3), which is also the answer
to there being no way to say Model3 vs ModelY anywhere in the UI before.

A procedure basename already carries its tree (Gen3/tasks/PROC_...), so
requirements and run needed nothing. The proc cache did: one cache for all
trees would have served Model3s list for Gen3.

list_procedures now raises on a missing entries dir rather than globbing an
absent directory into an empty list, which read as "no procedures" instead of
"wrong place".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A car is a thin specialisation over shared trees, not an alternative to them:
Model3s own tasks reference Gen3/ 432 times and Common/ 84. So picking Model3
now lists Model3 + Gen3 + Common together, which is what makes the DIR/DIF
resolver-learn procedures reachable at all -- they live in Gen3, and with
Model3/tasks hardcoded nothing ever scanned it (625 entries to Model3s 569).

The names collide and the content differs, so precedence matters: Model3 and
Gen3 share 217 task names on 2022.45.15 and 172 of them DIFFER. The cars own
tree wins, and runnability is filtered AFTER that choice so a blocked proc
cannot fall through to another trees version of the same name.

The picker sits in the ODIN sidebar and is remembered in localStorage -- also
the first way to say Model3 vs ModelY anywhere in the UI. Basenames already
carry their tree (Gen3/tasks/PROC_...), so requirements and run needed nothing.

Correcting 15b9f0d: I claimed Common/ was a library pulled in by reference. It
is not -- it has 212 entries of its own. The rule is the tasks/ dir, not the
name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
It has tasks/ (6 sample graphs) and is not one of the shared trees, so it fell
through the vehicle list and the picker offered it alongside Model3 and ModelY.
It is neither a car nor something a car layers on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
REAR/FRONT Drive Unit Resolver Error Learning died with "'str' object has
no attribute 'values'". Its task wrapper calls Gen3/scripts/PROC_DIX_X_
RESOLVER-ERROR-LEARN, and that file's `network` is not a dict of nodes at
all -- it is a Python SOURCE STRING defining

    async def odin_script_test(api, nodeName, hasEsp=True, hasPgear=True)

which drives the same UDS/ODX/CID interop as awaited calls on an `api`
facade instead of wired ports. 154 of the 1,955 files on 2022.45.15 are
this form; the runner assumed every one was a graph, so _has() called
.values() on the source text.

odin_script_api is that facade, mapped onto the same Backend the graph
handlers use, so both forms run against one bench with one set of
adapters. Verified against the real bundle: the rear task now walks the
whole choreography -- drive-rail gate, tester-present context, extended
session, security access LEVEL_5, RESOLVER_LEARNING start, the ESP 0xf00a
velocity-limit disable, the gear-D poll loop, teardown -- and exits 0.

Two things a straight call-forward would not have given it:

* A virtual clock. Scripts import `sleep` and `time` themselves and write
  real deadline loops (poll for up to 95 seconds, sleeping 3 between).
  Speeding up sleep alone leaves such a loop spinning against the wall
  clock for the full 95 seconds, hammering the ECU. A sleep not taken in
  real time is instead added to an offset the script's own clock reports,
  so the loop advances and exits. Installed through the script's own
  __builtins__.__import__ -- per script, no sys.modules mutation.
* run_coroutine, because a script can reach another script through a
  WIRED graph in between, and that middle hop is synchronous.

api.http.*, api.modem.* and api.cid.car_server_request are deliberately
NOT implemented: they reach the Tesla mothership or the car's LTE modem,
and a stub would make a connectivity test pass for the wrong reason.
required()/supported() name what a script needs against what is here, so
those scripts report BLOCKED the way an unhandled node type does --
including the one script the firmware itself ships with a syntax error.

The same honesty applied to the requirements readout, which said "no
static CAN reads detected" for every script: script_requirements reads
the api calls, resolving both the aliases scripts bind up front and the
literal the CALLING task bound to a parameter. PROC_DIR_ and PROC_DIF_
are one script with nodeName='DIR'/'DIF', so without that binding the
readout would omit the one ECU the task is about. The rear task now
reports nodes DIR+ESP and MCU values VAPI_driveRailOn/VAPI_shiftState.

Also landed the graph-side node types the same audit surfaced -- all of
them, bar the ones needing http/modem: reporting.ServiceOutput (the wired
twin of a script's verdict, terminal) + ServiceOutputExitReason,
scripts.ScriptTest, cid.IsFused/SetFactoryMode/GetPlatform/
SetVehicleConfig/ShowAuthoredPopup/GetGatewayFile, cidupdater.Command,
odin.GetMetadataPath, misc.WhitelistDict, enum.EnumInput,
control.ConcurrentSplit, testing.PingOutcome, reporting.FileOutput,
messages.Broadcast, lists.ExtendMulti, apupdater.ClearCache.

cid.StartHRL/StopHRL now record a REAL capture instead of no-opping: a
python-can Logger on the vehicle channel's existing notifier, so it rides
the socket can_read already has open, written to TM3_HRL_DIR. The upload
step reports where the file landed rather than claiming an upload that
has nowhere to go; an operator-specified endpoint is the next step.

cid.IsFused is operator-declared, not inferred -- it reads GUI_isFused
from the CID store, which run_procedure's new cid_values seeds.

Model3 goes 868 -> 918 runnable of 1,021, the four resolver-learn
procedures among them. What stays blocked is http (91 procedures), the
car server (8), the modem (4), an OTA install, Tesla PKI, and that one
broken script.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
cid.IsFused asks whether the gateway is fused -- a production car -- or
unfused, a factory/service unit. Several procedures branch on it before
writing factory mode or a vehicle config, and nothing on the bus answers
it, so it is the operator's to declare rather than ours to guess.

A "fused gateway" checkbox sits under the vehicle picker. It posts to
/api/odin/bench-state, which run_procedure seeds into the CID store as
GUI_isFused before the procedure reads it. The server owns the truth and
starts on its defaults; the browser remembers the last declaration and
re-asserts it on load, so a reload does not silently revert what the
operator said. An unknown key is a 400, not a quiet no-op.

The endpoint is a map, not a single flag -- BENCH_STATE_CID pairs each
declared fact with the CID data value it sets, so the next one that turns
up is a row, not a new route.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
FRONT Drive Unit Resolver Error Calibration listed "UDS target nodes DIR,
ESP" -- the REAR inverter, on the front task -- and "1 dynamic signal
unresolved".

Both come from the same place. The shared libs are written once and
parameterised: Gen3/lib/DI_RESOLVER_LEARNING takes a node_name Input
defaulting to 'DIR', its odx call is {'connection': 'node_name.value',
'value': 'DIR'}, and its CAN signal is a strings.Concat of that input
with '_axleSpeed'. The readout only ever looked at a field's own value,
so it reported the lib's declared default and could not fold the concat.
PROC_DIF_X_RESOLVER-LEARN binds node_name='DIF'; nothing consulted it.

The walk now carries what the caller bound, and a field's connection is
followed through the node types that actually source one -- surveyed, not
guessed: constant.Constant (147 of the bundle's node_name connections),
networks.Input (129), strings.Concat. A field carrying both a connection
and a value keeps that value as the FALLBACK, so nothing that resolved
before stops resolving. A loop item or a run-time variable still counts
as dynamic rather than being guessed at, which is what keeps the list an
honest lower bound.

Across the 1,021 Model3-reachable procedures: UDS node rows 221 -> 566
(a constant-sourced node_name resolved to nothing before, so it was
dropped entirely), signals 63 -> 72, unresolved 34 -> 25, and 371
readouts change. Twenty of those named the WRONG ECU, not merely too
few -- the panel was telling the operator to put the wrong unit on the
bench:

    PROC_DIF_X_RESOLVER-LEARN       DIR,ESP        -> DIF,ESP
    PROC_PMF_X_STORE-DATA-BOOT      DIR,PMR        -> PMF
    PROC_PMR_X_RESTORE-DATA-BOOT    DIR            -> PMR
    PING-DTC_IBST_X_PARTY           ESP            -> IBST
    PROC_VCRIGHT_..._TRUNK-LED      VCLEFT         -> VCRIGHT

_static_inputs also learned the bare binding form ("node_name": "DIF"),
which is how the front resolver task writes it -- the wrapped form alone
would have missed exactly the case that started this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The 55 UPDATE_* procedures -- UPDATE_PMR, UPDATE_PCS, UPDATE_VCSEC-WITH-
BOOTLOADER, the lot -- all reach the drive unit through one binary:
Gen3/scripts/UPDATE_MODULE shells out to the MCU's flasher as

    /sbin/smashclicker -h <hwidacq list> -u <update list> -j <job id>

Every one of them reported RUNNABLE and returned PASS, because
cid_execute answers any shell-out with canned success. UPDATE_MODULE
reads exit_status and scrapes stdout, so each one declared a flash that
never happened. That is worse than being blocked.

Two things made the substitution small. Gen3/lib/FIRMWARE_DOWNLOAD
downloads nothing -- nine inputs and a pass-through -- so the per-ECU
path has no Tesla-network dependency anywhere in it. And smashclicker's
component names ARE FirmwareEntry.component, so dfu's signed-metadata
selection takes them directly, with no translation table: `-h pmr -u
pmr,dir` selects the two entries under pmr:<packed_key>.

flash_scripts._headless is the seam. It reuses dfu's own identity read,
firmware selection and phase-4 sequence; what it adds is choosing by
component instead of by prompt, and reporting precisely what it could not
place -- a component this ECU has no firmware for, or firmware with no
validated flash sequence. Either fails the request. A flasher that
silently does nothing is the failure this whole path exists to remove.

Worth recording: ODIN independently corroborates the flash model dfu
derived from firmware. It appends bl/bu to an explicit parent list rather
than testing for an 'bl' suffix -- the same way, and for the same reason,
dfu does (epbl/epbr are apps that end in 'bl'). Its ordering matches too:
UPDATE_VCSEC-WITH-BOOTLOADER is bu, bl, app, ramapp. UPDATE_PCS is the
dual-CPU pair. dfu covers 46 of the 71 components the tasks name; the gap
is mostly body ECUs, plus `dif` and `pmf` -- we built the rear path only.

FLASHING IS ARMED, NOT DEFAULT. A UPDATE_* procedure asks for a flash
with no confirmation of its own, so BenchBackend.allow_flash starts false
and the ODIN pane carries an "allow ECU flashing" toggle, red while on.
Unarmed, the seam refuses in smashclicker's own format, so the procedure
reports a real failure with a real reason.

Progress, because a flash takes minutes and silence reads as a hang:

  * StatusDisplay grew set_progress(current, total, label) and both
    transfer loops call it, so a non-terminal driver gets real byte counts
    instead of scraping them back out of a rendered bar.
  * QuietDisplay forwards those, and each step, to the run's event stream.
  * messages.ProgressUpdate / StatusUpdate (and their api.* twins) emit
    there too, so a procedure that narrates itself drives the same bar.
  * The UI shows it determinate when there is a number and indeterminate
    -- a sliding band, honestly meaning "working, no total" -- when there
    is only a step name.

A run holds a server-side lock, so it now runs behind a scrim over the
ODIN pane, with the step and the bar in a dialog. Scoped to that pane on
purpose: watching the live bus during a flash is often the point, so the
other tabs stay reachable.

Cancel is new, and honest about its reach: /api/odin/cancel sets a
run-level flag every frame shares, which ends the next WAIT -- a CID poll,
a routine's result poll, a timeout loop, including inside a nested
subnetwork. It does not interrupt work in flight, and a flash deliberately
never checks it, because stopping a transfer mid-write is how an ECU gets
bricked. The response says so rather than implying otherwise. The flag is
separate from each frame's own cancel, which only ends that graph's
background branches -- sharing one would have let the first graph to
finish stop the whole run.

Also: an unbound script input is now None rather than an error. ODIN's own
unbound networks.Input reads as None and the scripts are written for it --
UPDATE_MODULE takes `job_id: str` with no default and normalises a
non-numeric one to ''. Rejecting it refused a procedure the real tool runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three bugs from the first real bench flash of UPDATE_PMR.

**The ECU was flashed TWICE.** The result said "pmr update PASSED, pmr
update PASSED, dir update PASSED, dir update PASSED" because the signed
metadata carries a row per condition combination and most of those rows
name the same image: pmr:440467458 has two `pmr` rows differing only in
vdcType, both pointing at PMR_26-65-2_M3_Single_VDC_crc_lithium-signed.bhx.
select_entries took every match. dfu's interactive path never hits this --
it prompts when dest_names collide -- but headless it meant writing the
same image back to back, which is not harmless: a flash burns a count
against FLASH_COUNT_LIMITS. Identical images now collapse to one; rows
naming DIFFERENT images are a choice we cannot make, so they come back as
`ambiguous` and nothing is written. Nothing is written either if any
requested component could not be placed: a half-flashed set is a worse
state to leave an ECU in than an untouched one.

**The log was one line per transfer block.** A 216 KB image is ~865
TransferData blocks, and each one called set_detail with a freshly
rendered progress bar, so the collected log was 865 redraws of the same
bar with the actual steps buried in them -- and all of it went into the
service output and up to the browser.

The layering was wrong: a step and its progress are different things.
StatusDisplay.set_progress now renders the bar (terminal output is
byte-identical), the transfer loops call set_detail once per segment and
set_progress per block, and QuietDisplay does not record progress as a
log line at all. 865 blocks now leave 3 log lines.

**The modal hung after the run finished.** Same root cause, other end:
every one of those set_details emitted a status event, and _run cannot
return until the websocket pump has drained -- so the bar showed 100% and
then sat there pushing ~1,700 events. Status events are down to one per
step, and progress is throttled to a real change (1%) or a real interval
(250 ms), never per block: 865 blocks now emit 100 events, and the last
one is never dropped.

Which row applies is the car's own configuration, so it is read rather
than ignored. Every condition key in the map -- chassisType,
drivetrainType, vdcType and the sixteen others -- is a GTW_carConfig
(0x7FF) signal named GTW_<key>, so BenchBackend.vehicle_conditions decodes
them straight off the bus. A key the bus never carried stays ABSENT, which
find_firmware reads as "no constraint"; inventing a default would silently
pick somebody else's firmware. On a drive-unit bench there is no gateway
(and vehicle_sim is what transmits 0x7FF, so reading it back is reading
our own assertion), so an operator declaration overrides it -- a "car
config" field beside the arming toggle, `k=v` pairs.

And it is previewed, not inferred afterwards: the resolved plan --
component, image, crc, and the config that chose it -- goes out as a
flash_plan event BEFORE the first write, and the UI renders it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Car config was in the sidebar next to "fused gateway" and "allow ECU
flashing". Those two are standing facts about the unit on the bench;
which firmware row applies is not -- it is a decision belonging to one
flash, and putting it in the sidebar meant it was out of sight when Run
was pressed, with the plan arriving only once the first image was already
going down the wire.

It now sits in a preflight dialog that gates the run. Press Run on an
UPDATE_* procedure and the server resolves what would be written --
reading the ECU's identity and matching the signed metadata, writing
nothing -- and the operator confirms the actual images, or cancels. A
procedure that touches no firmware has nothing to confirm and runs
straight through.

Knowing WHICH procedures flash needed the static walk to carry a task's
bindings across a pass-through lib: Gen3/lib/FIRMWARE_DOWNLOAD is nine
Inputs relayed into UPDATE_MODULE, so UPDATE_PMR's update_list was
invisible one hop later. _static_inputs now follows a connection through
the caller's own bindings, using the same resolver the requirements
readout uses (moved into odin_coverage, where the graph structure lives).
flash_targets reads the component lists off that descent -- UPDATE_PMR ->
['pmr','dir'] through PMR at ACCESSORY_PLUS -- and list_procedures marks
each entry `flashes`, riding the walk it already does.

And the config is picked, not typed. A `k=v` text box asked the operator
to know both the key names and that drivetrainType 0 means RWD. The
dialog now shows one <select> per condition key that ACTUALLY decides this
flash -- of the map's nineteen, pmr:440467458 varies only in vdcType,
while an ESP varies in five -- with the firmware's own labels from
compact.json where it has a table (MODEL_3_CHASSIS, RWD/AWD) and the raw
value where it does not. Each carries an explicit "from the car" option,
because leaving a key unset (GTW_carConfig decides, or nothing does) is a
different answer from pinning a value. Changing one re-resolves, so the
plan on screen is always the plan.

Confirm is disabled unless the plan is complete AND flashing is armed --
an operator should not be able to start a flash already known to be
incomplete, and should learn that the backend cannot flash here rather
than at the point of no return.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three things wrong with the preflight modal on the bench.

The pickers showed bare numbers. A condition key IS a GTW_carConfig signal
-- the metadata's `vdcType` is `GTW_vdcType` on 0x7FF -- but the only
source consulted was compact.json, which carries just the subset Tesla
ships to the diagnostic tool: on 2026.8.3 it names chassisType and
drivetrainType and nothing else, so the ESP's five keys and the PMR's one
came out as 0/1. The generated DBC is built from that same firmware's
libQtCarCANData and carries the whole catalog
(`VAL_ 2047 GTW_vdcType 0 "BOSCH_VDC" 1 "TESLA_VDC"`), so it is preferred
and compact.json fills its gaps: 19 of the 20 keys the metadata actually
uses are now named, up from 2. The 20th resolves by folding case, because
the metadata writes brakeHWType on most rows and brakeHwType on the ESP's
while the signal is GTW_brakeHWType -- safe to fold HERE and only here,
since this decides what a value is CALLED, never which row matches.
Neither source is required and a key nothing names keeps its raw value: an
unlabelled number is still answerable, and better than withholding the
choice. Read as text rather than through cantools -- the DBC has ~26k
signals and this wants one message's value tables. dfu.py shares the
loader, so it gains the same names.

The modal said "flashing is not armed" for an armed bench. Preflight
builds its OWN backend from the factory, which starts disarmed, and only
the run path applied the operator's declaration -- so `armed` was always
false and Confirm could never enable. Both paths now go through one
_apply_bench_state, and flash_preflight takes allow_flash for the case
where the backend is named rather than passed, which is the only place
arming could be applied at all. It is reported, never acted on; nothing
here writes.

And the dropdowns were unreadable. `--panel2`, `--line` and `--dim` were
never defined, so the select fell back to a translucent grey the browser
composited against white -- a light popup with light text. The aliases now
resolve to the palette, the select uses the opaque --bg/--text pair the
rest of the pane uses, and color-scheme: dark points the native controls
at the page rather than the OS default.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
UPDATE_PMR flashes ['pmr', 'dir'] -- and that IS the PCS-family dual-CPU
pair (pmr is a primary, dir a secondary), so the whole flash goes through
run_pcs_dual_cpu rather than the per-image loop in _flash. That function
took no display and built its own, so every step and every transfer block
went to whatever terminal happened to be attached while the QuietDisplay
the caller passed saw only the one "flash (dual-CPU) ..." header it had
written itself. Hence a working bar in the server's terminal and nothing
but an indeterminate slider in the web UI: the progress events were never
emitted, because the object emitting them was not the object doing the
transfer.

It now takes `display` and threads it into the FlashContext, so the steps
report to the caller's stream. The three print() calls become set_detail,
for the same reason -- a bare print reaches a terminal and nobody else.

Each CPU's image opens with set_header rather than a detail line, which
starts a new progress span. Without that the second transfer's first
report (0% right after the first closed at 100%) is a -100 delta inside
the throttle's quarter-second window and gets dropped, leaving the bar
pinned at 100% for the whole of CPU1.

dfu.py passes its own display through too; it had been getting a second,
unrelated one for this leg.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
dfu.py asks this at the terminal (_prompt_bootloader_choice) and driving
it from ODIN did not, so a -WITH-BOOTLOADER procedure wrote its bu/bl with
nothing between the operator and a bootloader flash but the plan table.

The preflight dialog now carries the choice, alongside the car config and
for the same reason: it belongs to one flash, not to the bench. It appears
only when the resolved plan actually holds bootloader entries -- of the
real UPDATE_* lists, VCSEC has them and pmr/dir and pcs/pcscpu2 do not, so
most flashes are not asked a question they have no answer for.

The default runs the OTHER WAY from dfu's, deliberately. There a
bootloader turns up as a surprise from metadata matching, so NO is the
safe default. Here the operator ran UPDATE_VCSEC-WITH-BOOTLOADER and named
the bu/bl by hand; defaulting it off would silently downgrade that to a
plain app update, which is the same quiet no-op this path exists to
remove. The dialog is the confirmation dfu gets from its prompt.

What declining costs is spelled out rather than left to be recalled: each
entry is named with its role (bu = updater flashed into the app slot, bl =
the image the agent writes) and the note says which apps restore normal
operation afterwards -- or warns that none follow, so the ECU boots the
update agent until one is flashed. The phrasing takes a list because real
VCSEC has two (vcsec + vcsecramapp).

An exclusion is a narrowing the operator chose, not a component that could
not be placed, so it lands in `excluded` and leaves ok=True -- unlike
unplaced/ambiguous/no_script, which block the write. It is appended to the
run's log and shown on the plan, because nothing else would ever say the
procedure did less than its name promises.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Same gap as the bootloaders: dfu.py asks how to handle a procedure's RAM
apps (_prompt_ramapp_choice) and driving it from ODIN did not, so
UPDATE_VCSEC-WITH-BOOTLOADER pushed vcsecramapp with nothing between the
operator and it.

A select, not a checkbox, because there are genuinely three answers. A RAM
app is not part of a normal app update, so writing ONE WITHOUT rewriting
the app it rides on -- push the OPC gateway applet, leave the PMR app
alone -- is a real bench request that "include/exclude" cannot express.
`only` is offered solely when the plan holds something else to leave out;
where a procedure names nothing but RAM apps, include and only are the
same list and choosing between them is a choice about nothing.

The note carries dfu's OPC caveat, and only where it applies: pushing
pmramapp to service an OPC is a wasted step on gen26, where the CAN<->LIN
gateway is already resident in the PMR app and opc/opcs flash THROUGH it.
It is attached to pm-family RAM apps rather than every RAM app, so it
means something when it appears.

And dfu's own prompt was wrong. It offered two labels for three branches:
"RAM app(s) only -- flash just the RAM app" returned app + RAM app, and
the `return rams` branch that meant it was unreachable, so the one option
promising just the RAM app flashed the app as well. The third label is
restored, which is also what the ODIN side mirrors.

An exclusion stays a narrowing, not a failure -- ok=True, recorded in
`excluded` and on the plan. The message is now "excluded by the operator"
rather than naming bootloaders, since either control can be the reason.

`ramapps` is the first bench-state key that is not a flag, so the endpoint
grew BENCH_STATE_ENUMS: an unknown value is rejected with a 400 instead of
being coerced to a bool, which would have read "only" and "yes" alike as
True and silently meant something else.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The signed 2026 PMR resident bootloader auto-boots into the app on a timer the
older permissive bl never had: FUN_00086bb0 hands off when now - DAT_b03e >=
DAT_b03c, armed to 0x32 (~50 ms from a bare reset) and bumped to 2000 (~2 s) on
every UDS request, at a 1 ms CPU-Timer0 tick (PRD 200000 @ 200 MHz). The 3E80
flood only covers wait_for_bootloader itself, so a bootloader-context read
(0xF180 identity, a slow erase) issued >2 s later silently landed in the app.

- wait_for_bootloader() starts the 2 Hz TesterPresent keep-alive on the 7E
  confirm so the ~2 s window stays open through the rest of the bl session.
- ecu_reset()/ecu_reset_no_wait() stop the keep-alive first -- a running
  keep-alive would feed the bl and fight the reset (hold it / delay app boot);
  wait_for_bootloader re-establishes it only when we want to hold the bl.
- start_tester_present() is now idempotent so flows that also start it don't
  double-launch.

The old 12603 bl left the window functions as empty stubs (never timed out),
which is why the same host tooling held the bootloader before the signed flash.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
cpPlcFw/cpPlcPib were mapped to module byte 0x00, which the CP bootloader
validates against the app window [0x8000,0xe0000]; a RequestDownload at the
modem regions (cpPlcFw @0x100000, cpPlcPib @0xe0000) fails NRC 0x31.

Per the CP bootloader RequestDownload window validator (cpbl 2026,
FUN_00001d3c) the WDBI 0x0102 module byte gates the allowed address range:
0x08 -> [0x100000,0x200000] (cpPlcFw), 0x06 -> [0xe0000,0x100000] (cpPlcPib).

Resolves through dfu.py phase4 and the can_live ODIN UPDATE_PLC path
unchanged (both take the module byte from get_script). Fails safe: a wrong
byte NRCs, it cannot mis-target another region.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Reviewed-on: #15
Merge private/main into feat/vapi-layout
Some checks failed
Tests / test (pull_request) Has been cancelled
7f8f45eba6
Clears the conflicts from main's can_live -> tm3web and tm3diag.py -> tm3cli.py
renames, which landed after this branch forked:

- .env.example: keep this branch's TM3_ALERTLOG_LAYOUTS text. The layouts are an
  offline fallback here, not the primary decode path -- fields now come from the
  firmware's own decoder.
- docs/can_live.md: main deleted it in favour of docs/tm3web.md, so this branch's
  alertLog-decode paragraphs are ported into that file rather than kept alongside.
- tests/test_can_live.py -> tests/test_tm3web.py: this branch's tests (which add
  the _decode_batch cases) under main's name; main's own edit was purely the rename.

Also updates stale can_live prose references in files main's rename never saw, and
combines a nested if ruff flags as SIM102 (pre-existing on both tips).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
outlandnish/tm3diag!16
No description provided.