feat(vapi): recover CAN layouts from the MCU decoder, and default every tool to the generated DBC #16
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "feat/vapi-layout"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Recovers CAN signal layout from the MCU's own decoder and makes the generated DBC
the default signal database everywhere.
vapi layout recovery —
vapi_layout.pywalks libQtCarVAPI's dispatch tree torecover bit layout for ~98% of 26227 signals;
vapi_emu.pyruns the firmware'sdecoder under Unicorn rather than modelling it. Validated against compact.json at
100% bit agreement on 2020, with mux pages cross-checked against alert names on
every extraction.
The catalog is the database — layouts are an optional overlay, and every tool
now defaults to the generated DBC instead of the partial compact.json (which
carries only a fraction of the catalog). alertlog decodes its fields from the .so
by default, with packer JSON as fallback.
ODIN runner — flashes through dfu for real instead of faking success, picks the
platform tree and lists a car's procedures, confirms a flash with pickers before it
starts, and lets the operator decline bootloader updates or declare on the bench
what the bus cannot report.
Also: CP PLC subcomponents route to their own flash regions, the signed
bootloader's handoff window is held open with TesterPresent, and a batch of
config/test/UI fixes.
53 commits, 52 files, +12207/-309.
Note: the branch is 9 commits behind
mainat time of opening.🤖 Generated with Claude Code
libQtCarCANData.so carries the complete signal catalog but no bit-layout, and Model3_ETH.compact.json -- which does -- is stripped down every release (2022: 347 signals vs the catalog's 26227). candata_to_dbc was therefore borrowing layout from OLDER revisions, which is unsound: layouts move between releases. 0x118 DI_immobilizerState is 27|3 in 2020 but 13|3 in 2022, DI_driveBlocked 12|2 -> 24|2, and bits 27-28 became DI_hvilSystemStatus -- so the 2020 table read hvilSystemStatus AS immobilizerState, silently, since both are 0 on a dormant bench unit. The layout does exist same-revision, as code, in libQtCarVAPI.so: CAN frame -> GUICanCracker::crackMessage(bus, msgId, payload) -> CANDataManager::storeSignalValue(key, double, valid, bus) crackMessage is one generated function (2022: ~1.1 MB) whose cases load payload bytes, shift/mask the field, scale it, and call storeSignalValue with a 32-bit KEY -- and that KEY is the u32 at +0x08 in the catalog's signal struct, which recovers the name. The call count equals the catalog signal count exactly, so every catalogued signal has exactly one decode site. vapi_layout.py disassembles it with capstone and models the field arithmetic, emitting nothing rather than a guess wherever the model does not fully follow the sequence. Multiplexed messages are most of the catalog, and GCC compiles their page switch three ways -- jump table, if-else compare chain, and a chain that walks the page number down with `sub` -- all three are recognised, as is `ja default` being a binary split rather than the end of the chain. Measured against the same-revision 2020 DBC (real ground truth): 98.4% of start|width exact, 99.99% of signedness, 99.5% of scale+offset, 22 genuine Intel disagreements in 8767 signals. Coverage 98.0% of the 2022 catalog, 86.0% of 2020; the generated 2022 DBC goes from 21368 to 25768 signals. candata_to_dbc auto-detects the sibling VAPI library and prepends it as the highest-priority donor (--no-vapi opts out), and di.py's 0x118 overlay becomes revision-aware, selecting floor-with-clamp on the DIR's firmware. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>bswapas Motorola byte order, and validate against compact.json 236f876cbcEvery one of the 117 Motorola signals the 2020 extraction got wrong was a single unmodelled instruction: mov eax, dword ptr [rbx + 4] bswap eax ; ESP_infoApplicationCRC -> 39|32 @0 All 117 already had the right WIDTH and were emitted LITTLE with an LSB start -- the byte reversal states the byte order outright, rather than it having to be inferred from which byte supplies the high half of a split field. The DBC start of a Motorola field is its MSB, which after the swap is bit 7 of the first payload byte. A swap of anything already shifted or masked is refused, since then the field is not the whole load. Motorola disagreements 117 -> 46, and all 46 that remain cover exactly the same BITS as compact.json -- a field inside a single byte has an equivalent Motorola spelling (start + width - 1, @0), so they decode identically and there is nothing in the binary that says which notation to prefer. Validation no longer goes through a generated DBC. It now compares against the raw same-revision Model3_ETH.compact.json: Tesla's own layout DATA, shipped in the same image as the CODE being decoded, so the two are independent. Of the 8781 signals in both: 8706 exact (99.15%), 8759 same bits (99.75%), 22 genuinely different. The docstring now also records what this cannot check -- compact.json covers 8781 of 10193 catalogued signals, so the ~1400 only this recovers are exactly the ones no DBC could verify. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>sarandcwde-- 100% bit agreement with compact.json 45a9c99f13The `or` handler required the low side to sit at bit 0, which only expresses a TWO-way join. PMR_bootGitHash takes bytes 1, 2-3 and 4-7, and at every intermediate step of a chain like that BOTH sides carry a shift: mov eax, dword ptr [rbx+4] ; shl rax,0x18 or rdx, rax ; <- rdx already holds bytes 2-3 at bit 8 movzx eax, byte ptr [rbx+1] or rax, rdx so the whole field was refused. Each side occupies result bits [shl, shl+width); they join iff those runs are adjacent, and the result keeps the lower side's payload origin and offset. The two-way case is the same rule with lo.shl == 0, so it is unchanged. Unmodelled decode sites 2022 140 -> 66, 2020 249 -> 165. Coverage 2022 98.4% -> 98.5%, 2020 86.5% -> 87.3%. Agreement with compact.json holds at 100% of the bits over 80 more comparable signals (8900). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>The overlap guard had one response to a message it could not make consistent: discard it. On 2026.8.3 that cost 1243 signals across 22 messages -- including all 458 of VCFRONT2_alertLog, of which 444 had been recovered perfectly. The condemning signals were 14 stragglers whose mux page we failed to attribute. An unpaged signal in a multiplexed message is read on EVERY page, so one whose page we missed collides with all of them and trips the pervasive-overlap rule on its own. But a genuinely mux-independent signal -- a counter, a checksum -- sits in bits no page uses and conflicts with nothing. Conflict is what tells those two apart, not the absence of a page, so prune on that. Guards now prune in order of how well corroborated a signal is: unpaged-and-conflicting < ordinary signal < the multiplexor The multiplexor comes last because it is the best-evidenced signal in the message: the dispatch we read the pages from is a branch on exactly its bits. It also cannot be replaced -- without it the pages cannot be expressed at all and the message has to go, which is what happened to SDCR_info when SDCR_infoAppCrc landed on the selector's byte and both were dropped. Only when conflicts survive all of that is the message handed to a donor. 2026.8.3 gaps 1346 -> 647 VCFRONT2_alertLog 0 -> 444 signals VCFRONT1_alertLog 0 -> 162 VCBATT2_alertLog 0 -> 85 SDCR_info 0 -> 8 2022.45.15 gaps 366 -> 334 UI_alertLog recovers 32 signals Both revisions load clean; no message is left with stranded pages. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>A message whose pages hang on ONE payload bit gets neither of the shapes we looked for. There is nothing to build a jump table from and nothing to compare, so GCC emits a bare branch: test byte ptr [rbx + 6], 1 ; VC_pcsInterfaceMuxIndex, bit 48 je page0 ; ... and page 1 is the fall-through With no dispatch found, every signal landed on one flat page, the pages collided, and the whole message was discarded. VCFRONT_vehicleStatus loads the byte first and tests the register (`test al,1`), so both operand forms count; the existing `test` rule only matched `test dl,dl` on an already-identified selector, and worse, dropped the selector register on anything else. Only a single-bit mask qualifies. `test` separates zero from nonzero, which is an unambiguous two-way split exactly when the field is one bit wide. Two details the real code forces: * The bit test REDEFINES the selector, so a preceding `and` must not bound it. 0x441 masks byte 6 with 0xf to read an unrelated signal immediately before testing bit 0 of that same byte; inheriting that 0..15 bound leaves 15 pages unaccounted for and the fall-through page is never attributed. * The walk continues past the branch into the page body, where further `and`s are ordinary field extraction. The selector is pinned when it branches so they cannot be folded into the reported mask. name_multiplexors now locates the selector signal from any contiguous mask rather than assuming it is low-aligned in its byte. TAS_axleData 0 -> 13 signals, and reads as it should: page 0 front (rawHeightFL/FR), page 1 rear (RL/RR), on TAS_axleIndex VC_pcsInterface 0 -> 13, VCBATT_pcsInterface 0 -> 13 VCFRONT_vehicleStatus 0 -> 34, VCLEFT_switchStatus 9 -> 75, VCLEFT_windowStatus 5 -> 31, UI_systemMonitor 0 -> 12 2026.8.3 gaps 647 -> 420 2022.45.15 gaps 334 -> 149 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>An alertLog's mux id IS its alert number, and the signal names carry it (VCBATT2_a192_hvState is alert 192), so the catalog is its own oracle. It said 13247 of 18592 alertLog signals had the wrong page -- every message off by its own constant. Not a missing page: a page numbered wrong, which silently decodes one alert's payload as another alert's. Confirmed against the mux_id compact.json does ship (21 signals across five revisions, all agreeing with the name). The constant is the table's low bound, and GCC subtracts it three ways. Only one was modelled: add ax, 0x3f1 / and ax, 0x3ff TRCM_alertLog -15 mod 1024 add eax, 0x6d / cmp al, 9 EPAS3S_alertLog -147 mod 256, the truncation to 8 bits IS the modulus sub eax, 0x79 OCS1P_alertLog base alert 121 The modular forms were not read at all, and the plain `sub` was guarded by str.isdigit(), which capstone's hex operands never satisfy -- so a bound of 0x79 parsed as 0 while a bound of 9 parsed fine. Arithmetic on rbx/rsp/rbp is excluded: `sub rsp,0x28` in a case body is a frame adjustment, not the selector being rebased. (_canon returns family names -- "b", "sp" -- not register names, which is what makes that guard actually fire.) The top-level message switch is untouched: its own bound stays 17 and all 446 / 580 cases still resolve to real catalog message ids in both revisions. 2026.8.3 13247 wrong -> 0 (18351 correct, 241 unpaged) 2022.45.15 0 (11946 correct, 22 unpaged) Every remaining unpaged alert lies OUTSIDE the id range of the pages we found, with none inside: those are the sibling subtrees of a binary search tree whose other branches we do not yet follow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>A big switch is not one table, it is a TREE of range splits with tables at the leaves. VCBATT2_alertLog tests alert 227 on its own, hands everything above it to one subtree and everything below 0x3b to another, and tables only the run in between: cmp ax, 0xe3 / je 0x6f16ad ; alert 227 ja 0x6c5e4e ; 228.. -> subtree cmp ax, 0x3a / jbe 0x70c277 ; ..58 -> subtree add ax, 0x3c5 ... jmp rax ; 59..192 -> table Reading only the root left 128 of that message's 209 signals with no page at all -- and an unpaged signal does not simply go missing, it inherits whichever mark precedes it in the address space, so VCBATT2 alerts were landing in regions owned by LS_lightShow and BMS_kwhCounter. The walker gains what the tree needs: `jbe`/`jb` split downward as `ja`/`jae` already split upward, each run now carries its whole [lo, hi] range so a sibling inherits the right bound, selectors may be read at any WIDTH (an alert id is a 10-bit field read as a word), and a table met mid-walk is expanded in place rather than ending the walk. Three things this forced, each found by regression rather than reasoning: * Comparisons are reduced modulo the selector mask only PAST a rebase, i.e. when the bias is negative. Reducing unconditionally folds ordinary comparisons in page bodies onto real pages -- with a one-bit selector `cmp al,5` becomes page 1 -- which cost TAS_axleData and both PARK_psc*Slot. * A second register holding the same payload bytes is added to the selector set, not swapped in. PARK_pscEnvSlot reads byte 0 into eax and then word 0 into ecx before comparing `al`; replacing the set loses the register the chain actually branches on. * The selector is pinned when the first page resolves, and pages above its mask are refused. The running selector drifts as the walk continues into the page bodies, and a bogus page is worse than a missing one because it claims a code region belonging to something else. Also: a rebase's modulus is the mask wherever the mask is visible, even when the selector load is not -- a subtree is entered below it, and falling back on the register width read `add ax,0x2d4` as base 64812 rather than 300. 2026.8.3 40035 -> 40287 of 40455 (99.6%) gaps 420 -> 168 alertLog pages 18351 -> 18568 correct, 241 -> 24 unpaged 2022.45.15 26078 -> 26127 of 26227 (99.6%) gaps 149 -> 100 DAS_alertLog, FC_alertLog and SCS_alertLog recover whole 2020.8.1 94.9% (was 87.3%) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>2022.45.15 reaches 100.0% of its catalog (26220/26227) and 2026.8.3 99.8%. Three fixes, all found by looking at what the code actually does rather than at what the model assumed. **A left shift bounds the width.** Bits shifted off the top of the register are gone: `shl eax,0x1f` on a byte load keeps ONE bit, not eight. Without that the piece is read at full width and the join runs past the end of the field -- GTW_nmDebugWakeUp came out 39 bits instead of 32 and swallowed the three signals above it, which cost the whole of GTW_vehNm. The 56-bit bootGitHash fields are unaffected: GCC shifts those through RAX, and knows why. **`sar` smaller than the preceding `shl`.** With M < K the pair does not move the field down, it places the high half at sub-register bit (K-M) and sign-extends from the top of the byte. Read as a right shift of M it kept a stale `shl` and modelled nothing at all, which was most of the unmodelled sites. The width comes from the ORIGINAL shift -- `shl eax,6` leaves 2 bits of the byte inside al and `sar` moves those 2 back down -- so sizing it by the net shift over-reads: DAS_TE_accMinJ 8 bits where compact.json says 5. **Prune the worst offender, not the conflicted set.** One over-wide field collides with everything beneath it, so dropping every signal involved discards its victims too. Highest conflict degree, then widest, is the one least likely to be right; bound the DROPS rather than the conflicts, since a single bad field can conflict with many signals and still be one deletion from consistent. GTW_vehNm went to a donor over one bad signal, taking ten good ones. Validated against the same-revision 2020 compact.json, now over a bigger overlap than before because more signals are recovered at all: 9735 signals in both (was 8900) same decoded bits : 9735 (100.00%) genuinely different : 0 2022.45.15 26170 -> 26220 of 26227 (100.0%), 7 unmodelled sites and nothing else; GTW_vehNm, GTW_info and DAS_gpsStatus recover whole 2026.8.3 40327 -> 40384 of 40455 (99.8%) 2020.8.1 9711 -> 9736 of 10193 (95.5%) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>2020.8.1 goes 95.5% -> 99.2%, and the check against Tesla's own layout data now covers 10111 of its 10193 catalogued signals -- all of them decoding to exactly the same bits. Every <NODE>_alertMatrix reaches page 0 through a branch we were searching instead of reading: cmp dl, 1 je page1 jb page0 ; can only be taken when dl == 0 `jb` over a single value is not a subtree, it IS the page. Queued as a range to walk it found nothing, and page 0 is the biggest page in the message -- 287 of 2020's signals had no page for want of this one branch. Recording it also stops the walk wandering into a page body, where an ordinary `cmp al,5` on the selector register can invent a page. Two smaller ones: * `movbe` is a load that byte-reverses on the way in -- the same Motorola statement `bswap` after a plain load makes, folded into one instruction. It was not modelled at all, so SCCM_infoAppCrc, ESP_infoApplicationCRC and GTW_uptimeSeconds emitted nothing rather than a big-endian 32-bit field. * On a tie, prefer the dispatch model that WALKED the code. It resolved the same pages and also knows which payload loads were still live at each branch; a page body that opens on a register loaded before the dispatch (`shr al,4` with no load in the block) models as nothing without those seeds. 2020.8.1 9736 -> 10112 of 10193 (99.2%), and both the page-attribution and overlap gaps close entirely: 81 unmodelled sites, nothing else 2026.8.3 40384 -> 40397 of 40455 (99.9%) 2022.45.15 unchanged at 100.0% Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>2020.8.1 now recovers all 10193 catalogued signals with no gaps of any kind, and 10192 of them decode to exactly the same bits as Tesla's own compact.json. A dispatch preamble often narrows the payload byte it loaded and leaves the RESULT live for the page bodies. DAS_telemetryEvent computes its jump table in rdx/rcx precisely so that `eax = byte0 >> 5` survives the branch, and all ten of its pages open by using it: movzx eax, byte ptr [rbx] ... shr al, 5 ; <- the value the pages read lea rdx, [rip + table] ; table in rdx, NOT rax jmp rdx The tracker treated that `shr` as destroying the load, so every page was seeded with nothing and its first signal modelled as nothing. Carry shifts, masks and self-truncations through instead, and drop the register only on something the model genuinely cannot follow (a shift by a register, say). That exposed a second bug: the per-page snapshot was `dict(pristine)`, which copies the dict but SHARES the Field objects. It never showed while the tracker only ever deleted fields; now that it mutates them, a later `shr` reached backwards into pages already recorded, and page 1 was seeded with page 2's state. Snapshots now copy the fields. 2020.8.1 10112 -> 10193 of 10193 (100.0%), nothing dropped at all 2026.8.3 40397 of 40455 (99.9%) 2022.45.15 26220 of 26227 (100.0%) Validation against the same-revision compact.json grew with the coverage: 10192 signals compared (was 8900 when this started), 100.00% same bits, 0 different. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>BMS_kwhCounter is the last case in address order, so no following case bounds it and find_dispatch scanned the 1.2 MB to the end of the function, latching onto VCBATT2_alertLog's table. It claimed nine pages at VCBATT2's own addresses; VCBATT2's stores then landed in a region owned by the wrong message and 22 of them were left with no page at all. Two guards, either of which would have caught it: * Bound the case-body scan. An in-line case body is small -- the largest in 2026.8.3 is under 1.3 KB -- so 4 KB is generous, and only the LAST case was ever unbounded. * A message cannot have more pages than it has signals to put on them. BMS_kwhCounter has two signals and was claiming nine. The catalog knows the count, so pass it in and use it. 2026.8.3 40397 -> 40419 of 40455 (99.9%), pages unattributed 24 -> 2 alertLog pages 18584 -> 18606 correct, still none wrong 2020 and 2022 are unchanged, and both still agree with compact.json on every signal they share (10192 and 192, 100.00% same bits). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>Every message case ends by telling the manager the frame arrived, and a mux whose selector has a single live value is written as one branch over exactly that block: movzx eax, byte ptr [rbx] test al, 0xf je page0 mov edx, 0x3e6 ; the frame arrived, nothing was decoded mov esi, 3 call CANDataManager::messageArrived A lone comparison is refused, and rightly -- a plausibility check on a payload byte looks identical. What tells them apart is where the OTHER side goes: over the sign-off block, the case has given up, which no ordinary check does. So find_compare_chain now takes the storeSignalValue PLT and accepts a single target when the branch's other side reaches a call having stored nothing. 29 messages in 2026.8.3 were affected. Their signals were emitted as if unconditional when they are only valid on one selector value, and the twelve whose page body opens on a register loaded BEFORE the branch (`shr al,4` with no load in the block) modelled nothing at all for want of a seed. Also read `test al,0xf` as the selector mask it is -- the `and` form without the write-back, which is all GCC needs when there is one page -- and stop inferring a fall-through page when the fall-through is the sign-off: with a one-bit selector that handed page 1 to a block that decodes nothing. 2022.45.15 26220 -> 26227 of 26227 (100.0%), no gaps of any kind 2026.8.3 40419 -> 40432 of 40455 (99.9%), unmodelled sites 27 -> 14 alertLog pages 18606 -> 18608 correct, unpaged 2 -> 0, none wrong 2020.8.1 stays at 10193/10193. Checked against Tesla's own compact.json in the same image: 10192 + 192 + 281 signals compared across the three revisions, all still 100.00% on the decoded bits, and the mux ids now agree on 10609 of them with none disagreeing. compact.json is what settles this shape independently -- it gives GTW_hrl mux ids 1 and 2, and the branch says 2. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>TCU_log and TCU_alertLog convert straight out of memory: cvtsi2sd xmm1, dword ptr [rbx + 4] xorpd xmm0, xmm0 addsd xmm0, xmm1 ; a MOVE, written as an add Only the register form of cvtsi2sd was followed, so all six of those signals modelled as nothing. cvtsi2sd reads a SIGNED integer, which is what the load is. Two smaller things the same shape needed. `xorpd`/`xorps` on a register with itself is `pxor` by another name -- GCC picks between them by which unit is free -- and adding INTO a zeroed register is a move, not an unknown constant, so the scale survives it. And an xmm register stops being zero once something is written to it, or a later `mulsd` would be read against a value that has already moved on. 2026.8.3 40432 -> 40438 of 40455 (99.9%), unmodelled sites 14 -> 8 2020.8.1 and 2022.45.15 stay at 100%, and all three still agree with compact.json on every signal they share. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>Signal keys are unique, so two messages can never share a store. RCM_collision's compare chain walked out of its own case body and into GTW_info's dispatch, and resolved four perfectly good pages -- all four of them GTW_info's. GTW_info was then left flat, and with no pages to separate them every one of its seven signals overlapped every other: two were dropped and the whole message was handed to the donor. So read the first store at each claimed page and check whose signal it is. This is the same class as the BMS_kwhCounter fix, caught by a different invariant -- that one bounded the SCAN, this one checks the RESULT, and either would have caught the other's case. 2026.8.3 40438 -> 40445 of 40455 (99.9%), GTW_info fully recovered and no message deferred to the donor at all 2020.8.1 and 2022.45.15 stay at 100%. All three still agree with compact.json on every signal they share, and on the mux id of all 10609 that carry one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>`add ax,0x3ff` / `and ax,0x3ff` is ONE modular rebase, not a fresh field, so the bias has to carry through the `and`. VCFRONT1_alertLog bounds its table with the REBASED value: cmp ax, 0x10a / ja subtree ; alerts above 266 test ax, ax / je default ; alert 0 add ax, 0x3ff / and ax, 0x3ff ; -1 mod 1024 cmp ax, 0x109 / ja default ; ...so this is alert 266, not 265 Zeroing the bias put that bound one short. The single-value split then had exactly one alert unaccounted for and handed 266 to the sign-off block, so when the jump table was expanded a moment later 266 was already claimed and its two signals were left with no page at all. find_dispatch has read the pair correctly since the table-low-bound work; this is the same rule, in the walker that follows the chain rather than the one that reads the table. 2026.8.3 40445 -> 40447 of 40455 (99.9%) alertLog pages 18608 -> 18610 correct, still none wrong 2020.8.1 and 2022.45.15 stay at 100%, and all three still agree with compact.json on every signal and every mux id they share. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>GCC merges page bodies that decode the same field under different keys, leaving stubs that do nothing but name the signal: movzx eax, byte ptr [rbx+1] cvtsi2sd xmm0, eax test byte ptr [rbx], 1 ; the mux bit mulsd xmm0, [rip+X] je left_stub ; mov esi,<key> / call storeSignalValue The linear sweep arrives at the stub long after the value has been reset, so UI_ventPanelLeftPositionX modelled as nothing. The branch that reached it is the only place the value can have come from, and a stub is recognisable outright: it stores without loading, converting or scaling anything. Also separate the last class from the gaps. The generator emits one store per CATALOGUED signal, so a field this revision's frame does not carry is still stored -- from a constant, with no payload read at all. Six of DI_alertLog's track voltages are that, and reporting them as an unmodelled decode said we had failed to follow something the code plainly does not do. 2026.8.3 40447 -> 40448 of 40455 (99.98%) What is left in 2026.8.3 is 6 signals the decoder itself stores as constants and ONE genuinely unmodelled: VCFRONT_isAdaptiveHighBeamAvailable, whose page is chosen by a 3-bit selector split across byte 4 bit 7 and byte 5 bits 0-1. That field is not contiguous in either byte order, so it cannot be written as a DBC multiplexor even once it is read. 2020.8.1 and 2022.45.15 remain at 100% with no gaps of any kind. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>VCFRONT_lightStatus picks its page with a 3-bit index that straddles a byte boundary: movzx eax, byte ptr [rbx + 5] mov edx, eax and edx, 3 ; byte 5 bits 0-1 lea ecx, [rdx + rdx] ; ...shifted up one movzx edx, byte ptr [rbx + 4] shr dl, 7 ; byte 4 bit 7 -- the LSB or edx, ecx sub dl, 1 ; ZF iff the index is 1 je page1 ; ...and the fall-through gives up In Intel numbering that selector is contiguous -- bits 39..41, which the extraction already recovers as VCFRONT_lightStatusMuxIndex -- so the message is perfectly expressible. The compare chain simply tracks a selector as a byte plus a mask, and this one is neither. Rather than teach that walker a second model (it is the most delicate code here and has cost regressions before), read this shape on its own terms and offer it as another candidate. Deliberately narrow: the assembled field must SPAN A BYTE BOUNDARY, which is exactly what the chain cannot express, and the branch's other side must be the case's bare sign-off. A single-byte selector is left to the chain, so the two models never claim the same dispatch. The `or` join is now one helper shared with extract_stores rather than two copies of the same rule. 2026.8.3 40448 -> 40449 of 40455 (99.985%) VCFRONT_isAdaptiveHighBeamAvailable was the gap, but five more signals were worse off than missing: isAdaptiveHighBeamFaulted, isBendLightingAvailable, advancedDrivingBeamStatus, headlampBendAngle and headlampRoadClass were being emitted as unconditional when they are only valid on page 1. All that is left in 2026.8.3 is six signals the decoder itself stores from a constant. 2020.8.1 and 2022.45.15 stay at 100%, all three still agree with compact.json on every signal and every mux id they share, and the alert-name oracle still reports none wrong and none unpaged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>VapiDatabase was built on from_dbc, so the shim only engaged where someone had already run candata_to_dbc -- a 90-second build per host for something decoding never needed. It doesn't: libQtCarCANData carries every name, unit, enum table, DLC, cycle time and originNode the decode path renders, and the emulator supplies the values. No bit position is involved anywhere in it. So build from the catalog and layer a layout source on top only for the two things running the decoder cannot do: encode_frame, and decoding a frame whose emulation faulted. The generated DBC when there is one, else compact.json, and an unreadable one is now a warning rather than the loss of the shim. A message only the layout source has -- the *_udsRequest/*_udsResponse pairs -- is kept whole, so coverage never shrinks by preferring the catalog. Measured against a DBC-backed database on 2022.45.15, 446 shared messages: same 4875 different 5 only-catalog 0 only-dbc 0 The five are *Crc / *BootGitHash signals whose VALUES agree exactly; only the hex label is absent, because rendering bytes needs a width and the catalog has none. That also fixes a latent crash: _is_hex_signal reads a missing width as 0 % 8 == 0, so those signals took the hex path and then KeyError'd on it. The catalog's node names are the canonical spelling, so a catalog-first database cannot reproduce the case-duplicate nodes fixed ina232168at all: 50 nodes, 0 duplicate pairs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>_overlay_0x118 re-read DI_systemState, DI_immobilizerState, accel pedal and traction mode straight out of the raw bytes using a per-revision table in di.py, and it ran AFTER the decode -- so it overwrote whatever the database had produced. The table is picked by TM3_FW, which is unset on every bench, and an unset rev clamps to the 2020 baseline. On the 2022 unit here that put DI_immobilizerState at bit 27, where 2022 keeps DI_hvilSystemStatus: the HUD was showing one signal labelled as another. It only looked right because both read 0 on a dormant unit, exactly as di.py's own comment warned. There is nothing left for it to recover. The database decodes 0x118 with the loaded firmware's OWN decoder, which cannot drift from the firmware the way a table here can -- all six signals, with their enum labels: DI_systemState 2.0 STANDBY DI_immobilizerState 0.0 INIT_SNA DI_driveBlocked 3.0 FALCON DI_hvilSystemStatus 0.0 DISABLED DI_accelPedalPos 1.6 DI_tractionControlMode 5.0 DYNO_MODE So the overlay goes, and with it _DI_0X118_LAYOUTS, di_0x118_layout and _DI_0X118_RECOVERED, which had no other caller. A layout borrowed from another revision is worse than a gap: the gap is visible. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>TM3_ROOT is the extraction being read; TM3_FW is the revision vehicle_sim TRANSMITS, and may deliberately be older -- an older car emulated for an older ECU under test. config._resolve_eth_dbc already had that precedence, and says so, but the two places that WRITE the artifacts it reads back had it inverted: candata_to_dbc._resolve_rev args.rev or FW_VERSION or _rev_from_lib(...) so_alerts a.rev or FW_VERSION or _rev_from_lib(...) So TM3_FW=2020.8.1 made a DBC built from a 2026 root come out named Model3_ETH.2020.8.1.dbc, which config would then load as though it were the 2020 database -- every layout borrowed from the wrong release, silently. Now: $ TM3_FW=2020.8.1 TM3_ROOT=.../2022.45.15.ice.extracted candata_to_dbc extract libQtCarCANData.so: rev=2022.45.15 ... _rev_from_lib returned the string "unknown" when it could derive nothing, which is truthy, so simply reordering would have made TM3_FW unreachable rather than a fallback. It now returns None and each caller supplies its own last resort. alert_log._layouts_path keeps FW_VERSION deliberately: those layouts come from DRIVE-UNIT firmware, so the bench ECU's revision is the right one to match, not the MCU extraction's. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>REAR/FRONT Drive Unit Resolver Error Learning died with "'str' object has no attribute 'values'". Its task wrapper calls Gen3/scripts/PROC_DIX_X_ RESOLVER-ERROR-LEARN, and that file's `network` is not a dict of nodes at all -- it is a Python SOURCE STRING defining async def odin_script_test(api, nodeName, hasEsp=True, hasPgear=True) which drives the same UDS/ODX/CID interop as awaited calls on an `api` facade instead of wired ports. 154 of the 1,955 files on 2022.45.15 are this form; the runner assumed every one was a graph, so _has() called .values() on the source text. odin_script_api is that facade, mapped onto the same Backend the graph handlers use, so both forms run against one bench with one set of adapters. Verified against the real bundle: the rear task now walks the whole choreography -- drive-rail gate, tester-present context, extended session, security access LEVEL_5, RESOLVER_LEARNING start, the ESP 0xf00a velocity-limit disable, the gear-D poll loop, teardown -- and exits 0. Two things a straight call-forward would not have given it: * A virtual clock. Scripts import `sleep` and `time` themselves and write real deadline loops (poll for up to 95 seconds, sleeping 3 between). Speeding up sleep alone leaves such a loop spinning against the wall clock for the full 95 seconds, hammering the ECU. A sleep not taken in real time is instead added to an offset the script's own clock reports, so the loop advances and exits. Installed through the script's own __builtins__.__import__ -- per script, no sys.modules mutation. * run_coroutine, because a script can reach another script through a WIRED graph in between, and that middle hop is synchronous. api.http.*, api.modem.* and api.cid.car_server_request are deliberately NOT implemented: they reach the Tesla mothership or the car's LTE modem, and a stub would make a connectivity test pass for the wrong reason. required()/supported() name what a script needs against what is here, so those scripts report BLOCKED the way an unhandled node type does -- including the one script the firmware itself ships with a syntax error. The same honesty applied to the requirements readout, which said "no static CAN reads detected" for every script: script_requirements reads the api calls, resolving both the aliases scripts bind up front and the literal the CALLING task bound to a parameter. PROC_DIR_ and PROC_DIF_ are one script with nodeName='DIR'/'DIF', so without that binding the readout would omit the one ECU the task is about. The rear task now reports nodes DIR+ESP and MCU values VAPI_driveRailOn/VAPI_shiftState. Also landed the graph-side node types the same audit surfaced -- all of them, bar the ones needing http/modem: reporting.ServiceOutput (the wired twin of a script's verdict, terminal) + ServiceOutputExitReason, scripts.ScriptTest, cid.IsFused/SetFactoryMode/GetPlatform/ SetVehicleConfig/ShowAuthoredPopup/GetGatewayFile, cidupdater.Command, odin.GetMetadataPath, misc.WhitelistDict, enum.EnumInput, control.ConcurrentSplit, testing.PingOutcome, reporting.FileOutput, messages.Broadcast, lists.ExtendMulti, apupdater.ClearCache. cid.StartHRL/StopHRL now record a REAL capture instead of no-opping: a python-can Logger on the vehicle channel's existing notifier, so it rides the socket can_read already has open, written to TM3_HRL_DIR. The upload step reports where the file landed rather than claiming an upload that has nowhere to go; an operator-specified endpoint is the next step. cid.IsFused is operator-declared, not inferred -- it reads GUI_isFused from the CID store, which run_procedure's new cid_values seeds. Model3 goes 868 -> 918 runnable of 1,021, the four resolver-learn procedures among them. What stays blocked is http (91 procedures), the car server (8), the modem (4), an OTA install, Tesla PKI, and that one broken script. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>FRONT Drive Unit Resolver Error Calibration listed "UDS target nodes DIR, ESP" -- the REAR inverter, on the front task -- and "1 dynamic signal unresolved". Both come from the same place. The shared libs are written once and parameterised: Gen3/lib/DI_RESOLVER_LEARNING takes a node_name Input defaulting to 'DIR', its odx call is {'connection': 'node_name.value', 'value': 'DIR'}, and its CAN signal is a strings.Concat of that input with '_axleSpeed'. The readout only ever looked at a field's own value, so it reported the lib's declared default and could not fold the concat. PROC_DIF_X_RESOLVER-LEARN binds node_name='DIF'; nothing consulted it. The walk now carries what the caller bound, and a field's connection is followed through the node types that actually source one -- surveyed, not guessed: constant.Constant (147 of the bundle's node_name connections), networks.Input (129), strings.Concat. A field carrying both a connection and a value keeps that value as the FALLBACK, so nothing that resolved before stops resolving. A loop item or a run-time variable still counts as dynamic rather than being guessed at, which is what keeps the list an honest lower bound. Across the 1,021 Model3-reachable procedures: UDS node rows 221 -> 566 (a constant-sourced node_name resolved to nothing before, so it was dropped entirely), signals 63 -> 72, unresolved 34 -> 25, and 371 readouts change. Twenty of those named the WRONG ECU, not merely too few -- the panel was telling the operator to put the wrong unit on the bench: PROC_DIF_X_RESOLVER-LEARN DIR,ESP -> DIF,ESP PROC_PMF_X_STORE-DATA-BOOT DIR,PMR -> PMF PROC_PMR_X_RESTORE-DATA-BOOT DIR -> PMR PING-DTC_IBST_X_PARTY ESP -> IBST PROC_VCRIGHT_..._TRUNK-LED VCLEFT -> VCRIGHT _static_inputs also learned the bare binding form ("node_name": "DIF"), which is how the front resolver task writes it -- the wrapped form alone would have missed exactly the case that started this. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>The 55 UPDATE_* procedures -- UPDATE_PMR, UPDATE_PCS, UPDATE_VCSEC-WITH- BOOTLOADER, the lot -- all reach the drive unit through one binary: Gen3/scripts/UPDATE_MODULE shells out to the MCU's flasher as /sbin/smashclicker -h <hwidacq list> -u <update list> -j <job id> Every one of them reported RUNNABLE and returned PASS, because cid_execute answers any shell-out with canned success. UPDATE_MODULE reads exit_status and scrapes stdout, so each one declared a flash that never happened. That is worse than being blocked. Two things made the substitution small. Gen3/lib/FIRMWARE_DOWNLOAD downloads nothing -- nine inputs and a pass-through -- so the per-ECU path has no Tesla-network dependency anywhere in it. And smashclicker's component names ARE FirmwareEntry.component, so dfu's signed-metadata selection takes them directly, with no translation table: `-h pmr -u pmr,dir` selects the two entries under pmr:<packed_key>. flash_scripts._headless is the seam. It reuses dfu's own identity read, firmware selection and phase-4 sequence; what it adds is choosing by component instead of by prompt, and reporting precisely what it could not place -- a component this ECU has no firmware for, or firmware with no validated flash sequence. Either fails the request. A flasher that silently does nothing is the failure this whole path exists to remove. Worth recording: ODIN independently corroborates the flash model dfu derived from firmware. It appends bl/bu to an explicit parent list rather than testing for an 'bl' suffix -- the same way, and for the same reason, dfu does (epbl/epbr are apps that end in 'bl'). Its ordering matches too: UPDATE_VCSEC-WITH-BOOTLOADER is bu, bl, app, ramapp. UPDATE_PCS is the dual-CPU pair. dfu covers 46 of the 71 components the tasks name; the gap is mostly body ECUs, plus `dif` and `pmf` -- we built the rear path only. FLASHING IS ARMED, NOT DEFAULT. A UPDATE_* procedure asks for a flash with no confirmation of its own, so BenchBackend.allow_flash starts false and the ODIN pane carries an "allow ECU flashing" toggle, red while on. Unarmed, the seam refuses in smashclicker's own format, so the procedure reports a real failure with a real reason. Progress, because a flash takes minutes and silence reads as a hang: * StatusDisplay grew set_progress(current, total, label) and both transfer loops call it, so a non-terminal driver gets real byte counts instead of scraping them back out of a rendered bar. * QuietDisplay forwards those, and each step, to the run's event stream. * messages.ProgressUpdate / StatusUpdate (and their api.* twins) emit there too, so a procedure that narrates itself drives the same bar. * The UI shows it determinate when there is a number and indeterminate -- a sliding band, honestly meaning "working, no total" -- when there is only a step name. A run holds a server-side lock, so it now runs behind a scrim over the ODIN pane, with the step and the bar in a dialog. Scoped to that pane on purpose: watching the live bus during a flash is often the point, so the other tabs stay reachable. Cancel is new, and honest about its reach: /api/odin/cancel sets a run-level flag every frame shares, which ends the next WAIT -- a CID poll, a routine's result poll, a timeout loop, including inside a nested subnetwork. It does not interrupt work in flight, and a flash deliberately never checks it, because stopping a transfer mid-write is how an ECU gets bricked. The response says so rather than implying otherwise. The flag is separate from each frame's own cancel, which only ends that graph's background branches -- sharing one would have let the first graph to finish stop the whole run. Also: an unbound script input is now None rather than an error. ODIN's own unbound networks.Input reads as None and the scripts are written for it -- UPDATE_MODULE takes `job_id: str` with no default and normalises a non-numeric one to ''. Rejecting it refused a procedure the real tool runs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>