diff --git a/RESEARCH.md b/RESEARCH.md index 3c9c7e9..e562c53 100644 --- a/RESEARCH.md +++ b/RESEARCH.md @@ -1462,3 +1462,1306 @@ New `U-BOOT.md` in this repo: boot chain, payload layout constants, build/flash/verify command sequences, recovery paths, U-Boot tree state (kept features vs WIP/temporary), rotation and keyboard notes. The user owns the U-Boot config and further development. + +## Round 38 — real display pipeline: DSI host + MIPI TX PHY + BOE panel drivers + +Goal: replace the firmware-handoff revival with a real cold bring-up of +the MT8183 display pipeline (MMSYS → OVL0 → OVL0_2L → RDMA0 → COLOR0 → +CCORR0 → AAL0 → GAMMA0 → DITHER0 → DSI0 → panel), upstreamable, on +krane-updates (commits on top of 1ae771d9f90). + +### Sources ported (numbered findings) + +1. **MMSYS clock gates** (`clk-mt8183.c`): the in-tree clock driver had + NO mmsys provider (grep CLK_MM/mmsys empty). Added CG_CON0 (0x100) + / CG_CON1 (0x110) gate groups with set/clr at +4/+8 and the full + CLK_MM_* gate list, ported from Linux `drivers/clk/mediatek/ + clk-mt8183-mm.c` (v6.6). Legacy vs upstream binding headers number + CLK_MM_* identically (OVL0=19, DSI0_MM=31…), so the C-side gate ids + match the DTB cells. Parents: mm_sel → legacy CLK_TOP_MUX_MM(85), + dpi0_sel → CLK_TOP_MUX_DPI0(110), f26m → CLK_TOP_F26M_CK_D2(4). +2. **GPIO**: no MT8183 pinctrl/gpio driver in-tree (pinctrl-mtk-common + has no mt8183 table). Wrote a minimal `drivers/gpio/mt8183_gpio.c` + (dir/dout/din only, pinmux left to firmware) from the device-era + depthcharge `src/drivers/gpio/mt8183.h` GpioRegs layout: dir[6], + dout[6], din[6] as GpioValRegs (val@0, set@4, rst@8, 16 B/group), + blocks at 0x000/0x100/0x200. Pin 43 = DISP_PWM, 45 = LCM_RST, + 66/166/36 = the three panel rail enables, 176 = PERIPHERAL_EN13. + The set/rst semantics for pins 43/176 were already proven on device + (Round 4). +3. **MIPI TX PHY** (`drivers/phy/phy-mtk-mipi-tx.c`): UCLASS_PHY, + PLL programming + analog lane bring-up ported from Linux + `drivers/phy/mediatek/phy-mtk-mipi-dsi-mt8183.c`: txdiv/txdiv0 by + rate (956.55 Mbps → txdiv0=2, VCO≈3.83 GHz), SDM_PWR_ON → ISO_EN + clear → pcw = rate*txdiv<<24/26MHz → POSDIV → PLL_EN; lane_con + BG_CORE_EN/BG_LPF_EN sequence; lanes switched off until the host + enables them; CK_CKMODE_EN set. Data rate passed via + generic_phy_set_mode(PHY_MODE_MIPI_DPHY, bps). The Linux efuse lane + calibration is NOT ported — coreboot 4.14 runs uncalibrated lanes + on this device (its mtk_mipi_dphy.c programs fixed values), so + calibration is an optional refinement. drive-strength-microamp + defaults to 4600 uA (Linux default). +4. **DSI host** (`drivers/video/mtk_dsi.c`): UCLASS_DSI_HOST. + Register map from Linux `mtk_dsi.c` = coreboot `dsi_common.h` + (identical offsets). D-PHY timing formulas ported from both (same + math). Video timing (VSA/VBP/VFP/VACT, HSA/HBP/HFP word counts, + PSCTRL custom header 0xb<<26, SIZE_CON, D-PHY turnaround taken out + of HFP/HBP) ported from coreboot `dsi.c` — the code proven on this + device. Command FIFO programming (short/long packets, BTA for + reads, HSTX bit for non-LPM) from Linux mtk_dsi_cmdq(). Data rate = + pixelclock*bpp/lanes (no mipi ratio; coreboot uses 100/100, + mainline Linux dropped the ratio entirely). Flow: clocks + (mmsys gates via clk uclass) → phy set_mode/init/power_on → DSI + reset (FORCE_COMMIT USE_MMSYS|ALWAYS + CON_CTRL pulse) → phy + timing → rxtx → 1 ms → DPHY reset pulse → HS clk off → video + timing → HS clk on; [panel init commands in command mode]; enable() + → video mode + start. NOTE: coreboot never sets DSI_EN (CON_CTRL + bit 1) and works on this hardware; followed coreboot. +5. **Panel** (`drivers/video/panel_boe_tv101wum.c`): UCLASS_PANEL + for "boe,tv101wum-nl6". Timing = Linux boe_tv101wum_nl6_default_mode + (159.425 MHz, 100/40/24 / 10/14/4). Init command stream ported + VERBATIM from coreboot 4.14 `panel_params/panel-BOE_TV101WUM_NL6.c` + (packed {cmd,len,data…} stream, ends with 150 ms delay; NO explicit + sleep-out/display-on — the panel comes up in video mode, as proven + by every coreboot boot on this device). Power order from Linux + boe_panel_prepare(): avdd/avee/pp1800 (fixed regulators, GPIO + enables now real through the new GPIO driver), 10 ms, reset pulse + high 1.5 ms / low 8 ms, init DCS, then backlight phandle + (pwm-backlight) if compiled in. On this board the DCS stream is the + coreboot one, so no 0x11/0x29 are sent. +6. **Display driver** (`drivers/video/mt8183_display.c`): UCLASS_VIDEO + on the ovl0 node. Stage machine [dsi] phase prints: clocks (CG_CON0 + ALL + CG_CON1 DSI0 clears + SMI_LARB0 +0x380 = 0) → panel (uclass + probe) → dsi-init → dsi-enable → pipeline (MMSYS routing/mutex, + OVL ROI/bgclr, RDMA0 size+FIFO 5 KiB, COLOR0 bypass, PQ relay, + OVL L0 con/src_size/pitch/addr, engines on, backlight GPIOs). All + routing bits/mutex/registers from coreboot ddp.h/ddp_common.h. + **Naming trap**: coreboot's `DISP_OVL1_BASE` (0x14009000) is the + upstream DT's `ovl_2l0` (OVL0_2L) — OVL0_2L is the second engine of + the main path, which is why coreboot programs ROI on "disp_ovl[0]" + and "[1]". Framebuffer at 0xFC000000 (above the 2 GiB DTB DRAM + window; same region class as the firmware scanout at 0xFD536000), + FDT memreserve + LMB (commit 1ae771d9f90 pattern). Fallback: any + failing stage → mt8183_disp_setup_handoff() (the old revival path, + now shared code in mt8183_disp.c) with a log line naming the stage. +7. **Kconfig**: choice VIDEO_MT8183_SCANOUT (old behavior) vs + VIDEO_MT8183_DISPLAY (default; selects VIDEO_MTK_DSI, + PHY_MTK_MIPI_TX, PANEL_BOE_TV101WUM; needs VIDEO_MIPI_DSI, PANEL, + PHY). Defconfig adds MT8183_GPIO, POWER, DM_REGULATOR, + DM_REGULATOR_FIXED (POWER was explicitly off in the krane + defconfig; without it the fixed regulators cannot bind). + +### Known gaps / decisions + +- SCPSYS display power domain: no MT8183 power-domain driver in-tree + (mtk-power-domain.c has no mt8183 compatible). The bring-up relies + on the display MTCMOS being on (true on every path that reaches + U-Boot here). Documented in the driver. +- Backlight is still the two GPIOs (DISP_PWM 43 + EN_LCD_BL 176) + driven as board glue, not the pwm-backlight node: driving a real PWM + on pin 43 would need the MT8183 pinctrl mux (no pinctrl driver + in-tree), and the backlight node's power-supply chain (reg_vsys ← + mt6358) has no PMIC driver. The pwm-backlight phandle is wired and + used when BACKLIGHT_PWM is available. +- The panel node in the upstream DT has no pp3300-supply; Linux enables + a dummy there. Only avdd/avee/pp1800 are handled. +- OVL layer input format = coreboot's RGBA8888 value for the 32-bit + XRGB surface (proven on device); naming is MTK-internal. + +### Payload + +`krane-uboot-payload.bin` sha256 +`d29192c06266663f6b2bb2fa683a8acdb180a9d9049358d1c7fb6f17b28bb39c` +(the hash differs between rebuilds — U-Boot embeds a build timestamp; +verify via vbutil instead), vbutil body verification succeeded, `_start == 0x4C001000 == +__image_copy_start` verified. Flash recipe unchanged (U-BOOT.md). +Serial now shows `[dsi] phase 0/9/F` lines describing which path ran. + +## Round 39 — cold bring-up: stream dies at DSI re-init and never returns + +User observations, three flashes: + +- Flash A (initial): sub-ms white band top (portrait), then black lit. +- Flash B (reset polarity fixed: pulse ends released, of_to_plat + releases reset at panel probe): sub-ms wide dark-gray artifact while + backlight ramps, then black, backlight on, NO reset (PANIC_HANG=y + works — no abort → no magenta). +- Flash C (PHY analog → coreboot LANE_CON 0x3fff0180/0x00c0, timing → + coreboot HSA24/HBP40/VSA4/VBP14): "blinking band then black". + +Decoded so far: + +- The brief artifact = the firmware scanout still running while we + paint/mid-fill; the stream then dies for good. +- Black + backlight + no magenta = U-Boot console runs blind: bring-up + "succeeds", banner drawn into the new 0xFC000000 fb, but the DSI link + never re-transmits after our re-init. The failure is inside the + DSI/PHY re-init itself (link dead), not an abort. +- The DSI re-init kills the firmware stream the moment we stop/start + the host (mtk_dsi_reset → command mode → video restart), so after + that point ALL panel-side diagnostics are invisible: our bands paint + into the old fb, which nothing scans anymore. Instrumentation blind + spot: any post-first-DSI-touch failure looks like "black, lit". + +Audit findings during Round 39 (both fixed in flash C): + +1. My PHY used Linux-style analog init (per-lane RTCODE + HSTX LDO + ref). Linux's per-lane RTCODE regs are written from efuse + calibration data; uncalibrated Linux path != coreboot path. + Replaced with coreboot's proven LANE_CON sequence (embeds bandgap + + lane impedance defaults). drive-strength property dropped. +2. Sync/back-porch split: Linux HSA40/HBP24/VSA14/VBP4 vs coreboot + HSA24/HBP40/VSA4/VBP14 (same totals). Sync-pulse video mode is + sensitive to this split; aligned to coreboot. + +Open questions for flash D (not yet ruled out): + +- mm_sel (TOP mux 0x40[10:8]) must be set to a source ≥ 956.55 Mbps × + lanes/4 ... actually DSI0_IF digital clock comes from mm_sel; the + DT "hs" clock is mipi_tx0 PLL (a clock OUTPUT of the PHY). We never + program the mm_sel mux parent — we rely on firmware's setting. + coreboot never touches muxes either (relies on defaults), so + probably fine. +- CG_CON1 bit 7 (DISP_26M / CLK_MM_26M): coreboot does NOT clear it + (CG_CON1_DISP_DSI0 = bits 0|1 only); we match. +- MIPID0_26M: an apmixedsys 26M gate feed consumed ONLY by mipi_tx0. + The U-Boot clk driver has NO MIPID0 26M gate (apmixed_plls only). + On this firmware generation it is on at handoff. NOT a suspect for + the dead link (panel/DSI stays alive until WE touch it). + +## Round 40 — cold bring-up WORKING, diagnostics removed, series cleaned + +Final boot log on device: "[dsi] phase 9: full bring-up done", U-Boot +console on panel (landscape, rot=3 from DT), kernel boots after +bootefi bootmgr with logs visible. Serial console via Suzy-Q works. + +Root causes found this round (in order): + +1. **Panel driver NULL deref (the big one)**: boe_panel_send_init_ + sequence() reads plat->device, but nothing ever set mplat->device. + EVERY previous build aborted (PANIC_HANG) at the panel init stage, + right after the reset pulse. Fixed: mt8183_display.c publishes + mplat->device = &priv->device in STAGE_PANEL. +2. **CONFIG_BAUDRATE=921600**: payload reprogrammed the console to + 921600 (8x) while BootROM/terminal run 115200. minicom -b does not + reliably switch speeds; stty/picocom do. Fixed to 115200. +3. **mtk_serial HS0 sample regs**: _mtk_serial_setbrg wrote + sample_count=0/sample_point=0xffffffff in low speed mode; coreboot + leaves them untouched. Fixed upstreamable (serial: mtk commit). +4. **bpix line eaten by diag edits**: a temporary-diag edit removed + uc_priv->bpix = VIDEO_BPP32 from mt8183_display_bringup(); with + bpix=0 the console glyph path returns -ENOSYS ("Error: 1 bit/pixel + mode, but BMP has 256"), no text renders, video_clear mis-fills. + Restored. +5. DSI re-init kills the firmware scanout mid-boot: all panel-side + diagnostics after the engine reset are invisible. Workaround during + bring-up: minimal-touch DSI path (no engine/DPHY resets, no PLL + reprogram; firmware state + panel init + video restart). The final + cold path resets everything and works (panel reset + full init). + +Diagnostic techniques that worked: color bands into the live firmware +scanout (visible only until the DSI reset), backlight blink post-codes +(panel-independent), Suzy-Q serial (decisive). + +Final series on krane-updates (base 1ae771d9f90, checkpatch 0 errors): +- fab1110ff35 clk: mt8183 MMSYS display gates +- 7de03c4fa39 gpio: mt8183 GPIO driver +- 58d57350b43 phy: MT8183 MIPI TX D-PHY +- 0c6a4de7ce2 video: MT8183 MIPI DSI host +- c91f8f93d5a video: BOE TV101WUM-NL6 panel +- eaab63c84cb video: mt8183 display pipeline bring-up +- 36c0a9919eb krane: select full display pipeline bring-up +- 68b992ad2c2 serial: mtk sample-registers fix +- 357608ac481 arm: mediatek: krane: 115200 serial console + +Known issues / follow-ups: +- infra_clks array order vs legacy header IDs mismatch after index 51 + (pre-existing upstream): needs id_offs_map; bites CLK_INFRA_USB, + AP_MSDC0, I2C4/6/7/8 gate addressing. +- efi_add_memory_map rejects 0xfc000000 (above DTB DRAM window): + "reserving fdt memory region failed ... -22"; -17 for payload region + is benign. Matters for kernel GOP handoff quality. +- Kernel (pmOS) takes over the panel via fbcon/GOP; panel node is + status=disabled in the kernel DT, so no kernel DSI driver probe. +- Kernel "stuck at Running /init as init process" on serial: pmOS-side + init issue, not display. +- MAINTAINERS warnings from checkpatch: add entries when submitting. + +## Round 41 — [ROOT CAUSE] serial stall = nonexistent console=ttyMT0; fixed to ttyS0,115200 + +Symptom recap: Ubuntu kernel (7.0.0-30-generic, grub/U-Boot path) printed +early logs on Suzy-Q serial until a fixed point mid-log-line, then silence. +Looked like a hang; was a console handoff, not a hang. + +Root cause: `console=ttyMT0,921600` in GRUB_CMDLINE_LINUX_DEFAULT names a +device that does not exist on this kernel stack. Mainline 8250_mtk +(CONFIG_SERIAL_8250_MT6577) registers its ports on the universal 8250 +driver, device/console name "ttyS" — NOT "ttyMT" (ttyMT0 is the MTK +*vendor* driver, drivers/tty/serial/mtk-serial.c, not in mainline). +Ground truth from this same machine's pmOS kernel (6.12.87-mt81, same +driver, same uart node): + 11002000.serial: ttyS0 at MMIO 0x11002000 ... is a ST16650V2 + printk: legacy console [ttyS0] enabled + +Failure sequence on the Ubuntu boot: +1. earlycon=uart8250,mmio32,0x11002000 prints at firmware rate (115200) — + early logs readable. +2. console_init(): vt console enables (console=tty0 matched, CON_CONSDEV); + printk unregisters all boot consoles → earlycon dies, cutting output + mid-line (the "stall point", deterministic). +3. The pending ttyMT0 entry is retried at every later console + registration and never matches (univ8250_console.match only accepts + console=uart8250,... earlycon-style entries). No serial console is + ever enabled → silence for the rest of boot. Kernel keeps running on + tty0 only. + +Doc's open question answered: `console=ttyS0` WITHOUT a baud argument is +NOT firmware-rate. serial8250_console_setup (8250_port.c) defaults +`int baud = 9600` when options are absent and probing is off; only the +console=uart8250,mmio32,... match path probes the hardware divisor +(probe_baud) to keep the firmware rate. So the fix must state 115200 +explicitly. (mtk8250_set_termios handles 115200 fine: uartclk 26 MHz, +quot=14, highspeed register 0.) + +Fix applied (on this rootfs, U-Boot/grub boot path only — the running +depthcharge boot is unaffected): +- /etc/default/grub: GRUB_CMDLINE_LINUX_DEFAULT changed from + "console=tty0 console=ttyMT0,921600 earlycon=uart8250,mmio32,0x11002000" + to "console=tty0 console=ttyS0,115200 earlycon=uart8250,mmio32,0x11002000" +- sudo grub-mkconfig -o /boot/grub/grub.cfg (backup: + /etc/default/grub.bak-round41). Verified regenerated grub.cfg carries + console=ttyS0,115200 and the devicetree /boot/dtb line (10_linux keeps + it as long as /boot/dtb exists). + +Boot mechanics after fix: earlycon (115200) → dies when vt enables at +console_init() → ttyS0 console matched/enabled at console_initcall +(generic "ttyS" name match + .setup → uart_set_options 115200 on the +legacy placeholder port, harmless on arm64; Rockchip/rk3399 +console=ttyS2,1500000 uses the same path) → real port registers via +mtk8250_probe ~0.7 s later, console follows serial8250_ports[0] +automatically; hardware stays at U-Boot's 115200. Small window (~0.7 s) +between console_init and 8250_mtk probe where serial messages are lost +to the placeholder port; add keep_bootcon if that gap matters (duplicates +output; also useful as a diagnostic: with keep_bootcon the whole log +stays on serial regardless of handoff). + +Verification checklist (user, needs Ctrl+U in depthcharge): +1. Boot to grub → Ubuntu with terminal at 115200 8N1. +2. Expect: early logs (earlycon), brief gap, then full kernel log through + userspace at 115200; /dev/console = ttyS0 (last console= wins). +3. If silence still occurs: add keep_bootcon to the cmdline and compare — + if output continues, it is purely a console handoff issue; if not, + capture the last ~50 lines and triage initcalls (initcall_debug). +4. Panel check (separate bug, believed fixed): frozen U-Boot console + should show no noise blocks; kernel output on tty0 via the EFI + framebuffer may appear on the panel once vt/simpledrm come up — that + is expected, not a regression. + +No U-Boot-side change needed; no payload reflash required. krane-updates +untouched. + +## Round 42 — stall survives console fix: real hang; debug cmdline staged + +Round 41 fix (ttyS0,115200) did NOT change the symptom: output still cuts +at the same point (~2.3s, mid-line, right after the "evm: security.evm +attributes" message). Conclusion: the ttyS0 console WAS working from +~0.8s (mtk8250 probe registers port line 0, which is the same struct the +console attached to, so the console follows the real port automatically; +output between 0.8s and 2.3s already went through the working console at +115200). The stall is a genuine hang, not a console handoff artifact. + +Cut-point analysis: "evm: security.evm attributes reinitialized" is a +late_initcall (evm_init_xattrs). What runs next, in order: +1. remaining late(_sync/_rosync) initcalls, +2. "Freeing unused kernel memory", free_initmem, +3. smp_init() — secondary CPU bring-up via PSCI (BL31) — AFTER all + initcalls, immediately before "Run /init", +4. "Run /init as init process". +The mid-line cut is consistent with either a late-initcall hang or a +hang inside smp_init/PSCI cpu_on (nbcon printk kthread can be preempted +mid-line by a hard hang). U-Boot-handoff state is a candidate for both +(U-Boot payload leaves devices in non-reset state; secondary CPUs were +parked by BL31 either way). + +Debug boot staged (2026-09-02): +- /etc/default/grub GRUB_CMDLINE_LINUX_DEFAULT now: + "console=tty0 console=ttyS0,115200 earlycon=uart8250,mmio32,0x11002000 + keep_bootcon ignore_loglevel initcall_debug maxcpus=1" + (backup: /etc/default/grub.bak-round42; grub.cfg regenerated). +- keep_bootcon: earlycon survives the whole boot -> any output loss after + this point is hardware/hang, never handoff. +- initcall_debug + ignore_loglevel: last "calling " without a matching + "initcall ... returned" names the hang. +- maxcpus=1: tests the PSCI/secondary-CPU bring-up path. + +User checklist (needs Ctrl+U, terminal 115200 8N1, CAPTURE TO FILE): +1. Boot grub -> Ubuntu. WARNING: boot is now slow and verbose (initcall + trace over 115200 serial, ~30-60s extra; earlycon + ttyS0 duplicate + every line — expected, same UART). +2. Capture the FULL serial log to a file (e.g. picocom -b 115200 + /dev/ttyUSB0 | tee boot42.log) from before U-Boot output starts. +3. Report the last ~50 lines. Key reads: + - Last "calling ..." with no "returned" -> hung initcall (name it). + - Last line = "smp: Bringing up secondary CPUs ..." -> PSCI/secondary + CPU hang; next test drops maxcpus=1 and we go after the U-Boot + handoff / TF-A PSCI state (compare with depthcharge boot of the + same kernel; MT8183 is 4xA73+4xA53, all 8 boot fine via depthcharge + on the pmOS kernel with the same BL31). + - Output now survives past 2.3s to login -> the single-core change + fixed it; then bisect smp/hotplug vs initcall. +4. If it STILL cuts mid-line at the same spot with earlycon kept alive: + hang is in whatever the last complete "calling" line names, or in + free_initmem/smp_init if no initcall_debug lines trail. A hard hang + with a fully drained console that cuts mid-line would point at the + printing path itself dying with the machine (clock/powerdomain kill + during the print) — then we check whether the system is actually + alive (add a delayed "reboot" via kernel cmdline? no — check via + serial ping input: type a key; if the tty echoes, userspace is alive + and only output died). + +No U-Boot changes made. krane-updates untouched. + +## Round 43 — debug boot #2 result "nothing changed"; aliveness/panic probes staged + +User reports debug boot (keep_bootcon ignore_loglevel initcall_debug +maxcpus=1) = "nothing changed", still cut mid-line at ~2.3s. Suspicion: +initcall_debug would visibly flood the log from ~0.5s — an IDENTICAL log +suggests the cmdline may not have been applied. Need the "Kernel command +line:" line from the user's capture (kernel prints it early under +ignore_loglevel) plus the last ~100 lines. + +New cmdline (added this round; backup /etc/default/grub.bak-round43): +... keep_bootcon ignore_loglevel initcall_debug maxcpus=1 panic=10 +softlockup_panic=1 hung_task_panic=1 +Rationale: CONFIG_LOCKUP_DETECTOR / CONFIG_DETECT_HUNG_TASK are on; +hung task default timeout 120s. If the kernel is alive-but-stuck, a +panic + 10s reboot follows within ~2min; the reboot brings back the +U-Boot banner on serial — visible aliveness proof even if the UART dies +mid-boot. Nothing after many minutes = hard hang or dead UART path. +(Type a key on the terminal: tty echo = userspace alive.) + +ramoops dead end (for now): CONFIG_PSTORE_RAM=m (module, loads too late +to catch an early panic) and the Ubuntu /boot/dtb has no ramoops node; +pmOS kernel's own DT reserves 1 MiB ramoops at 0xffedb000. Could add a +ramoops node to /boot/dtb + load ramoops.ko from initramfs as a later +panic-capture path. + +Candidate explanations for an identical cut across 3 cmdlines: +A. cmdline never applied (verify via "Kernel command line:" in capture). +B. hard hang independent of cmdline content at a fixed early point + (U-Boot handoff state: xHCI/eMMC/DSI left active; or a driver probing + a device in non-reset state — initcall_debug will name it). +C. system alive, UART dies at a fixed point (clock/mux/pinctrl) — + indistinguishable from hang on serial alone; panic probes + key-echo + test split this. + +## Round 44 — fb-log.txt analyzed: boot never had the new cmdline; true root cause + +fb-log.txt (user capture, mtime 2026-09-02 15:05) shows the OLD cmdline +boot: 0 "calling" initcall_debug lines (would be thousands), no +"printk: legacy console [ttyS0] enabled" at the 1.787s ttyS0 port +registration, no earlycon disable message. The Kernel command line +printk itself is truncated mid-line ("root=0") — the capture pipeline +drops bytes (other lines spliced: "[0x410fd034]28f3628b...", +"0x...bbc00000ce(s) found"). So all Round 42/43 debug params were never +exercised; and the Round 41 fix has plausibly never been tested either. + +The log instead pins the REAL stall mechanism: +- 1.787s: 8250_mtk probes, registers ttyS0. Old cmdline has console= only + tty0+ttyMT0 → port is NOT a console → uart_configure_port() powers it + OFF (serial_core: "power down all ports by default, except the console", + uart_change_pm(UART_PM_STATE_OFF)) → 8250_mtk runtime suspend gates the + UART clock. +- earlycon keeps printing raw MMIO into a now-unpowered UART → output + dies silently at a fixed point (~2.38s, probe+autosuspend delta), + mid-line. Kernel continues fine on tty0. +This explains: identical cut across all previous boots, why it starts +exactly at 8250_mtk probe + delay, and why pmOS (no console= → all +consoles default-enabled → ttyS0 becomes console → port stays powered) +never stalls. The Round 41 fix (console=ttyS0,115200) attaches the +console at probe → port stays powered → serial should survive. It just +has never actually been booted. + +Next boot protocol (Round 45): +1. Lossless capture, no terminal in the path: + stty -F /dev/ttyACMx 115200 raw -echo + cat /dev/ttyACMx > boot45.log + (Ctrl+C after). Verify: grep -c "calling " boot45.log (expect + thousands); grep "console \[ttyS0\] enabled" (expect present right + after the 11002000.serial line). +2. If the cmdline STILL doesn't apply (no "calling" lines): grub is + serving a stale config — at the grub menu press "e" on Ubuntu and + boot the edited entry (Ctrl-X), or move the params directly into + /boot/efi/EFI/BOOT/grub.cfg. +3. If params verified and output still dies: check key-echo + panic + reboot probes (Round 43) — then it is genuinely the UART path/hang, + not console power-off. + +## Round 45 — [ROOT CAUSE #2] standalone grub image with hardcoded cmdline; grub-install redone + +picocom capture (fb-log.txt, 15:23) finally delivered a CLEAN "Kernel +command line:" line: it read "console=ttyMT0,921600" — the ORIGINAL +cmdline, no Round 41-43 params ever reached the kernel in any boot. + +Root cause of the delivery failure: /boot/efi/EFI/BOOT/BOOTAA64.EFI was +a grub-mkstandalone image (905 KB, built Aug 31) with a memdisk-embedded +grub.cfg containing hardcoded menuentries ("linux (hd0,gpt3)/boot/ +vmlinuz-7.0.0-30-generic ... console=ttyMT0,921600 earlycon=..."). It +never read the ESP stub nor /boot/grub/grub.cfg — every grub-mkconfig +since was a no-op. The 209-byte ESP stub existed but was dead code +(standalone image prefix = (memdisk)/boot/grub). This also explains the +BOOT_IMAGE=(hd0,gpt3)/... form in the kernel log (matches the embedded +entry verbatim). + +Fix: sudo grub-install --target=arm64-efi --efi-directory=/boot/efi +--boot-directory=/boot --bootloader-id=BOOT --no-nvram (grub 2.14), +then cp grubaa64.efi over BOOTAA64.EFI (fallback path). New image: +2.9 MB monolithic, zero embedded cmdline occurrences, plus grub-install +wrote a fresh EFI/BOOT/grub.cfg stub (search.fs_uuid e362f850 -> +configfile /boot/grub/grub.cfg). Chain now: U-Boot bootmgr -> +BOOTAA64.EFI -> /boot/grub/grub.cfg (ext4) -> Ubuntu entry with +console=ttyS0,115200 keep_bootcon ignore_loglevel initcall_debug +maxcpus=1 panic=10 softlockup_panic=1 hung_task_panic=1. +Backup of the standalone image + old stub: +/boot/efi/EFI/BOOT-standalone-bak45 (restore by copying back if ever +needed). + +Round 44's serial-path analysis stands as the expected outcome: with +console=ttyS0,115200 the port is a console at 8250_mtk probe time, so +uart_configure_port keeps it powered and the "power down non-console +ports" path that killed earlycon at ~2.4s never runs. + +Round 46 test protocol (user, picocom OK — capture was lossless): +1. Ctrl+U -> U-Boot -> grub -> Ubuntu. +2. First marker: "Kernel command line:" line must contain console=ttyS0, + 115200 keep_bootcon initcall_debug. +3. Thousands of "calling ..." lines; "printk: legacy console [ttyS0] + enabled" right after the 11002000.serial ttyS0 line. +4. Boot will be slow/verbose (115200 flood, earlycon+ttyS0 duplicate + lines). If it reaches login: serial console fixed; then trim cmdline + back (drop debug params) and re-verify a clean boot. +5. If output still dies: key-echo test + wait for panic-reboot probes + (~2 min, hung_task 120s + panic=10). + +## Round 46 — serial stall FIXED (kernel reaches initrd); initrd missing mmc devices + +Round 45 grub-install fixed the delivery: kernel boots past the old +2.4s stall all the way to the dracut initrd. Serial stall ROOT CAUSE +confirmed as the console power-off path (Round 44): ttyS0 console +attached at 8250_mtk probe keeps the port powered. + +Cmdline trimmed per user request — now: + console=tty0 console=ttyS0,115200 earlycon=uart8250,mmio32,0x11002000 +(backup /etc/default/grub.bak-round46). earlycon kept temporarily while +the initrd issue is debugged; drop it at the end. + +NEW ISSUE: dracut initrd cannot find root by UUID; /dev has no mmc*. +Static analysis says the initrd is complete: +- usr/lib/modules/.../drivers/mmc/host/mtk-sd.ko.zst PRESENT (note: file + is mtk-sd.ko, module name mtk_sd — earlier grep with "mtk_sd" missed it) +- mmc_block, cqhci, mmc_hsq present; modules.alias has + of:N*T*Cmediatek,mt8183-mmc -> mtk_sd; modules.dep lists deps +- PMIC chain present: mtk-pmic-wrap, mt6397 (MFD), mt6358-regulator, + mt6397-regulator, rtc-mt6397 +- pinctrl-mt8183 / clk-mt8183 / infracfg are built-in (=y) +- vermagic matches kernel image (same Jul 31 build), module signed +- /boot/dtb (custom krane-fb-stub DTB) is byte-identical in structure to + upstream /boot/efi/mt8183-kukui-krane-sku176.dtb (full-file diff + EMPTY): mmc0 @11230000 okay, compatible mediatek,mt8183-mmc, clocks + topckgen+infracfg phandles valid +So the failure is runtime: either mtk_sd never got loaded by udev, or +its probe fails/defers. Diagnostics for the dracut emergency shell: + cat /proc/modules | grep -Ei 'mtk|mmc|pmic' + modprobe mtk_sd && ls /dev/mmc* + dmesg | grep -iE 'mtk-sd|msdc|mmc|pmic|regulator' + ls /sys/bus/platform/devices | grep mmc +If modprobe succeeds and /dev/mmcblk0 appears: just "exit" — dracut +resumes, mounts root, boot completes; fetch dmesg/journal from the +booted system afterwards to pin the root cause (probe defer vs error). +Round 46 backup: /etc/default/grub.bak-round46. + +## Round 47 — mt6358_regulator was the missing initrd load; native display path completed + plymouth + +User confirmed: modprobe mt6358_regulator in the dracut shell unblocked +the initrd (regulators registered, mmc0 deferred probe resolved, root +mounted). Boot then proceeded to systemd (Ubuntu 26.04.1 userspace on +this rootfs) and stopped after "Starting wpa_supplicant.service" — open +issue, suspected mt7663s/mt76 SDIO path (initrd dmesg showed msdc cmd52 +errors on mmc1). Discriminator: press Enter on serial — login prompt = +system alive, wpa-supplier-only stuck. + +Panel goes blank when the kernel takes over ("graphical console +disappear"): NOT a reason to blacklist the display stack (rejected — +native display is the goal). Real cause found: the initrd contained +mediatek-drm/mtk_mmsys/mtk_mutex/DSI-phy (which reset the DSI link U-Boot +left running -> panel dark) but NOT the panel/backlight/PWM modules, so +nothing could re-light it. The DTB (/boot/dtb) has the full native path +ENABLED: panel@0 boe,tv101wum-nl6 (avdd/avee/pp1800 fixed GPIO +regulators), pwm-backlight on SoC pwm@11005000. The old "DSI/panel +disabled in distro DTB" note does not apply to this DTB. + +Fixes applied: +- /etc/dracut.conf.d/50-display.conf: + add_drivers+=" panel-boe-tv101wum-nl6 pwm-mediatek pwm_bl " + force_load="mt6358_regulator" + (force_load because udev failed to load the already-present module at + runtime in the previous initrd — root cause unknown, worked around.) +- plymouth + plymouth-theme-spinner + plymouth-label installed via apt; + dracut now embeds plymouthd (50plymouth) — initrd rebuilt + (53 MB, 16:43). NOTE: apt's dracut trigger also runs update-initramfs, + so future kernel/apt operations keep the config. +- /etc/default/grub: cmdline now + console=tty0 console=ttyS0,115200 earlycon=uart8250,mmio32,0x11002000 splash + ("splash" only, no "quiet" — serial stays verbose for the wpa debug). + grub.cfg regenerated. + +Expected next boot: panel lights during initrd (native DSI panel takes +over from U-Boot firmware scanout — brief flicker), plymouth splash on +the panel, serial stays verbose. If panel still dark: capture +dmesg | grep -iE 'panel|dsi|drm|backlight' from serial login and check +panel bind/defer. + +## Round 48 — regulators still missing with force_load; deterministic pre-udev hook + +fb-log.txt (16:52 boot): plymouth-start ran, but the deferred tree was +back — mmc0 (ldo_vio18), usb (ldo_vusb), gpu (buck_vgpu), i2c (vcn18/ +vcamio), AND the whole MT8183 power-controller: + mtk-power-controller: power-domain@2 failed to get power supply + (domain-supply = MT6358 buck, coupled vproc pair) + -> iommu, all larbs, ovl/rdma/dsi/mutex/aal/ccorr/color/gamma, pwm, + backlight_lcd0 all defer on "supplier 10006000.syscon:power- + controller not ready" +So the ENTIRE deferred forest (eMMC + display + iommu + backlight) has a +single root: MT6358 regulators not registering. force_load="mt6358_regulator" +did NOT load it (no evidence of any generated load mechanism in the +initrd). + +Fix (deterministic): dracut pre-udev hook. Gotchas found: +- dracut 110-11 does NOT copy host /usr/lib/dracut/hooks into the image. +- Runtime hookdir = /var/lib/dracut/hooks (dracut-lib.sh:367); stage dir + is pre-udev (dash), per source_hook pre-udev in usr/bin/dracut-pre-udev. +- Placed /var/lib/dracut/hooks/pre-udev/01-mtk-pmic.sh (host) with: + modprobe mtk_pmic_wrap; modprobe mt6397; modprobe mt6358_regulator + and /etc/dracut.conf.d/50-display.conf: install_items+=" " + (install_items preserves the path; survives apt-triggered + update-initramfs). +Verified in rebuilt initrd: var/lib/dracut/hooks/pre-udev/01-mtk-pmic.sh +present, executable. Manual escape if a boot still lands in dracut +shell: modprobe mtk_pmic_wrap mt6397 mt6358_regulator, then exit. +Expected: regulators register ~4s into initrd; mmc0, power-controller, +display/iommu/backlight all unblock; panel lights; plymouth splash. +Still open: wpa_supplicant hang (previous boot; suspect mt7663s/mt76 +SDIO after msdc cmd52 errors on mmc1). Check with Enter-on-serial for +login prompt, then journalctl. + +## Round 49 — [ROOT CAUSE] soft lockup = live scanout faulting through re-enabled M4U; U-Boot quiesce committed + +fb-log.txt (17:13): eMMC fixed (regulators registered via pre-udev hook), +boot went further than ever: initrd pivot, real-root systemd, +wpa_supplicant [OK] (previous hang gone). Two issues surfaced: + +1. UBSAN shift-out-of-bounds mt6358-regulator.c:384 — ffs(0)-1 = -1 in + mt6358_get_buck_voltage_sel. Root: mt6358_volt_fixed_ops routes + get_voltage_sel through the buck helper, but the fixed LDOs + (vio18, vrf12, ...) never initialize da_vsel_reg/da_vsel_mask + (v7.0 mainline has the same code). Non-fatal: read-only path, selector + 0 == nominal for these LDOs. Upstreamable fix: use + regulator_get_voltage_sel_regmap for fixed ops (they have valid + vsel_reg/vsel_mask). NOT the lockup cause. +2. mtk-iommu fault storm: reads at iova 0xbe000xxx (the U-Boot + framebuffer at 0xBE000000!) from master larb0/port0 — the display + engine kept scanning out the U-Boot console while the kernel's M4U + enabled translation; the region has no IOMMU mapping. Interrupt storm + starved timer handling: CPU#1 rcu_exp_gp_kthr soft lockups (26/52/89s), + rcu_preempt GP kthread starved on CPU4 (first A73, "timer wakeup + didn't happen"). Boot wedged around 40s, right at + NetworkManager/ModemManager startup. + +Fix (commit d3d3502de0a on krane-new-panel-driver, checkpatch 0/0): +"video: mt8183: quiesce the display pipeline at ExitBootServices". +- drivers/video/mt8183_display.c: board_quiesce_devices() — stops the DSI + video stream + powers the D-PHY down (mtk_dsi_disable), stops the + pipeline engines (OVL0/OVL0_2L/RDMA0/COLOR/PQ blocks/mutex), turns the + backlight off and gates the MMSYS display clock domains. The kernel + display driver does a full cold bring-up (Round 40), so nothing of the + handoff state needs preserving. +- drivers/video/mt8183_disp.c: mt8183_disp_disable_backlight() (inverse + of enable; DOUT clear registers at +8 in the GPIO dout block). +- drivers/video/mt8183_disp.h: DOUT_CLEAR macro + prototypes. +U-Boot rebuilt; payload rebuilt and vbutil_kernel-verified: +krane-uboot-payload.bin sha256 771a0fc6dfda12af9d6779b7637787dc5177b6db3e3f7cba7443be9911a09bb5. +PENDING: dd to /dev/mmcblk0p1 (user confirmation per protocol), then +boot via Ctrl+U. + +Expected next boot: panel goes dark after the kernel's EFI stub calls +ExitBootServices (U-Boot hands over with the pipeline quiesced — no more +frozen console, no fault storm), kernel brings the panel up natively +(~5-10s), plymouth splash, full boot. The UBSAN warning remains (harmless; +module fix is a follow-up). + +Round 49 addendum: payload flashed to /dev/mmcblk0p1 (dd verified with +cmp against the source file, 860160 bytes). Ready for Ctrl+U boot test. + +## Round 50 — [ROOT CAUSE] U-Boot quiesce works; new oops = mtk_smi larb runtime-resume before iommu bind; patched modules installed + +fb-log.txt (17:54): the IOMMU fault storm + RCU soft lockup are GONE (the +U-Boot quiesce commit d3d3502de0a works). Boot got to coldplug, then: + + Internal error: Oops 0000000096000004, FAR=0x0, pc + mtk_smi_larb_config_port_gen2_general+0xf0 [mtk_smi], lr + mtk_smi_larb_resume+0xb8, via pm_runtime_get_suppliers from + mtk_drm_init (mediatek_drm module load, udev-worker PID 266). + Code bytes match mainline v7.0 drivers/memory/mtk-smi.c exactly: + `ldr x1,[x28,#144]` (= larb->mmu, offset 144) then `ldr x1,[x1]` at + +0xf0 -> NULL because larb->mmu is only set by mtk_smi_larb_bind(), + the IOMMU component bind, which ran at 8.168s — AFTER the oops at + 8.155s. mediatek_drm's probe runtime-resumes the larb through the + device link/genpd before the IOMMU binds it. Unfixed in upstream + master (checked mtk-smi.c master == v7.0). + +Also confirmed this boot: mt6358 UBSAN fires from +mt6358_regulator_probe->regulator_register->machine_constraints_voltage +(ops->get_voltage_sel on register: mt6358_get_buck_voltage_sel derefs +da_vsel_mask which MT6358_REG_FIXED never sets). Fixed LDOs DO have +valid vsel_reg/vsel_mask (MT6358_*_ANA_CON0 / GENMASK(3,0)), so the +correct ops is regulator_get_voltage_sel_regmap (as +mt6358_volt_range_ops uses for regmap reads elsewhere). + +Fix: rebuilt both modules out-of-tree against the Ubuntu headers +(/usr/src/linux-headers-7.0.0-30-generic, Module.symvers, MODVERSIONS +OK, vermagic matches, unsigned load = taint only, MODULE_SIG not +forced). Source validated against the shipped modules before patching: +rebuilt unpatched mtk-smi.ko reproduces the oops Code bytes at +0xf0. + +- mtk-smi.ko: guard in mtk_smi_larb_resume: if (!larb->mmu) return 0; + after enabling clocks (no IOMMU master attached yet -> nothing to + configure). Upstreamable: "memory: mtk-smi: skip MMU port config on + larb runtime-resume before the IOMMU binds". +- mt6358-regulator.ko: mt6358_volt_fixed_ops.get_voltage_sel -> + regulator_get_voltage_sel_regmap (line 495; vproc/vsram buck ops + untouched). + +Installed to /lib/modules/7.0.0-30-generic/kernel/drivers/{memory/ +mtk-smi.ko.zst,regulator/mt6358-regulator.ko.zst}; originals kept as +*.orig-round49 (NOTE: named round49, stamped before analysis); depmod +run. BTF skipped (no vmlinux) — same as many Ubuntu modules. + +Expected next boot: no oops; mediatek_drm probes; panel lights +natively; plymouth; boot to login. + +## Round 51 — false alarm: same oops; initrd ships stale module copies; initrd rebuilt + +fb-log.txt (next boot): IDENTICAL oops (same pc +0xf0, same Code bytes, +same UBSAN from mt6358_get_buck_voltage_sel). Cause: dracut initrd +contains its own module copies (usr/lib/modules/7.0.0-30-generic/...) +built Jul 31 — the modules we replaced under /lib/modules never load; +initrd modules are what run during coldplug (before root pivot). The +"more errors" = the oops printed twice (a second udev worker retried +the mediatek_drm finit_module and hit the same fault) + dracut +initqueue hang, collateral of the oops killing the worker handling the +mmcblk uevent chain (root node never settled). + +Fix: dracut -f rebuild — initrd now carries the patched mtk-smi.ko.zst +(114115 bytes, mine keeps DWARF that Ubuntu strips to dbgsym; loads +fine) and patched mt6358-regulator.ko.zst. Stale .orig-round49 backup +files also got copied into the initrd by dracut (harmless, never +loaded). + +Next boot expectation: no oops, no UBSAN, initqueue completes, root +mounts, mediatek_drm probes, panel lights natively. + +## Round 52 — [WEDGE] boot reaches real root; CPU2 kworker spin + multi-CPU timer death; evidence-led prep for next boot + +fb-log.txt (post-Round-51 initrd): no oops, no UBSAN, no IOMMU storm — +the module fixes hold. Boot reaches systemd, NetworkManager, +wpa_supplicant. Display: backlight comes back (pwm_bl) but screen stays +BLACK and mediatek_drm never registers an fbdev. Then CPU#2 soft lockups +(26/52/119s, kworker/2:3), RCU stalls on 4/5/7, rcu_preempt kthread +(cpu3) "timer wakeup didn't happen". Start of wedge ~13.8s. + +Key observations: +- The soft-lockup STACK DUMPS never appear on the serial console (only + the header lines) — dumps are lost somewhere in printk/console path. + Rely on ramoops next boot instead. +- The current initrd was MISSING mediatek_drm (my Round-51 rebuild + dropped it vs the Jul 31 build). So this boot loaded mediatek_drm from + the real root at ~13s — exactly when the wedge started. Both full-boot + wedges (49, 52) time-correlate with mediatek_drm activity; the boot + where its probe oopsed early (51) never wedged the CPUs. +- PSCI CPUidle EXonerated: the running pmOS kernel uses psci_idle with + the SAME WFI/cpu-sleep/cluster-sleep-0 states and the same stock ATF, + up 17min+ fine. Not the cause despite first suspicion. +- No unbounded loops found statically in mtk_crtc/mtk_dsi/mtk-mutex/ + cmdq-mailbox/cpufreq/mtk-coupler; only the DSI IRQ handler + do{}while(tmp & DSI_BUSY) (unbounded, irqs-off) — but a CPU stuck + there could never report its own soft lockup, so it is not the + reported kworker spin. U-Boot quiesce now clears DSI INTEN/INTSTA + anyway (commit c112424b983, checkpatch clean; payload sha256 + 7ed2de1e0ab3ff80cb6f461889d06bbf3d9d773f555308d02107dcb124cea92b, + flash PENDING user confirmation). +- grub2-common/grub-initrd-fallback FAILED at exactly the wedge moment + (collateral; recordfail cleared via grub-editenv). + +Prepared for the next boot (all in place): +1. cmdline: timer_migration=off (targets the timer-migration/tick + failure class matching "timer wakeup didn't happen"; nohz/timer + rework landed 6.13..7.0 while pmOS runs 6.12.87 stable with the + backported fixes), sysrq_always_enabled, softlockup_panic=1 + panic=10 (auto-evidence: wedge -> panic -> stacks -> warm reboot). +2. ramoops via DTB: /boot/dtb-krane-ramoops.dtb adds a reserved-memory + region at 0xBFF00000 (1 MiB) + ramoops node (console 512K, dmesg + 128K, pmsg 128K). grub entries use it via devicetree (10_linux picks + /boot/dtb-7.0.0-30-generic first). pmOS kernel has PSTORE=n and a + different DTB so it ignores the region. +3. initrd rebuilt: mediatek_drm + full display stack (mtk_mutex, + mtk_mmsys, mtk_smi, mtk_iommu, dsi phy, cmdq) + ramoops (force_load) + + scp.img.zst restored/added. scp remoteproc should now bind at + initrd coldplug instead of failing with -2. +4. grub experiment entries: "Ubuntu 7.0 EXP-B: cpuidle.off=1" and + "pmOS kernel via U-Boot (wedge bisect)" (pmOS kernel + pmOS DTB + + pmOS initrd from the ESP, under our U-Boot). Ladder: default entry + (timer_migration=off) -> if wedged+panicked, second boot archives + /sys/fs/pstore via systemd-pstore (enabled) to + /var/lib/systemd/pstore; read stacks from there. If still wedging + without evidence, try EXP-B, then the pmOS-under-U-Boot entry to + separate bootloader state from kernel regression. + +Black display analysis: pipeline is quiesced at ExitBootServices (by +design), simpledrm fb0 exists but nothing scans it out; mediatek_drm +did not complete bind in this boot (late load + wedge). With the display +stack back in the initrd and the wedge fixed, the kernel should bring +the panel up natively. If the wedge turns out to be INSIDE mediatek_drm +probe, the ramoops stacks will show it. + +Round 52 addendum (evidence path locked in): CONFIG_PSTORE_CONSOLE and +PSTORE_PMSG are NOT set in the Ubuntu kernel, so ramoops only produces +a dmesg-ramoops record on PANIC. That is exactly what the new cmdline +gives: softlockup_panic=1 -> full ring buffer (incl. the lockup stacks +that never reached the serial console) -> dmesg-ramoops -> panic=10 -> +warm reboot. Each boot, systemd-pstore (enabled, runs ~12.5s, before +the 13.8s wedge point) archives the previous panic to +/var/lib/systemd/pstore. Read results from there (or /sys/fs/pstore on +a boot that completes) after the test. + +## Round 53 — ramoops region collided with U-Boot's runtime data at DRAM top; moved to 0x60000000 + +fb-log.txt (early crash, 1.31s): efi_call_rts oops — "Unable to handle +kernel paging request at 0xbff29f30", x0=0xbff29ee0. 0xbff2xxxx is +U-Boot's EFI runtime services data: U-Boot relocates to the TOP of DRAM +(0xBFF00000..0xC0000000 for 2 GiB), exactly where I placed the ramoops +no-map region. The no-map carve-out removed those pages from the +kernel's linear map; the first EFI runtime call after boot (efi_rts_wq, +rtc-efi probe) dereferenced U-Boot's runtime data and the whole runtime +services path died with it. Boot never reached the wedge test. + +Fix: ramoops moved to 0x60000000 (mid-DRAM; clear of the 0x50000000 +shared-dma-pool at 0x50000000-0x52900000, the low kernel image, and the +top-of-RAM U-Boot runtime area). Both /boot/dtb-krane-ramoops.dtb and +/boot/dtb-7.0.0-30-generic rebuilt. Everything else (cmdline, initrd, +grub entries) unchanged. Lesson: never reserve anything at the top of +DRAM on this platform — that is U-Boot's relocation + EFI runtime + +variable-store area. + +Round 53 addendum: U-Boot payload 7ed2de1e0ab3ff80cb6f461889d06bbf3d9d773f555308d02107dcb124cea92b +(quiesce + DSI INTEN/INTSTA clearing, commits d3d3502de0a + c112424b983) +flashed to /dev/mmcblk0p1, cmp-verified. + +## Round 54 — no-watchdog silent lock both boots; new prime suspect: mt7663s wifi fw download deadlocking mtk-sd/eMMC I/O; hung_task_panic wired + +Two boots (default + EXP-B cpuidle.off=1) both lock at the same point: +last kernel line = sbs uevent at ~12.8/13.2s, services continue to +ModemManager start, then TOTAL silence — no softlockup, no RCU stall, +no panic. A silent (sleeping) deadlock, not a spin: cpuidle is +exonerated, and timer_migration=off turned out to be an UNKNOWN param +on 7.0 ("will be passed to user space") so it never applied anyway. + +New leading theory: NetworkManager brings wlan0 up right there -> +mt76 mt7663s firmware download over the SDIO link that shows CRC +errors from boot (msdc cmd52 host->error=0x2) -> mtk-sd driver wedges +-> eMMC I/O hangs (grub2-common/grub-initrd-fallback grubenv writes on +eMMC FAIL in every wedged boot!) -> system sleeps forever. Round 52's +spinning kworker/2:3 = mt76 fw download busy-wait; the current +silent shape = same trigger, deeper sleep. Not yet proven. + +Prepared: +- hung_task_panic=1 added to all Ubuntu entries (CONFIG_DETECT_HUNG_TASK + + HUNG_TASK_BLOCKER are on): a 120s-stuck D-state task now panics with + full stacks AND the blocker name into the ring buffer -> ramoops + dmesg-ramoops -> auto-reboot -> systemd-pstore archives it. +- EXP-C entry: module_blacklist=mt76,mt76_sdio,mt7663s, + mt7663_usb_sdio,mt76_connac_lib,mt7615_common (wifi off) + same + panic params. If EXP-C boots past the wedge point, wifi/SDIO is + the trigger. +- pmOS bisect entry fixed: the pmOS system moved to the USB drive + (sda1 kernel FIT, sda2 /boot, sda3 root); the running kernel is + 6.12.87-mt81, gzipped Image decompressed and staged as + /boot/vmlinuz-pmos-6.12.87 (PE/EFI stub verified) + initramfs + + dtb on eMMC, entry boots it via U-Boot. +- NOTE: our shell session is a chroot into eMMC p3; the real running + pmOS boots from the USB drive (pmos_root_uuid=ddc5b150 = sda3). + +## Round 55 — EXP-C panicked with FULL STACKS: the wedge is CPUs going dead to IPIs at coldplug settle + +EXP-C (wifi blacklisted) locked like the others, but this time the +softlockup detector FIRED and we finally have stacks: +- watchdog: CPU#3 soft lockup 26s, udev-worker (PID 257), stack: + smp_call_function_many_cond <- kick_all_cpus_sync <- + flush_module_icache <- load_module <- finit_module. A udev module + load broadcast an IPI and got no answer for 26s. +- panic path: "SMP: failed to stop secondary CPUs 0-2,5-7" — SIX of + eight CPUs were already unreachable when it panicked; only CPU3 + (the loader) and CPU4 responded to the stop IPI. +- Timeline: last normal log 8.69s (ccifreq deferral spam ending = + coldplug settling); CPUs died in the ~8.7-10.3s window; report at + 36.3s. The dead CPUs never softlockup-report themselves (no + watchdog ticks at all — deeper than an IRQs-off spin: no timer + interrupts / IPIs reaching them). +- wifi-blacklist did NOT prevent the wedge -> mt76/SDIO is NOT the + trigger. The wedge family across all boots = multi-CPU death at + initrd coldplug settle; the visible symptom (spin vs silent sleep) + depends on which task notices first. +- ramoops was broken all along (-22 "failed to locate DT + /reserved-memory resource"): v7.0 of_device_alloc creates MEM + resources only from `reg`; a root-level ramoops node with + memory-region never gets one. FIXED: moved the node into + /reserved-memory with compatible="ramoops" + reg (upstream exynos + pattern), record/console/pmsg sizes inside the node. Recompiled + /boot/dtb-krane-ramoops.dtb + /boot/dtb-7.0.0-30-generic. Next + panic will be archived to /var/lib/systemd/pstore by systemd-pstore. +- New EXP-D entry: maxcpus=1 (if it boots fully, the death is in the + per-CPU idle/PSCI/PM layer, not in drivers). +- Available next: pmOS-kernel-via-U-Boot bisect entry (env vs kernel + split), EXP-B cpuidle.off=1 (already shown insufficient alone). + +Interpretation candidate for the dead-CPU signature: CPUs stopped +servicing IPIs AND their own timer ticks — PSCI/ATF-level CPU state +(suspend that never returns) or clock/power gated out from under +running CPUs around sync_state/coldplug settle. No proof yet. + +## Round 56 — pmOS-6.12-via-U-Boot bisect made actually runnable without the USB drive + +Constraint discovered: the USB-C port is shared between the serial +cable and the pmOS USB drive — both cannot be attached at once, so the +pmOS rootfs (sda3) is unavailable for the bisect boot. Workaround: +the running pmOS rootfs IS reachable via /proc/1/root, so: +- Copied /proc/1/root/lib/modules/6.12.87-mt81 (18 MiB) to + /lib/modules/ on the eMMC Ubuntu root. In the pmOS kernel mtk-sd, + mtk-smi and mediatek-drm are BUILT-IN (its initramfs has only 24 + modules), so the pmOS initramfs can mount eMMC p3 with no modules; + wifi (mt7663s) etc. load from the copied tree at full coldplug. +- Rewrote the 'pmOS 6.12 kernel via U-Boot (wedge bisect)' grub entry: + /vmlinuz-pmos-6.12.87 (decompressed Image, PE/EFI stub verified) + + /initramfs-pmos-6.12.87 + /dtb-pmos-krane.dtb, all on eMMC, + with pmos_root_uuid=e362f850 (Ubuntu eMMC root) so the FULL coldplug + window runs under the 6.12 kernel + Ubuntu userspace + U-Boot + handoff. Dropped pmos_boot_uuid (this initramfs was rebuilt for the + USB layout; FAT ESP mount could stall it). Added + softlockup_panic=1 hung_task_panic=1 panic=10 so a 6.12 wedge + panics with stacks on serial. +- Interpretation: bisect boots fine -> 7.0 kernel bug. Bisect wedges + the same way -> U-Boot handoff / ATF / DTB environment issue. +- NOTE: the bisect boots UBUNTU userspace under a pmOS kernel — it is + NOT the real pmOS; do not confuse the two after boot. depmod of the + copied tree was done on pmOS originally; modules.dep present (935). + +## Round 57 — EXP-D (maxcpus=1) BOOTS FULLY: wedge requires SMP; pmOS bisect entry had /boot path bug (fixed) + +- EXP-D maxcpus=1 boots through coldplug to serial login (panel still + black — display issue is separate). Wedge does not occur with one + CPU. Combined with EXP-B (cpuidle.off=1 wedged): the death is tied + to multi-CPU bring-up/coupling, not to the idle framework itself. +- Unexplained: sudo hang at the EXP-D login prompt (no kernel output, + log ends at "[sudo: authenticate]"). Ask user to retry and WAIT + >=2-3 min: hung_task_panic should panic with the D-state stack + + blocker on serial (single CPU means nobody reports a CPU0 death, + but a sleeping task is still catchable). +- ramoops DID NOT register in the EXP-D boot (no probe message at + all; module ramoops.ko.zst IS in the initrd; v7.0 has OF match + table + reserved_mem_matches entry so the /reserved-memory node + should get a device). Unresolved — have user check + `ls /sys/fs/pstore` and `modprobe -v ramoops; dmesg|grep -i ramoops` + on the next EXP-D boot. +- pmOS-6.12-via-U-Boot entry: grub spammed file-not-found then fell + through — ROOT CAUSE: my rewritten entry used root-level paths + (/vmlinuz-pmos-...) but on eMMC the files are in /boot/. Fixed: + /boot/vmlinuz-pmos-6.12.87, /boot/initramfs-pmos-6.12.87, + /boot/dtb-pmos-krane.dtb. +- Added EXP-E maxcpus=4 (big A73 cluster only, no LITTLE cpus): + discriminates LITTLE-cluster involvement (cpufreq policy4, CCI, + cpus 4-7) from big-cluster SMP. +- Current best theory family: something in the multi-CPU bring-up/ + cluster-coupling path (cpufreq/CCI/SVS/power) kills CPUs dead to + IPIs at coldplug settle on 7.0; absent with maxcpus=1; not cpuidle + (EXP-B); not wifi (EXP-C); not ramoops region (existed in wedging + boots only since round 53, wedge predates it). + +## Round 58 — pmOS 6.12 via U-Boot DIES too (env confirmed!); EXP-E hung later; cleanup-trio suspicion + +- pmOS 6.12 via U-Boot (all 8 CPUs, own DTB, eMMC root): all 8 CPUs + boot (0.088s), eMMC enumerates (HS400 2.28s), mediatek-drm binds, + fb1 created — then SILENT at ~3.02s: last prints clk banner (2.994) + / genpd banner (3.002) / ALSA list (3.017); "Freeing unused kernel + image (initmem) memory" never printed. THE SAME KERNEL BOOTS FINE + VIA DEPTHCHARGE. => U-Boot handoff is a necessary condition. + Environment, not (only) 7.0 kernel. +- Window analysis: death sits in the late_initcall_sync tail, right + where clk_disable_unused -> genpd_poweroff_unused -> + regulator_init_complete run. regulator_init_complete silently + force-disables boot-on-but-unclaimed regulators — U-Boot (display + bring-up) leaves regulators/clocks/domains ON that depthcharge + does not; the kernel then tears down something CPUs depend on. + CPUs 1-7 die, CPU0 freezes shortly after (no free_initmem print). + Unifying with 7.0: EXP-C full-SMP "CPUs dead to IPIs" and the + ~10s module-load IPI spin = same teardown, different notice time; + maxcpus=1 survives (nothing to tear down under other CPUs); + EXP-E (maxcpus=4) PASSED the 2.49s cleanup (Freeing initmem + + Run /init seen) and hung later at NM/ModemManager (~15s, wifi not + blacklisted — mt76 fw download is back as a candidate for THAT + hang, possibly a second, separate deadlock). +- ALSO: U-Boot hands off at EL2 ("All CPU(s) started at EL2"), + depthcharge at EL1 — another handoff delta to keep in mind. +- Prepared: pmOS bisect entry now has initcall_debug + + clk_ignore_unused + pd_ignore_unused + regulator_ignore_unused. + Boot it: if it reaches login, the teardown trio is the killer and + we bisect which of the three; the initcall_debug tail pins the + exact hung function if it still dies. +- Note: EXP-D ramoops still silent — check `ls /sys/fs/pstore` + + `modprobe -v ramoops` on a working boot sometime. + +## Round 59 — ignore-params FIXED the 3s death; two separate bugs now cleanly separated + +pmOS 6.12 via U-Boot with clk_ignore_unused + pd_ignore_unused + +regulator_ignore_unused: SAILS through the 3s teardown death into +full userspace (systemd starting Ubuntu services at 15s), then hangs +at ~16.1s at NetworkManager/ModemManager start WITH wifi active — +the classic point. initcall_debug lines did not appear on serial +(KERN_DEBUG vs console loglevel mystery — unresolved, moot now). + +Bug matrix across experiments (wifi = mt76 bring-up): +- teardown death (clk/genpd/regulator cleanup under U-Boot handoff, + CPUs 1-7 killed, CPU0 freezes): pmOS 6.12 @3s pre-params. Fixed by + the three ignore params. +- Bug B (mt76 fw download deadlocks with >1 CPU): EXP-B (cpuidle.off, + wifi on) ~14s; EXP-E (maxcpus=4, wifi on) ~15s; this boot (8 CPUs, + wifi on) ~16.1s. Absent with wifi blacklisted (EXP-C reached the + OTHER bug at 9-10s) and with 1 CPU (EXP-D wifi came up fine). +- EXP-C (8 CPUs, wifi off): died 9-10s = first all-8-idle window + => deep idle/domain-sleep death on 7.0 (cpuidle.off=1 should fix). +So: Bug A' on 7.0 = deep idle (cluster/domain sleep) kills CPUs; +Bug A on 6.12 = the teardown kills CPUs (only seen on clang-built +pmOS kernel — Ubuntu gcc 7.0 passed teardown in EXP-D/E). +maxcpus=1 avoids both (no domain idle states, no other CPUs). + +New entries (this round): +- EXP-F (7.0): cpuidle.off=1 + mt76 module_blacklist, full 8 CPUs. + If it boots fully -> both bugs confirmed, working 8-CPU system. +- pmOS bisect entry: same three ignore params + cpuidle.off=1 + + mt76 blacklist. If it boots fully -> 6.12 also working via U-Boot. +Next after confirmation: live root-causing on the working system +(disable cpuidle states one by one via sysfs to find the killer +state; bisect mt76 with 8 CPUs), plus decide the real fix (DTB +always-on marks? U-Boot handoff cleanup? mt76 fix?). + +## Round 60 — second full panic nails the shape: individual CPUs die silently in hardirq context + +EXP-F reboot: same panic shape as EXP-C — udev module load spinning +in kick_all_cpus_sync (CPU#5, started ~10.3s), but this time +"SMP: failed to stop secondary CPUs 4,7": only cpus 4 and 7 were +dead; 0-3,5,6 answered. Cp 4,7 = LITTLE cluster members, but 5,6 +(same cluster) alive => NOT a cluster-wide clock/regulator kill. +Individual random CPUs go silent at ~8-11s (initrd coldplug window). + +Interpretation: CPUs stuck in HARDIRQ context (explains: no IPI +service, no timer ticks, no self softlockup report, no panic; a +spinning hardirq handler never returns so hrtimers never fire). +Trigger candidate: an IRQ handler with an unbounded wait loop +(mtk_dsi irq do{}while-DSI_BUSY, cmdq mailbox, cros-ec rpmsg/spi) +arming at coldplug under the U-Boot handoff hardware state. +Downstream effects now unified: module-load IPI spins -> panics; +mt76 fw-download work queued on a dead CPU's kworker -> the +NM/ModemManager-era hangs (EXP-E, pmOS round 58 boot); silent +freeze when no spinner reports. +maxcpus=1 survival remains consistent (no IRQ spreading). + +Prepared EXP-G: EXP-F + irqaffinity=0 -> all external IRQs on CPU0; +if a handler spins, CPU0 dies first/visibly. Ask user for sysrq +(BREAK + w/t) during any wedge: 'l' backtrace of all CPUs would show +the stuck hardirq handler directly. + +## Round 61 — irqaffinity=0 did NOT protect: 7 of 8 CPUs died (0-2,4-7); cpuidle confirmed OFF; cascade model + +EXP-G (EXP-F + irqaffinity=0): CPU#3 spun in kick_all_cpus_sync +(module load, started ~14.3s, further than before — real-root modules +loading), "SMP: failed to stop secondary CPUs 0-2,4-7": SEVEN CPUs +dead including CPU0 — but with all device IRQs pinned to CPU0 a +spinning device handler would have killed only CPU0. => +- cpuidle.off=1 IS effective ("failed to register cpuidle driver", + "CPUidle PSCI: Failed to create psci-cpuidle device") — no PSCI + suspend path exists in these boots at all. +- Simple device-IRQ-storm-as-primary is dead too (CPU0 died anyway). +Working model now: PRIMARY = CPU(s) stuck in a hardirq handler +(any CPU; can hit several — EXP-F had 4,7); CASCADE = a stop_machine +(jump-label/text patch during module probes) parks every other CPU's +stopper thread in multi_cpu_stop with IRQs masked, waiting forever +for the stuck one -> whole-machine silent death; the innocent +module-load CPU then spins in kick_all_cpus_sync and softlockups. +Consistent with 7-dead (EXP-G) and 2-dead (EXP-F) variants. + +picocom correction: C-a C-b = "set baudrate" (that was the prompt!). +Serial BREAK in picocom = C-a C-j (pulse BREAK), then the sysrq +letter (l = all-CPU backtrace, t = task dump) quickly after. + +Prepared EXP-H: EXP-F + threadirqs -> handlers run as kernel threads; +a spinning handler becomes schedulable and the softlockup/hung-task +detector NAMES it (stack + handler identity) instead of silently +killing CPUs. This is the experiment that should finally reveal the +killer function. + +## Round 62 — EXP-H (threadirqs): same crash, new victim; pseudo-NMI prepared as the stack-revealing tool + +EXP-H: CPU#2 kworker/2:2 stuck 26s in smp_call_function_single <- +rcu_barrier <- fqdir_free_fn (netns frag teardown work — another +ALL-CPU barrier wait, not the cause). "failed to stop 0-1,3-7" — 7 +CPUs dead again. threadirqs didn't change the class => the stuck +CPUs are NOT in a plain device-IRQ handler (those would have become +visible as threaded tasks). All panics share: reporting CPU waits in +an smp_call/rcu_barrier on other CPUs that never service IPIs. +Death window ~2.5-14s (coldplug storm), every full-SMP U-Boot boot, +both kernels. maxcpus=1 immune. cpuidle confirmed off. irqaffinity=0 +confirmed ineffective (CPU0 died). Display/iommu never bind on 7.0 +under U-Boot (deferred), so the DSI-IRQ-loop theory is weakened. +Ubuntu 7.0 has CONFIG_ARM64_PSEUDO_NMI=y but disabled by default +("watchdog: NMI not fully supported"). +Prepared EXP-I: + irqchip.gicv3_pseudo_nmi=1 nmi_watchdog=1 -> +hard lockup detector becomes live; CPUs stuck with IRQs masked +(multi_cpu_stop, hardirq, anything) will SELF-REPORT their stacks +via pseudo-NMI on serial. This should finally show where the dead +CPUs are. Removed softlockup_panic from EXP-I so hardlockup reports +print repeatedly instead of one soft-lockup panic cutting the dump +short. sysrq during the ~26s wedge window also works: picocom +C-a C-j (pulse BREAK) then 'l' (all-CPU backtrace). + +## Round 63 — EXP-I null test (quirk), ftrace dump-on-panic prepared + +EXP-I: pseudo-NMI refused by an UPSTREAM QUIRK: the krane DTB's GIC +node carries "mediatek,broken-save-restore-fw" ("broken MediaTek +firmware that doesn't properly save and restore GIC priorities") and +cpufeature.c disables pseudo-NMI on it — printed at 0.000000. +NOT our wedge cause (pmOS idles/suspends constantly on the same DT +without pseudo-NMI and never dies; the breakage only matters for +priority-programmed NMI). No stacks obtained. + +Prepared EXP-J: ftrace=function + ftrace_dump_on_oops (both =y in +Ubuntu kernel). Function tracing records every CPU's executed +functions into per-CPU ring buffers; at the softlockup panic the +kernel dumps ALL CPUs' buffers to serial — INCLUDING the frozen +CPUs' last executed functions before they died. This should name +the code the dead CPUs were running, no timing luck needed. +Boot is slower (function tracing on); panic dump is LARGE (serial +@115200 — let it run, could take minutes; do not interrupt). + +## Round 64 — EXP-J dump partially captured: only CPU 6 (alive); logfile capture next + +ftrace dump-on-panic WORKS (trace lines after "SMP: stopping secondary +CPUs"). User's terminal-buffer paste contained ONLY CPU 6's section +("6" in "6d.h3." = CPU 6 hex): alive & normal (timer, mmc, idle) +through trace-ts 75720287-75726786us (~26s window before panic). +Dump order = CPU 0 first -> dead CPUs' sections (0-5,7) were the +EARLIEST output, lost to terminal scrollback while surviving CPUs +kept dumping for 5+ min at 115200. +This wedge: CPU#3 rcu_exp_gp_kthr stuck 26s (started ~78s, after +login) — later + less deterministic than coldplug wedges (ftrace +overhead shifts timing); same all-CPU-IPI-dead mechanism. +NEXT: rerun EXP-J with picocom --logfile /tmp/fb-full.log (or +| tee). Dead CPUs' final functions = first sections of dump. + +## Round 65 — EXP-J full analysis: CPUs die ENTERING WFI idle; nohlt test prepared + +Full ftrace dump (131k lines captured, timestamp-merged across CPUs, +covering trace-ts 34.82-34.925s — ~100ms before the mass freeze): +Panic 60.7s CPU#1 rcu_exp_gp_kthr waiting on CPU6. Dead: 0,4-7. +CONFIRMED DEATH POINTS (last trace event before silence): +- CPU4: 34.880s check_and_switch_context <-__schedule (entering idle) +- CPU7: 34.924s timer_base_try_to_set_idle <-tick_nohz_stop_tick + (programming wake timer, entering NOHZ idle) +- CPU5, CPU6: still running at capture cutoff (death later, uncaptured) +CPUs 0-3 sections were beyond the cutoff. NOT killed mid-execution: +cores go into WFI idle and NEVER WAKE — no trace, no IPI response +(failed-to-stop SGIs), no timer wake, no watchdog. The GIC stops +delivering to a WFI'd core. +WHY cpuidle.off=1 DOESN'T HELP: it only removes the cpuidle +framework; default idle is still cpu_do_idle (WFI). +KEY HANDOFF DELTA (arm_arch_timer.c arch_timer_select_ppi): +- depthcharge: EL1 entry -> hyp unavailable -> VIRT timer (CNTV) +- U-Boot: EL2 entry -> hyp available -> PHYS NONSECURE timer (CNTP) + (log: "cp15 timer running at 13.00MHz (phys)") — wake PPI path + never exercised by depthcharge-launched kernels on this board. +Firmware context: DT GIC node carries mediatek,broken-save-restore-fw +(upstream quirk; kernel only uses it to disable pseudo-NMI) — known +broken firmware save/restore of GIC state on this SoC family. +858921 note: workaround (= Cortex-A73 counter read) active ONLY on +CPUs 4-7; dead sets always include big cores but also little cores +(CPU0 has no workaround and dies too) — 858921 not the trigger. +maxcpus=1/4 survive: big cluster offline AND fewer idle cores. +NEXT: EXP-K = + nohlt (cpu_idle_force_poll=1; do_idle busy-polls, +NEVER executes WFI). If boot reaches userspace with 8 CPUs and no +wedge => WFI confirmed as trigger. Then: make U-Boot hand off at +EL1 (kernel would pick CNTV like depthcharge) or find GIC/SPM wake +fix. nohlt is power-hungry — diagnostic/permanent stopgap only. + +## Round 66 — EXP-K: BUG A CONFIRMED (WFI trigger); silent hang at network.target = Bug C + +EXP-K (nohlt = cpu_idle_force_poll=1, CPUs busy-poll in idle, never +WFI): the ~8-14s coldplug wedge DID NOT HAPPEN — boot sailed through +to full userspace: NM up, network.target reached (~20s+), ccifreq +spam ended normally ~8-9s. => CPUs die while EXECUTING WFI. Bug A +root cause class: core enters WFI and its GIC redistributor/timer +wake never fires again (U-Boot EL2 handoff -> kernel uses PHYS +nonsecure timer PPI as wake source; depthcharge EL1 -> CNTV virt; +broken-save-restore-fw firmware context). +REMAINING silent hang: log stops right after "Reached target +network.target" (no systemd-user-sessions line). Same point as +EXP-F's silent hang (13.5s) and pmOS-6.12-via-U-Boot 16.1s — +BUT mt76 is blacklisted in EXP-K => NOT Bug B (wifi-independent). +Called it Bug C: SMP-dependent silent deadlock after NM start; +silent because hung_task_panic was NOT set (softlockup needs a +spinning CPU; this is a blocked/deadlock state) and sysrq BREAK+l +on ttyS0 got no response in that state (serial IRQ possibly dead +too, or full freeze). +EXP-K2 prepared: same + hung_task_panic=1 hung_task_timeout_secs=10 +-> 10s after a task hangs, panic prints ALL CPU stacks = names the +deadlock. Boot EXP-K2 next; when it stops, WAIT ~15s for the +auto-panic dump (no sysrq needed). + +## Round 67 — VFS root panic was a grub entry mistake (mine), fixed + +Both reboot attempts: "Cannot open root device ... unknown-block(0,0), +available partitions: (EMPTY)" + prepare_namespace in the panic +trace = kernel got NO initramfs (prepare_namespace never runs when +an initrd is present). Without initrd there are no modules -> no +mtk_sd -> no block devices. EFI banner in both boots lacks the +INITRD=0x... line. Cause: my EXP-K edit accidentally replaced the +entry's initrd line with a second linux line and then deleted the +original linux line -> entry booted with NO initrd at all. Fixed: +initrd line restored, grub regenerated and verified (entry now has +linux+initrd+devicetree). NOT a kernel regression. Also: the entry +label stayed "EXP-K" (I edited in place; "EXP-K2" never existed as +a label — the hung_task params were active in both boots). +Next boot: EXP-K again (nohlt + hung_task_panic) — when output +stops, wait ~15s for the auto hung-task panic dump. + +## Round 68 — Bug C named: QCA Bluetooth firmware download (hci_uart); bt blacklist test ready + +eMMC journal (hostname duet) preserved the EXP-K death context that +serial never showed: ath10k_sdio wifi loaded, ModemManager + +wpa_supplicant started, then: + Bluetooth: hci0: QCA Downloading qca/rampatch_00440302.bin + kernel: ------------[ cut here ]------------ <- freeze, no more +Bug C = firmware-download deadlock, same class as Bug B (mt76 SDIO +fw download at NM time) but via QCA BT UART. With wifi blacklisted, +the BT path (hci_uart) hits the equivalent bug at user-sessions +time. Re-explains EXP-E hang and pmOS 16.1s hang. +Also: hung_task_timeout_secs=10 is NOT a valid boot param (only +sysctl kernel.hung_task_timeout_secs; default 120s in Ubuntu) — +hung_task_panic WAS active but needs 120s to fire; reboots were +too early. Boot param valid: hung_task_panic only. +NEXT: EXP-K entry now blacklists hci_uart,btqca,bluetooth on top of +mt76 + nohlt. If it reaches serial login with 8 CPUs => both +remaining boot bugs are firmware-download deadlocks under U-Boot +handoff (wifi SDIO + BT UART). Then: bisect WHICH SMP interaction +breaks fw download (candidates: SDIO/UART DMA + per-CPU IRQ wake +marginality, GIC-to-SPM wake path, or sg_table/DMA vs IOMMU-off). + +## Round 69 — BT blacklist did NOT fix the user-space hang; sysctl dump prepared + +EXP-K with hci_uart/btqca/bluetooth blacklisted: same silent hang, +stops around NM/hostnamed/ModemManager start (slightly earlier than +the network.target stop of the previous EXP-K boot — placement +varies). So Bug C is NOT (only) the QCA BT download — the journal's +cut-here during rampatch download was likely collateral, not the +cause. Firmware-download-class theory: NOT yet confirmed for C. +hung_task_timeout_secs boot param is invalid (only sysctl exists; +default 120s — user reboots too early to ever see the dump). +Prepared: /etc/sysctl.d/99-krane-hungtask.conf +(kernel.hung_task_timeout_secs=10, hung_task_all_cpu_backtrace=1, +hung_task_warnings=100) — applies to every boot of this rootfs; +10s after a task hangs, panic prints ALL CPUs' stacks and names +the blocker (v7.0 has debug_show_blocker = mutex owner). +NEXT: reboot EXP-K (nohlt, wifi+bt blacklisted). At the hang WAIT +AT LEAST 3 MINUTES. The panic dump is the deliverable. + +## Round 70 — WFI theory DEAD; new unified suspect: mtk-cci-devfreq first rate switch + +EXP-K (nohlt) definitive panic: CPU#2 kworker/2:0 stuck 22s in +rcu_barrier <- fqdir_free_fn, "failed to stop secondary CPUs +0-1,3-7" — ALL SEVEN OTHER CPUS DEAD WITH NO WFI EVER EXECUTED +(deaths at ~14.3s). The "die at idle entry" ftrace reading was an +artifact: the round-64 capture ended 3s BEFORE the deaths (boot 1 +panicked at 104.7 = death ~78s; round-65 capture = boot 2, death +~34.6s). CPUs freeze mid-execution, idle mode irrelevant. +CORRECTION of the record: round 65's "CPU4/7 died entering idle" +was premature — last-traced-event != death point for CPUs 5/6/0-3. +NEW UNIFIED SUSPECT: drivers/devfreq/mtk-cci-devfreq.c probe: +each deferred retry raises VPROC to the highest CCI OPP voltage +(mtk_ccifreq_set_voltage BEFORE devfreq registration!), then +devm_devfreq_add_device defers -517 on CPUFREQ_PARENT_DEV +(mtk-cpufreq module) = the x71 spam. When mtk-cpufreq finally +registers (real-root module storm ~9-15s), ccifreq attaches and +the passive governor immediately syncs CCI rate+voltage on the +live system. A wedged CCI/PLL/VPROC switch freezes ALL cores +mid-instruction (shared resource) — matches every panic signature. +Spam-end -> deaths correlation holds in every boot (8.7s spam end; +ftrace boot spam 34.6, deaths 34.88+; nohlt boot spam 9.5, deaths +14.3). maxcpus=1/4 survive = big-cluster policy/transition absent. +Bug A and Bug C are probably THE SAME BUG. +NEXT: EXP-K now also blacklists mtk_cci_devfreq + mtk_svs (cpufreq +still allowed). If the boot reaches serial login with 8 CPUs => +CCI DVFS stack = killer. Then bisect: cpufreq vs ccifreq, and WHY +the switch wedges under U-Boot (clock state left by payload?). + +## Round 71 — CCI/SVS exonerated (still dies, NEW victim set {3,5,6,7}); cpufreq is the last DVFS suspect + +EXP-K + mtk_cci_devfreq/mtk_svs blacklisted: STILL dies at ~14.3s +(same rcu_barrier/fqdir_free_fn victim class) BUT the dead set +changed for the first time: 3,5-7 — big cores 5,6,7 + little core +3, with CPU4 ALIVE (first time ever with 8 CPUs online). CCI/SVS +exonerated as the trigger. +Remaining DVFS piece in the 14s module storm: mediatek-cpufreq +(module mediatek-cpufreq.ko; big-cluster policy init = switch to +intermediate clock, reprogram main PLL, change shared VPROC/VSRAM +rails). Explains: deaths spanning both clusters (shared rail), +survivor variation, maxcpus=1/4 immunity (big policy never inits — +EXP-E with 4 CPUs survived cpufreq and hit mt76 Bug B at NM +instead), depthcharge immunity (clock tree left in expected state; +Jul 28 journal: "CPU4: Running at unlisted initial frequency: +1199999 KHz, changing to 1248000" — U-Boot may leave a +non-OPP-listed rate -> fatal big PLL jump). No cpufreq messages at +all appear on serial in U-Boot boots before the freeze. +NEXT: EXP-K + mediatek-cpufreq blacklisted. Serial login with 8 +CPUs => big-cluster cpufreq policy init = the killer under U-Boot +handoff. Then compare with U-Boot's leftover MPU rate (bootefi +"CPU: ..." or /proc/cpuinfo) and design the real fix (U-Boot clock +cutover or cpufreq driver quirk). + +## Round 72 — cpufreq EXONERATED too; round-55 signature; EXP-L definitive trace prepared + +EXP-K + mediatek-cpufreq blacklisted: STILL dies — and the panic is +the round-55 signature EXACTLY: udev-worker CPU#1 stuck 23s in +smp_call_function_many_cond <- kick_all_cpus_sync <- +flush_module_icache <- load_module <- finit_module. Dead set {4,5,6} +(3 big cores; CPU7 alive for the first time). DVFS fully exonerated +(cci, svs, cpufreq all blacklisted — still dies). nohlt also ruled +out (round 70: dies without any WFI). The module-load IPIs reveal +the freeze, they don't cause it; several modules were mid-load in +parallel (cros_ec_keyb(+), hid_multitouch(+), hid_google_hammer(+), +extcon(+), cros_ec_dev(+)). +Invariants now: deaths at ~9-15s in the udev storm; victim sets +vary but always include big cores; U-Boot handoff required; +depthcharge immune; maxcpus=1/4 immune; idle mode irrelevant; +DVFS irrelevant; irqaffinity/threadirqs irrelevant. +EXP-L prepared: round-63 config EXACTLY (cpuidle.off + mt76 +blacklist + ftrace=function + ftrace_dump_on_oops, NO nohlt) + +trace_buf_size=128 (bounds each CPU ring -> dump ~30k lines, +completes in minutes; frozen CPUs' tails = their true death points; +normal WFI idle so idle CPUs still trace their entry). +Boot with picocom --logfile; capture EVERYTHING until reboot. diff --git a/U-BOOT.md b/U-BOOT.md index bff79ce..e6a2d93 100644 --- a/U-BOOT.md +++ b/U-BOOT.md @@ -88,15 +88,28 @@ Then: reboot → depthcharge dev menu → "Internal storage". - pmOS still boots from the USB stick (sda) via the depthcharge menu. - pmOS chroot lives on `mmcblk0p3`. -## U-Boot tree state (branch `krane`, on top of mainline 527115ef) +## U-Boot tree state (branch `krane-updates`, on top of mainline 527115ef) Kept features (upstreamable): - `drivers/video/mt8183_scanout.c` — UCLASS_VIDEO scanout driver on the upstream `ovl0@14008000` node: revives the depthcharge scanout - (OVL_EN/OVL0_2L_EN + backlight GPIOs), parses the coreboot table - LBIO at `0xffed9000`, falls back to `OVL_L0_ADDR` (0x14008f40), - sets `uc_priv->rot = 3` (270° CW landscape console). + (OVL_EN/OVL0_2L_EN + backlight GPIOs 43/176), parses the coreboot + table LBIO at `0xffed9000`, falls back to `OVL_L0_ADDR` (0x14008f40), + sets `uc_priv->rot = 3` (270° CW landscape console). The shared + handoff code lives in `mt8183_disp.c` / `mt8183_disp.h`. +- `drivers/video/mt8183_display.c` — full cold bring-up driver + (Kconfig choice `VIDEO_MT8183_DISPLAY`, now the defconfig default; + `VIDEO_MT8183_SCANOUT` keeps the old behavior): MMSYS display gates, + SMI LARB0, panel init via `drivers/video/mtk_dsi.c` + + `drivers/phy/phy-mtk-mipi-tx.c` + `drivers/video/panel_boe_tv101wum.c`, + overlay scanout of a framebuffer at 0xFC000000 (FDT memreserve + + LMB), backlight. Falls back to the handoff revival, logging the + stage. Serial prints: `[dsi] phase 0/9/F`. +- `drivers/gpio/mt8183_gpio.c` — minimal MT8183 GPIO driver (dir/dout/ + din; pinmux left to firmware) for the panel reset / supply-enable pads. +- `drivers/clk/mediatek/clk-mt8183.c` — now also models the MMSYS + display gates (CLK_MM_*) as a clock provider on `mmsys@14000000`. - `board/mediatek/mt8183/mt8183.c` — `get_page_table_size()` override (0x40000): the fb and coreboot-table dynamic mappings exhaust the default page-table budget (Round 31). Any new post-reloc