SIGSEGV in _tmxr_activate_delay after RESTORE when a telnet re-attach fails (UNIT_TM_POLL is restored, uptr->tmxr is not)
Summary
SAVE persists a unit's dynflags, including UNIT_TM_POLL, but uptr->tmxr is a runtime pointer that only a successful tmxr_attach can set. If a RESTORE's recorded telnet re-attach fails, the unit comes back tagged as a mux polling unit with a NULL tmxr backpointer, and the next GO/CONTINUE dereferences NULL in _tmxr_activate_delay.
Restore reports success, so nothing signals that the machine is now one command away from a crash.
Reproduced on a stock vax780 with no disk image and no operating system — the guest is a two-instruction counter loop.
Version
VAX 11/780 simulator Open SIMH V4.1-0 Current git commit id: a1f57fa3
Compiler: GCC Apple LLVM 21.0.0 (clang-2100.1.1.101)
Build Tool: simh-makefile (Release Build)
Host: macOS (Darwin 25.6.0), arm64
Asynchronous I/O support is compiled in (not enabled at run time — set asynch was never issued).
Repro
guest.ini — deposits INCL @#400 ; BRB .-8 at 0x200, attaches the console to telnet, saves. No telnet client ever needs to connect; the attach alone sets the flag.
set console telnet=7654
d -b 200 D6
d -b 201 9F
d -b 202 00
d -b 203 04
d -b 204 00
d -b 205 00
d -b 206 11
d -b 207 F8
dep pc 200
save crash.sav
exit
restore.ini:
./BIN/vax780 guest.ini # writes crash.sav, exits cleanly
# Occupy the console port so the restore's re-attach fails:
python3 -c "import socket,time; s=socket.socket();
s.setsockopt(socket.SOL_SOCKET,socket.SO_REUSEADDR,1);
s.bind(('127.0.0.1',7654)); s.listen(1); time.sleep(60)" &
./BIN/vax780 restore.ini # SIGSEGV
Exit status 11 (SIGSEGV). restore itself does not report failure here; a show console after the failed restore reports Connected to console window, i.e. the telnet console is gone while the TTI unit still carries UNIT_TM_POLL from the save file.
Crash
stop reason = EXC_BAD_ACCESS (code=1, address=0x0)
* frame #0: _tmxr_activate_delay + 48
frame #1: tmxr_clock_coschedule_tmr + 108
frame #2: tti_svc + 28
frame #3: sim_process_event + 952
frame #4: sim_instr + 1952
frame #5: run_cmd + 2544
frame #6: do_cmd_label + 2080
frame #7: main + 3412
The two frames on top pin both halves of the bad state simultaneously: reaching _tmxr_activate_delay at all requires UNIT_TM_POLL to be set (the guard in frame #1), and faulting at +48 requires uptr->tmxr to be NULL.
Mechanism
tmxr_attach_ex is the only thing that establishes the pair — uptr->tmxr (sim_tmxr.c:4456, and per-line at 4481/4485) together with UNIT_TM_POLL (4469, 4479, 4483).
SAVE writes uptr->dynflags (scp.c:8658) and RESTORE reads it back (scp.c:8924). UNIT_TM_POLL therefore returns in a fresh process where uptr->tmxr is still NULL — a pointer cannot be, and is not, part of the save file.
sim_rest re-attaches previously attached units (scp.c:9081) but treats attach failure as non-fatal (9083): it prints and moves on, leaving the restored flag in place.
- All three schedulers gate on the flag alone —
tmxr_activate (4815), tmxr_activate_after (4848), tmxr_clock_coschedule_tmr (4899) — and call _tmxr_activate_delay (4763), which does TMXR *mp = (TMXR *)uptr->tmxr; (4765) and immediately dereferences mp->lines (4769).
tmxr_detach clears UNIT_TM_POLL alongside the pointers (4745–4759), so an ordinary detach is safe. The restore path is the hole: it re-imports the flag without re-establishing what the flag implies.
How this is hit in practice
Any workflow that suspends with SAVE and resumes with RESTORE across process restarts. If the previous incarnation's telnet pairs are still in TIME_WAIT, the listener bind fails (bind error 48) because tmxr_open_master does not set SO_REUSEADDR; the restore then proceeds with polling flags set and no mux behind them, and the first CONTINUE segfaults. The occupied-port repro above is just a deterministic stand-in for that race.
Suggested fix
Two levels, either or both:
- Defensive — in
_tmxr_activate_delay, if (mp == NULL) return interval;, or have the three callers test uptr->tmxr alongside UNIT_TM_POLL. This removes the crash class wherever else a stale flag might arise.
- Root cause — make restore leave consistent state: when
scp_attach_unit fails in sim_rest, clear the polling state the flag implies for that unit (UNIT_TM_POLL, and the UNIT_ATT the save file implied), so it degrades to an ordinary non-mux unit. Alternatively, mask UNIT_TM_POLL out of the restored dynflags and let a successful attach re-establish it, since it is runtime-derived state rather than saved machine state.
Raising restore-time attach failure to a visible error (or a distinct return status) would also help — currently the session looks healthy right up to the crash.
Note on a workaround that looks right and isn't
reset <dev> re-links the unit pointers as a side effect and does make the crash disappear. It also zeroes device CSR state the restored guest OS had already configured — interrupt enable included — so the restored guest silently stops seeing console input. Anyone hitting this crash should not take reset as the fix.
SIGSEGV in
_tmxr_activate_delayafter RESTORE when a telnet re-attach fails (UNIT_TM_POLLis restored,uptr->tmxris not)Summary
SAVEpersists a unit'sdynflags, includingUNIT_TM_POLL, butuptr->tmxris a runtime pointer that only a successfultmxr_attachcan set. If aRESTORE's recorded telnet re-attach fails, the unit comes back tagged as a mux polling unit with a NULLtmxrbackpointer, and the nextGO/CONTINUEdereferences NULL in_tmxr_activate_delay.Restore reports success, so nothing signals that the machine is now one command away from a crash.
Reproduced on a stock
vax780with no disk image and no operating system — the guest is a two-instruction counter loop.Version
Asynchronous I/O support is compiled in (not enabled at run time —
set asynchwas never issued).Repro
guest.ini— depositsINCL @#400 ; BRB .-8at 0x200, attaches the console to telnet, saves. No telnet client ever needs to connect; the attach alone sets the flag.restore.ini:Exit status 11 (SIGSEGV).
restoreitself does not report failure here; ashow consoleafter the failed restore reportsConnected to console window, i.e. the telnet console is gone while the TTI unit still carriesUNIT_TM_POLLfrom the save file.Crash
The two frames on top pin both halves of the bad state simultaneously: reaching
_tmxr_activate_delayat all requiresUNIT_TM_POLLto be set (the guard in frame #1), and faulting at+48requiresuptr->tmxrto be NULL.Mechanism
tmxr_attach_exis the only thing that establishes the pair —uptr->tmxr(sim_tmxr.c:4456, and per-line at 4481/4485) together withUNIT_TM_POLL(4469, 4479, 4483).SAVEwritesuptr->dynflags(scp.c:8658) andRESTOREreads it back (scp.c:8924).UNIT_TM_POLLtherefore returns in a fresh process whereuptr->tmxris still NULL — a pointer cannot be, and is not, part of the save file.sim_restre-attaches previously attached units (scp.c:9081) but treats attach failure as non-fatal (9083): it prints and moves on, leaving the restored flag in place.tmxr_activate(4815),tmxr_activate_after(4848),tmxr_clock_coschedule_tmr(4899) — and call_tmxr_activate_delay(4763), which doesTMXR *mp = (TMXR *)uptr->tmxr;(4765) and immediately dereferencesmp->lines(4769).tmxr_detachclearsUNIT_TM_POLLalongside the pointers (4745–4759), so an ordinary detach is safe. The restore path is the hole: it re-imports the flag without re-establishing what the flag implies.How this is hit in practice
Any workflow that suspends with
SAVEand resumes withRESTOREacross process restarts. If the previous incarnation's telnet pairs are still inTIME_WAIT, the listener bind fails (bind error 48) becausetmxr_open_masterdoes not setSO_REUSEADDR; the restore then proceeds with polling flags set and no mux behind them, and the firstCONTINUEsegfaults. The occupied-port repro above is just a deterministic stand-in for that race.Suggested fix
Two levels, either or both:
_tmxr_activate_delay,if (mp == NULL) return interval;, or have the three callers testuptr->tmxralongsideUNIT_TM_POLL. This removes the crash class wherever else a stale flag might arise.scp_attach_unitfails insim_rest, clear the polling state the flag implies for that unit (UNIT_TM_POLL, and theUNIT_ATTthe save file implied), so it degrades to an ordinary non-mux unit. Alternatively, maskUNIT_TM_POLLout of the restoreddynflagsand let a successful attach re-establish it, since it is runtime-derived state rather than saved machine state.Raising restore-time attach failure to a visible error (or a distinct return status) would also help — currently the session looks healthy right up to the crash.
Note on a workaround that looks right and isn't
reset <dev>re-links the unit pointers as a side effect and does make the crash disappear. It also zeroes device CSR state the restored guest OS had already configured — interrupt enable included — so the restored guest silently stops seeing console input. Anyone hitting this crash should not takeresetas the fix.