Skip to content

SIGSEGV in _tmxr_activate_delay after RESTORE when a telnet re-attach fails (UNIT_TM_POLL is restored, uptr->tmxr is not) #576

Description

@ChristineTham

SIGSEGV in _tmxr_activate_delay after RESTORE when a telnet re-attach fails (UNIT_TM_POLL is restored, uptr->tmxr is not)


Summary

SAVE persists a unit's dynflags, including UNIT_TM_POLL, but uptr->tmxr is a runtime pointer that only a successful tmxr_attach can set. If a RESTORE's recorded telnet re-attach fails, the unit comes back tagged as a mux polling unit with a NULL tmxr backpointer, and the next GO/CONTINUE dereferences NULL in _tmxr_activate_delay.

Restore reports success, so nothing signals that the machine is now one command away from a crash.

Reproduced on a stock vax780 with no disk image and no operating system — the guest is a two-instruction counter loop.

Version

VAX 11/780 simulator Open SIMH V4.1-0 Current    git commit id: a1f57fa3
Compiler: GCC Apple LLVM 21.0.0 (clang-2100.1.1.101)
Build Tool: simh-makefile (Release Build)
Host: macOS (Darwin 25.6.0), arm64

Asynchronous I/O support is compiled in (not enabled at run time — set asynch was never issued).

Repro

guest.ini — deposits INCL @#400 ; BRB .-8 at 0x200, attaches the console to telnet, saves. No telnet client ever needs to connect; the attach alone sets the flag.

set console telnet=7654
d -b 200 D6
d -b 201 9F
d -b 202 00
d -b 203 04
d -b 204 00
d -b 205 00
d -b 206 11
d -b 207 F8
dep pc 200
save crash.sav
exit

restore.ini:

restore crash.sav
go 200
./BIN/vax780 guest.ini                     # writes crash.sav, exits cleanly

# Occupy the console port so the restore's re-attach fails:
python3 -c "import socket,time; s=socket.socket();
s.setsockopt(socket.SOL_SOCKET,socket.SO_REUSEADDR,1);
s.bind(('127.0.0.1',7654)); s.listen(1); time.sleep(60)" &

./BIN/vax780 restore.ini                   # SIGSEGV

Exit status 11 (SIGSEGV). restore itself does not report failure here; a show console after the failed restore reports Connected to console window, i.e. the telnet console is gone while the TTI unit still carries UNIT_TM_POLL from the save file.

Crash

stop reason = EXC_BAD_ACCESS (code=1, address=0x0)
  * frame #0: _tmxr_activate_delay + 48
    frame #1: tmxr_clock_coschedule_tmr + 108
    frame #2: tti_svc + 28
    frame #3: sim_process_event + 952
    frame #4: sim_instr + 1952
    frame #5: run_cmd + 2544
    frame #6: do_cmd_label + 2080
    frame #7: main + 3412

The two frames on top pin both halves of the bad state simultaneously: reaching _tmxr_activate_delay at all requires UNIT_TM_POLL to be set (the guard in frame #1), and faulting at +48 requires uptr->tmxr to be NULL.

Mechanism

  1. tmxr_attach_ex is the only thing that establishes the pair — uptr->tmxr (sim_tmxr.c:4456, and per-line at 4481/4485) together with UNIT_TM_POLL (4469, 4479, 4483).
  2. SAVE writes uptr->dynflags (scp.c:8658) and RESTORE reads it back (scp.c:8924). UNIT_TM_POLL therefore returns in a fresh process where uptr->tmxr is still NULL — a pointer cannot be, and is not, part of the save file.
  3. sim_rest re-attaches previously attached units (scp.c:9081) but treats attach failure as non-fatal (9083): it prints and moves on, leaving the restored flag in place.
  4. All three schedulers gate on the flag alone — tmxr_activate (4815), tmxr_activate_after (4848), tmxr_clock_coschedule_tmr (4899) — and call _tmxr_activate_delay (4763), which does TMXR *mp = (TMXR *)uptr->tmxr; (4765) and immediately dereferences mp->lines (4769).

tmxr_detach clears UNIT_TM_POLL alongside the pointers (4745–4759), so an ordinary detach is safe. The restore path is the hole: it re-imports the flag without re-establishing what the flag implies.

How this is hit in practice

Any workflow that suspends with SAVE and resumes with RESTORE across process restarts. If the previous incarnation's telnet pairs are still in TIME_WAIT, the listener bind fails (bind error 48) because tmxr_open_master does not set SO_REUSEADDR; the restore then proceeds with polling flags set and no mux behind them, and the first CONTINUE segfaults. The occupied-port repro above is just a deterministic stand-in for that race.

Suggested fix

Two levels, either or both:

  • Defensive — in _tmxr_activate_delay, if (mp == NULL) return interval;, or have the three callers test uptr->tmxr alongside UNIT_TM_POLL. This removes the crash class wherever else a stale flag might arise.
  • Root cause — make restore leave consistent state: when scp_attach_unit fails in sim_rest, clear the polling state the flag implies for that unit (UNIT_TM_POLL, and the UNIT_ATT the save file implied), so it degrades to an ordinary non-mux unit. Alternatively, mask UNIT_TM_POLL out of the restored dynflags and let a successful attach re-establish it, since it is runtime-derived state rather than saved machine state.

Raising restore-time attach failure to a visible error (or a distinct return status) would also help — currently the session looks healthy right up to the crash.

Note on a workaround that looks right and isn't

reset <dev> re-links the unit pointers as a side effect and does make the crash disappear. It also zeroes device CSR state the restored guest OS had already configured — interrupt enable included — so the restored guest silently stops seeing console input. Anyone hitting this crash should not take reset as the fix.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions