Skip to content

majestic lite: is WebRTC intended for 8 MB NOR devices? (+37% since April) #2344

Description

@bneigher

Is WebRTC intended for the lite majestic variant on 8 MB NOR?

Not a bug report — a sizing question, with measurements, from someone who may
simply be holding it wrong.

majestic.<family>.lite has grown 37% since April:

2026-04-22    685,672 B
2026-08-15    811,176 B   (+125,504)
2026-08-29    938,548 B   (+127,372)

Section-level, the last step is .text +86,084 and .rodata +34,712. The new
dynamic symbols identify it fairly clearly:

mbedtls_ssl_conf_dtls_srtp_protection_profiles
mbedtls_ssl_get_dtls_srtp_negotiation_result
mbedtls_ssl_conf_dtls_cookies
mbedtls_ssl_conf_export_keys_ext_cb
evhttp_send_reply_chunk_with_cb, evdns_base_load_hosts, ...

plus strings /ws/webrtc, Adapt bitrate to WebRTC viewers. So: DTLS-SRTP and
libevent-based signalling — a WebRTC stack.

The question. On a gk7202v300 with 8 MB NOR, the rootfs partition is 5632k.
For devices whose video transport is external (ours is a separate agent speaking
WebTransport/QUIC; others use RTSP into an NVR, or RTMP), that WebRTC stack is
dead weight — but it is not optional: BR2_OPENIPC_MAJESTIC defaults to "lite",
we are already on it, and majestic is a prebuilt binary so we cannot configure the
feature out. The published alternatives are ultimate (larger) and fpv (a
reduced build we cannot use — we need the full audio and encoder pipeline).

Is WebRTC in lite deliberate for 8 MB devices, or would a variant without it
make sense — something like lite-nowebrtc? I am happy to be told this is the
intended direction and that 8 MB NOR is simply reaching end of life; that is a
perfectly good answer and I would rather know than guess.

For context, not as a complaint: we absorbed it. Upstream's base actually
shrank 464K in the same window, but every byte came from files we already
trimmed locally (libnl*, jshn, jsonfilter, libblobmsg_json,
libjson_script), so it bought us nothing while majestic's growth cost us in
full. We have since reclaimed 64k by shrinking the kernel partition (the uImage
measured 1,883,236 / 1,883,236 / 1,883,228 across three consecutive nightlies, so
1920k was over-provisioned). That gives us roughly one more cycle at the current
rate.

Asking because if the answer is "a no-WebRTC lite is reasonable", that likely
helps every 8 MB NOR device, not just ours.

Activity

  1. widgetii commented on Aug 31, 2026

    @widgetii
    Member

    Your measurements are right, and this is a good question to have asked.

    Same numbers here, against today's nightly — note gk7202v300 builds the
    gk7205v200 majestic family, so we are looking at the same binary:

    lite       938,684 B
    fpv        515,920 B
    ultimate 1,439,200 B
    

    And yes, that is WebRTC: DTLS-SRTP, ICE, and the cloud signalling path. The
    mbedtls dynamic symbol count goes 28 → 88 between fpv and lite, which is the
    DTLS-SRTP surface arriving.

    Is it deliberate? Yes. The three flavours are curated by what the camera is
    for, not by how much flash it has — that is now written down in
    Majestic Streamer → Lite, Ultimate and FPV. Lite is the everyday
    build: audio, talkback, RTMP, WebRTC. FPV is the latency-first build with audio,
    WebRTC, SIP and RTMP compiled out. There is no size tier, and that is on purpose.

    Would a lite-nowebrtc make sense? No, and the reason is worth explaining,
    because it is about direction rather than bytes. WebRTC is no longer an optional
    extra sitting next to the others. It is the WebUI's default preview transport —
    MSE and MJPEG are the fallback path now, not the other way round — and it is how
    cloud.enabled reaches a camera without port-forwarding. The optimisation work
    queued behind it is aimed squarely at cellular and at OpenIPC Cloud, where
    WebRTC's congestion control and NAT traversal are the entire point. A build
    without it would get steadily less of what we are actually working on, and Lite
    is published for every supported SoC — so a fourth flavour that had to track
    Lite's coverage is the expensive one to add, not FPV's.

    Is 8 MB NOR reaching end of life? Not today, and there is no plan to drop it.
    For what it is worth, upstream is not comfortable on your board either — nightly
    2026-08-30, gk7202v300 lite:

    rootfs.squashfs  5,013,504 B / 5120 KB cap  →  224 KB headroom
    uImage           1,883,212 B / 2048 KB cap  →  209 KB headroom
    

    So you are not an outlier, and the caps are enforced per build, which means 8 MB
    cannot silently stop working — it would fail a build first. What I will not
    promise is that every future feature gets compiled out to keep it fitting. If
    there is a hardware refresh anywhere in your roadmap, 16 MB is cheap insurance.
    Your kernel-partition reclaim is sound reasoning, incidentally; make BOARD=<x> size-report will show you where the rest of it is going.

    On getting the build you actually want. You describe a device with its own
    transport agent and a private trim of the rootfs, which reads like a product you
    ship rather than a hobby camera. If that is right, then the route to a Lite
    without WebRTC already exists and it is not a public flavour: Majestic is
    Prosperity Public License 3.0.0 — free forever for personal and
    noncommercial use, thirty-day commercial trial — and commercial licenses include
    custom builds with feature selection. That is a better answer than
    lite-nowebrtc would be, because you would get exactly the set you want, with
    the full audio and encoder pipeline intact, rather than a compromise shared with
    every other 8 MB device. business@openipc.org if you want to talk about it.

    If it is not right and this is noncommercial, say so and ignore that paragraph —
    the rest of the answer stands either way.

  2. bneigher commented on Aug 31, 2026

    @bneigher
    ContributorAuthor

    I was wondering what your thoughts on using WebTransport for majestic over WebRTC? QUIC is quite stable now and HTTP3 is a much lighter technology to use, not to mention is great for Client Server exchanges

  3. bneigher commented on Aug 31, 2026

    @bneigher
    ContributorAuthor

    Thank you — that is a much better answer than the one I was fishing for, and the
    reasoning about direction rather than bytes is the part I needed to hear.

    On the licensing paragraph: it does not apply. This is a hobby project. Three
    boards on a bench, no product, nothing sold or planned to be. The "own transport
    agent and private rootfs trim" reads commercial because I have been over-engineering
    a home camera for months, not because anyone is paying for it. I would rather say
    that plainly than let an ambiguity sit in a public thread.

    So lite with WebRTC is simply what I get, and that is fine. I will stop trying to
    shrink my way out of it.

    Two corrections I am taking away, both mine:

    • fpv was on my shortlist until you said audio is compiled out. I had assumed it
      was lite-minus-latency-features. Audio is load-bearing here — the camera does
      on-device audio classification — so that would have been a slow, confusing
      discovery.
    • I had my kernel headroom backwards. I reclaimed 64k from the kernel partition
      (1920k → 1856k) after verifying my uImage fits, leaving 17,308 B spare. But you
      size against a 2048k cap and are at 1,883,212 — so you have ~214k of room you may
      legitimately grow into, and roughly 196k of that would not fit my partition. My
      reclaim is safe against today's kernel and fragile against yours. It fails a build
      gate rather than a board, but I would not have known to look without your numbers.

    make BOARD=<x> size-report is now on my list; I had not been running it.

    The part I actually want to ask about

    You mentioned WebRTC's congestion control and NAT traversal being the point for
    cellular and Cloud. That is adjacent to something I know well.

    My day job is transport: WebTransport and QUIC, in depth. The agent on these
    cameras is a Rust binary that publishes H.264/H.265 and Opus over WebTransport to a
    cloud relay, with a datagram path for intercom, and it has been running on 8 MB
    gk7202v300 hardware for months — so I have practical numbers on what QUIC costs on
    a 64 MB, single-core ARMv7 camera, not just opinions.

    Where that might be useful to OpenIPC, if any of it is interesting:

    • WebRTC and WebTransport share the hard parts — congestion control, pacing,
      loss recovery, MTU discovery. If the Cloud work is tuning any of that, I am happy
      to review, test on real constrained hardware, or contribute.
    • WebTransport as a publish path is genuinely simpler than WebRTC when the
      camera talks to your infrastructure rather than to a browser peer: one QUIC
      connection, no ICE, no SFU, no DTLS-SRTP. It is not a WebRTC replacement — it does
      not do browser-to-camera P2P — but for camera-to-cloud it removes most of the
      moving parts. Whether that is worth a flag in majestic is your call entirely; I
      mention it as something I could prototype rather than something I think you are
      missing.
    • Congestion control on cellular is where I would expect to be most useful.
      BBR-ish behaviour on a link that is simultaneously lossy and buffer-bloated is
      a place where good defaults matter more than features.

    Entirely understood if none of this fits your roadmap, or if a transport
    contribution from someone who has been here about three weeks is premature. But I
    have enjoyed the work — the DVP gate landing in #2276 came out of exactly this kind
    of exchange — and I would rather ask where help is wanted than turn up with an
    unsolicited PR.

  4. widgetii commented on Aug 31, 2026

    @widgetii
    Member

    Thanks for saying it plainly — and no harm done. I wrote that paragraph conditional
    precisely because I could not tell, and "over-engineering a home camera for months"
    is a description I recognise from the inside. Hobby it is.

    It also changes the answer in your favour, because nothing is stopping you stripping
    the image down to what you actually run.

    The thing you should delete is not WebRTC

    You have been trying to shrink the streamer, which is the one part you cannot
    change. Meanwhile you are shipping a web interface that, by your own description,
    nothing in your setup uses — your agent is the interface.

    I went and measured this properly rather than estimating, because uncompressed
    byte counts are misleading here and I did not want to hand you another wrong number
    after you already corrected one of mine. Real squashfs, your board, nightly
    2026-08-30, against the 5120 KB cap:

    gk7202v300 lite rootfs.squashfs headroom
    as shipped 5,013,504 224 KB
    WebUI removed 4,882,432 352 KB
    majestic → fpv 4,755,456 476 KB
    both 4,624,384 604 KB

    Dropping the WebUI is worth 131,072 B — exactly 128 KB of flash. That is half of
    everything the entire FPV downgrade would buy you, and you keep audio, WebRTC, RTMP
    and SIP. Given audio is load-bearing for your classifier, it is the whole of what you
    can actually collect.

    Same measurement on hi3516ev200, to check the number was not an artefact of one
    board: 4,837,376 → 4,706,304. Also 131,072 B. It is one squashfs block, which is why
    it lands so round.

    The method, since you will want to check it rather than take my word:

    fakeroot -s fr.state unsquashfs -d root rootfs.squashfs.gk7202v300
    # ... edit root/ ...
    fakeroot -i fr.state mksquashfs root out.sq -noappend -b 131072 -comp xz

    Rebuilding the unmodified tree that way reproduces the shipped image byte-exactly
    on both boards — 5,013,504 and 4,837,376 — so the deltas above are real flash, not a
    compression estimate.

    One incidental finding from doing this: between the 08-30 nightly and the 08-31
    publish, majestic grew 12,580 bytes of binary, which landed as 4,096 more bytes of
    flash. In a day. Your instinct that this is worth watching is correct.

    How to turn it off: one line in your board defconfig.
    gk7202v300_lite_defconfig has BR2_PACKAGE_MAJESTIC_WEBUI=y at line 63. haserl
    is selected by that symbol, so it leaves with it.

    What survives — the part worth checking before you pull the trigger:

    • Majestic's HTTP API, RTSP, ONVIF, snapshots, HLS, the WebSocket endpoints. All
      of that lives in the majestic binary, not in majestic-webui. Your agent keeps
      everything it talks to.
    • First-boot claiming. /setup.html and POST /setup are served by majestic
      itself — I checked, there is no setup* anywhere under /var/www. A camera with
      the WebUI removed can still be claimed from a browser, and openipc-claim over
      SSH is untouched regardless.

    What you lose is the browser UI: settings pages, preview, log viewer. If your agent
    covers that, it is 128 KB you are carrying for nothing.

    One caveat on the fpv row above, so you do not read more into it than it holds:
    that is WebRTC and audio and SIP and RTMP removed together. Majestic ships
    prebuilt, so I cannot isolate WebRTC on its own — 252 KB is an upper bound on the
    whole bundle, not a price tag for WebRTC.

    Why the protocol set is what it is

    WebRTC is the only sub-second path a browser gives you natively. No plugin, no
    polyfill, hardware-accelerated decode, congestion control running end to end to the
    actual viewer. Nothing else in a browser does that, which is why it became the
    default preview transport rather than one option among several.

    MSE over WebSockets is the other half, and it is deliberate. It works in every
    browser and every awkward network, costs the camera almost nothing, and buys a
    second or two of latency. MSE is the floor; WebRTC is the ceiling.

    Those two are what we can afford, and that is a flash decision rather than a taste
    one — every protocol is permanently resident in a binary with roughly half a megabyte
    of flash budget on the smallest boards we still ship. There is no third slot. Picking
    one would mean deleting one of those two, and neither can go.

    For the shrinking hobby specifically

    You will get more out of this than out of size-report:
    https://openipc.github.io/firmware-explorer/

    It renders each nightly × platform from the sizes.<plat>.json sidecar the build
    publishes: a treemap of packages by bytes, a sortable package table that expands to
    per-file detail, a module table flagging which .kos are autoloaded from
    /etc/modules versus sitting there on demand — that filter is a good compaction-
    candidate finder — a panel showing what the finalize hooks already stripped, and a
    drift view that diffs any two builds for a platform sorted by byte delta.

    The drift view is the one for you. It would have shown you the majestic step as a
    single row, and it will show you the next one before it costs you an evening of
    readelf.

    On WebTransport and QUIC

    Straight answer, since you asked straight: not as a third transport in majestic,
    and the reason is the flash paragraph above rather than anything about the design.
    You are right that camera-to-cloud over one QUIC connection removes most of the
    moving parts — no ICE, no SFU, no DTLS-SRTP — and if the budget were different that
    would be an interesting argument to have. It is not, so it is not.

    Where your expertise lands is the part you identified yourself. Congestion
    control, pacing, loss recovery and MTU discovery are shared between the two, and that
    is where the Cloud and cellular work actually is. BBR-ish behaviour on a link that is
    lossy and buffer-bloated at the same time is exactly the problem, and "good defaults
    matter more than features" is the correct instinct.

    The scarce thing there is not opinions — it is numbers from real constrained
    hardware, and you have months of them.

    So: please post the numbers. Concretely, what I would find most useful —

    • CPU cost of the QUIC path under sustained publish. What fraction of that single
      ARMv7 core goes to transport rather than to the encoder, at what bitrate.
    • Resident memory of the Rust agent, and what it peaks at when the link degrades.
    • What the congestion controller does when the cell link goes bad — how fast it
      backs off, how fast it recovers, whether it oscillates.
    • Anything that surprised you about QUIC on a 64 MB device. Handshake cost, crypto
      without AES acceleration, buffer sizing.

    Nobody has published what this costs on this class of chip. That write-up is worth
    something on its own, whether or not a line of it ever becomes a majestic patch, and
    it is the input that would make a congestion-control conversation concrete instead of
    theoretical. An issue here is a fine place for it, or the wiki if it grows.

    I am not going to promise you a roadmap slot in a GitHub thread. But three weeks in
    with #2276 already merged is not "premature", and if you want to go deeper on the
    Cloud side than an issue comment supports, d.ilyin@openipc.org.

    Leaving this open for follow-ups; the WebRTC question itself is answered.

  5. widgetii commented on Aug 31, 2026

    @widgetii
    Member

    Correction to my last comment. I said I could not isolate WebRTC from the prebuilt
    binaries and that the 252 KB fpv delta was an upper bound on the whole bundle. That
    was true of the published tarballs, but it was the wrong place to stop — I can build
    majestic both ways. So I did.

    hi3516ev200, lite, Release, stripped, same toolchain and the same dependency tree,
    with exactly one build option changed. The ENABLE_RTC=ON build comes out at
    946,188 bytes, the same size as the published lite, so this is the shipped
    configuration rather than an approximation of it.

    WebRTC costs 82,376 bytes of binary, and 44 KB of flash

    section RTC=ON RTC=OFF delta
    .text 626,278 573,934 +52,344
    .rodata 244,820 224,012 +20,808
    .data 7,476 3,172 +4,304
    dynamic linking¹ — — +3,696
    .bss 65,400 65,396 +4
    binary 946,188 863,812 +82,376

    ¹ .dynstr +1,264, .dynsym +800, .plt +616, .rel.plt +408, .got +204,
    .hash +200, .rel.dyn +104, .gnu.version +100, .data.rel.ro +56.

    In the rootfs, which is the number that decides whether your image fits — same
    squashfs method as before, baseline reproducing the shipped image byte-exactly:

    hi3516ev200 lite rootfs.squashfs headroom
    RTC=ON (as shipped) 4,841,472 392 KB
    RTC=OFF 4,796,416 436 KB

    45,056 bytes. 44 KB of flash — 8.7% of the binary, 0.9% of the rootfs.

    A footnote while I had the two builds side by side: 8,440 of those 82,376 bytes are
    the SoC crypto engine path behind the SRTP transform, which is compiled in on
    hi3516ev200 and hi3516cv500 only. It makes no difference to the squashfs at all —
    it is smaller than the 128 KB block granularity.

    What that changes

    The advice in my last comment holds, and gets stronger. Ranked by what each is
    actually worth on your board:

    flash
    The WebUI you are not using 128 KB
    Everything fpv removes (WebRTC + audio + SIP + RTMP) 252 KB
    WebRTC alone 44 KB

    WebRTC is about a sixth of the fpv delta, and the WebUI is worth nearly three times
    WebRTC on its own. If you had somehow got the lite-nowebrtc you originally asked
    for, it would have bought you 44 KB — less than a third of what deleting a web
    interface you never open buys you, and you would have given up the browser path
    permanently to get it.

    So the thing worth deleting was never the streamer. That was the answer I gave you
    before I had the number; I am glad the number agrees, because it could easily have
    gone the other way and I would have owed you a different comment.

    One correction to your figures

    You attributed .text +86,084 and .rodata +34,712 to the WebRTC step. WebRTC
    itself is .text +52,344 and .rodata +20,808 — so roughly 60% of that step was the
    WebRTC stack landing, and the other 40% was unrelated work that happened to land in
    the same window. Your identification was right; the attribution was a bit generous to
    WebRTC. Worth knowing if you are extrapolating a growth rate from that step, because
    it means the per-feature cost is lower and the general drift is higher than it looked.

    Still very much interested in the QUIC numbers whenever you get to them.

  6. bneigher commented on Aug 31, 2026

    @bneigher
    ContributorAuthor

    The QUIC numbers, as asked. All measured today on a GK7202V300 bench board; where
    something is inferred rather than measured I say so.

    Hardware. GK7202V300, ARMv7l Cortex-A7, one core, USER_HZ 100. CPU features:
    half thumb fastmult vfp edsp neon vfpv3 tls vfpv4 idiva idivt vfpd32 lpae evtstrm —
    no aes, no pmull, no sha1/sha2. The ARM crypto extensions are ARMv8-A;
    nothing on this class of chip has them. /proc/crypto shows aes-generic.

    One correction to the premise, and it matters for the memory numbers: these are sold
    as 64 MB boards, but MemTotal is 35,268 kB. The rest is reserved for the ISP/MMZ
    before Linux sees it.

    Stack. Rust agent: wtransport 0.7.2 → quinn 0.11.11 / quinn-proto 0.11.17,
    rustls 0.23.43 with the ring provider. No TransportConfig override, so this is
    quinn's defaults — Cubic, default windows. Only keep_alive_interval (3s) and
    max_idle_timeout (30s) are set. Publish is H.265 640×360@12 + AAC to a cloud relay;
    the same binary also serves 1080p H.264 over WebTransport on the LAN.

    Wire format, because it turns out to be the dominant cost: every access unit —
    each video AU, each audio frame — is sent as its own unidirectional QUIC stream.

    1. CPU: what fraction of the core is transport

    25-second windows, utime+stime from /proc/<pid>/stat, single core so percentages
    are of the whole machine. The publish-only row is n=3, reproduced within ±0.2.

    state agent majestic wlan0 TX
    agent stopped — encoder only — 19.5% 0
    cloud publish only (74 kbit/s media) 8.4% 21.9% ~105 kbit/s
    + 1 LAN viewer, substream 12.5% 23.8% 192 kbit/s
    + 1 LAN viewer, 1080p H.264 20.1% 24.4% 1292 kbit/s

    At a realistic load the split is roughly 20% encoder, 20% transport — transport is
    not a rounding error next to the encoder, it is the same order of magnitude.

    majestic costs ~2.4 points just to serve the loopback RTSP tap (19.5% → 21.9% with
    no network client involved). If you are counting the price of an agent-style
    architecture, that is part of it and it is not obvious.

    Solving for the marginal cost using the two viewer deltas as a 2×2 (per-stream α,
    per-byte β):

    • α ≈ 0.156% of core per stream/second
    • β ≈ 0.0046% of core per kbit/s — i.e. 4.6% of the core per Mbit/s

    Those two terms come out equal at ~1.35 Mbit/s for a 40 stream/s feed. Below that,
    stream count dominates the bytes. A 12 fps substream at 74 kbit/s spends more core
    on stream setup/teardown than on its own payload.

    That is the most useful thing I have learned running QUIC here, and it is a design
    lesson rather than a tuning one: one-stream-per-frame is clean, gives you per-frame
    framing for free, and lets a lagging reader drop whole frames — but on this class of
    CPU you pay for it per frame, not per byte. If I were starting again I would put
    frames on a small number of long-lived streams and frame them myself.

    Caveat on the 8.4% baseline: that is the whole agent, which also does RTSP demux
    and runs an on-device audio classifier. It is not 8.4% of pure transport. The α and
    β figures are clean, because they are deltas with everything else held constant.

    2. Memory

    VmRSS steady 4.3–4.5 MB
    VmHWM 6.0 MB
    VmSize 15.3 MB

    Under a degraded link VmHWM did not move. During sustained radio saturation RSS
    oscillated between 4,336 and 5,532 kB and the high-water mark stayed pinned at exactly
    6,000 kB — the same value it had before the test. Whatever else QUIC costs here, it
    does not balloon its buffers when the link goes bad. On a 34 MB box that was what I
    was most worried about, and it turned out to be a non-issue.

    A trap for anyone measuring this: ps reports %MEM from VSZ on busybox. It
    showed the agent at 43% of the box. Real resident is 12%.

    3. Handshake

    Full QUIC + WebTransport establishment, camera as the TLS server (so the camera
    does the server-side signature), over WiFi, software crypto only:

    trial 1: session=0.065s config=0.111s first-video=1.058s FIRST-KEYFRAME=1.103s
    trial 2: session=0.082s config=0.134s first-video=0.201s FIRST-KEYFRAME=0.201s
    trial 3: session=0.066s config=0.122s first-video=0.183s FIRST-KEYFRAME=0.183s
    trial 4: session=0.054s config=0.116s first-video=0.168s FIRST-KEYFRAME=0.168s
    trial 5: session=0.070s config=0.168s first-video=0.272s FIRST-KEYFRAME=0.272s
    trial 6: session=0.092s config=0.131s first-video=0.204s FIRST-KEYFRAME=0.204s
    

    54–92 ms, median ~68 ms. I expected worse without AES acceleration. Handshake cost
    is not an argument against QUIC on this hardware.

    (Trial 1's 1.06s first-frame is a cold replay ring waiting on a real IDR, not
    transport. Warm joins are 168–272 ms.)

    4. Crypto without AES acceleration

    No openssl on the image, so I cross-compiled a benchmark against the same ring
    0.17 the QUIC stack actually uses
    , sealing 1200-byte packets (QUIC packet sized, not
    bulk):

    AEAD throughput per packet
    AES-128-GCM 72.8 Mbit/s 131.9 µs
    AES-256-GCM 58.7 Mbit/s 163.4 µs
    ChaCha20-Poly1305 166.0 Mbit/s 57.8 µs

    ChaCha20-Poly1305 is 2.3× AES-128-GCM on this chip. NEON carries ChaCha20; AES has
    nothing to stand on and falls back to constant-time software. This inverts the usual
    server-side ordering, where AES-NI makes AES-GCM the obvious default.

    The actionable part: in TLS 1.3 the server picks the suite. A camera talking to
    a cloud that prefers AES-GCM pays 2.3× for record protection and has no say in it.
    Anywhere OpenIPC controls both ends — Cloud especially — preferring
    ChaCha20-Poly1305 for these SoCs is free performance. I would expect the same to hold
    for SRTP in the WebRTC path, though I have not measured that.

    But scale it before acting on it: at 1.3 Mbit/s, AES-128-GCM costs ~1.8% of the
    core and ChaCha20 ~0.8%. Against a 20% transport total, crypto is not the
    bottleneck
    — the per-stream state machine work is, by roughly an order of magnitude.
    I went in assuming crypto would dominate on a chip with no AES. It does not, and that
    surprised me more than any other result here.

    5. Behaviour when the link degrades — two caveats first

    The image has no tc, so I could not use netem. I degraded the link physically
    instead: three cameras streaming 1080p simultaneously on one 40 MHz channel (RSSI
    −60/−63/−68, negotiated 54–150 Mbps). That is real contention, but it is airtime
    starvation, not the lossy-and-bufferbloated cellular case you actually care about.

    More important: what I can observe is my application's send-window-exhaustion
    signal, which sits downstream of quinn's congestion control. I have not instrumented
    quinn's cwnd/RTT directly
    , so I cannot tell you what Cubic is doing underneath —
    only what the application feels.

    With that said, the signal oscillates, reproducibly:

    19:54:12  uplink pressure   → throttle
    19:54:35  pressure cleared  → restore      (23s)
    19:54:38  uplink pressure   → throttle     (3s later)
    

    An earlier independent run gave the same shape with a 1.5s re-entry. So: it backs off
    promptly, holds while contended, and re-enters pressure within 1.5–3 s of clearing
    — because my clear condition is a 5-second quiet window and the link re-saturates
    almost immediately once the bitrate is restored. That is my hysteresis being too
    optimistic rather than obviously a CC pathology, but it is exactly the "does it
    oscillate" question, and on a contended link the answer is yes, on a 1.5–3 s cycle.

    Instrumented cwnd/RTT/loss traces off this hardware are a small patch to my agent if
    that would be more useful than my application's shadow of the controller — including
    against a real cellular modem rather than contended WiFi.

    Summary

    • Transport is ~1× the encoder on this class of chip, not a rounding error.
    • Per-stream cost dominates per-byte cost below ~1.35 Mbit/s. Protocol shape
      matters more than bitrate at camera bitrates; anything that opens a stream or
      channel per frame will hurt here.
    • Prefer ChaCha20-Poly1305 wherever OpenIPC controls the server side. 2.3× on the
      crypto line, free.
    • Memory is not the problem. 4.3 MB resident, 6.0 MB peak, stable under
      degradation, on a box with 34 MB.
    • The 64 MB on the box is 35,268 kB to Linux.

    The crypto benchmark is ~40 lines if it is worth having in-tree — that table for every
    SoC OpenIPC ships would be more useful than my one row of it.

  7. widgetii commented on Sep 1, 2026

    @widgetii
    Member

    This is the best data anyone has put on this class of hardware, and the caveats are
    the part that makes it usable — you flagged every place the measurement stops short of
    the claim, which is rarer than it should be. Some of it I could check against our own
    bench, so here is what corroborates, what changes, and the one experiment I think you
    should run before acting on your own headline.

    The per-stream result needs one more experiment before you redesign

    Your α/β solve is internally consistent — I checked the arithmetic on all of it and it
    holds, including the 1.35 Mbit/s crossover and the 1.8%/0.8% crypto scaling. And the
    strong form of the conclusion does not even need the fit: 87 kbit/s costs 4.1 points
    and 1100 kbit/s costs 7.6, so a 12.6× increase in bytes bought a 1.85× increase in
    cost. Whatever dominates at these bitrates, it is not the bytes. That much is solid.

    But "per stream" is not the only thing that varied. The 2×2 attributes every
    non-byte cost to stream count, and the experiment cannot separate QUIC stream
    lifecycle from everything else that happens once per access unit — RTSP demux, the
    copy out of the tap, your classifier's per-frame path, allocator churn. Those all scale
    with AU rate and none of them go away when you multiplex frames onto a long-lived
    stream. If they are the bulk of α, your redesign buys you the framing you write
    yourself and very little CPU.

    The clean experiment is one binary with two modes, same feed, same bytes, same frame
    rate, changing only the mapping: N frames → N streams, versus N frames → 1 stream with
    your own length prefix. If α collapses, your instinct was right and it is worth the
    rewrite. If it barely moves, the cost is per-AU and lives somewhere else entirely.
    That is a couple of hours and it converts a design intuition into a measurement.

    One input missing from the table, and it is the one the fit depends on: the actual
    AU rates for each row. Your constants imply roughly 24 streams/s for the substream
    viewer and 16/s for the 1080p one — the substream opening more streams per second
    than the main one is exactly what makes the system solvable rather than degenerate, so
    if those rates are not right the constants move a lot. Worth publishing the per-row
    frame rates alongside the bitrates.

    You can have the 2.3× today, without anyone else changing anything

    You are right that in TLS 1.3 the server picks the suite. But it picks from what the
    client offered, and the client controls that list — in rustls, restrict the provider's
    cipher_suites to TLS13_CHACHA20_POLY1305_SHA256 and no cloud that wants to talk to
    you can choose AES-GCM, because you never offered it.

    So for your own agent this is not a request to anyone; it is a two-line change and you
    keep the 2.3×. The general recommendation for OpenIPC Cloud stands and is well made,
    but you do not have to wait for it.

    It will not transfer to SRTP, and the reason is worth knowing

    DTLS-SRTP negotiates protection profiles from a fixed list — AES-128-CM with
    HMAC-SHA1 (RFC 5764) and the AES-GCM profiles (RFC 7714). There is no deployed
    ChaCha20 SRTP profile. Both ends must support whatever is chosen, so unlike TLS there
    is no list you can trim to force the fast cipher. The WebRTC path is stuck with AES on
    a chip that has no AES.

    What we do about it instead is the SoC crypto engine. Measured on gen-4 HiSilicon
    parts, per 1100-byte packet, software → engine:

    86.4 µs → 38.4 µs
    87.6 µs → 42.0 µs
    

    A bit over 2× — the same order as your ChaCha win, arrived at from the other
    direction. Note this lands close to your software AES-128-GCM figure of 131.9 µs per
    1200-byte packet, on a different vendor's silicon, which is a nice independent check
    that software AES on this generation is simply expensive.

    Two things that will interest you given your "scale it before acting on it" instinct.
    A gen-3 part measured slower through its engine than in software, so we do not build
    it there. And it is enabled on two boards only — Goke, including your GK7202V300, is
    not one of them. Your chip has neither AES acceleration nor an engine we use.

    Memory: reproduced on our bench, with one correction that matters

    Our lab GK7205V200 — same family as your board — reports MemTotal 35,232 kB
    against your 35,268. Within 36 kB, so your board and ours are configured the same way.

    That configuration is not the shipped default. osmem falls back to 32M when
    unset, which gives about 27 MB of userspace; you are on 40M, as our bench is. Anyone
    budgeting from your numbers on a stock image has roughly 8 MB less than you did, which
    makes your 4.3 MB resident a materially bigger slice of the box than it looks here.
    Your result is still the reassuring one — it just deserves that asterisk.

    An incidental echo: majestic on that same board sits at VmRSS 4,704 kB / VmSize
    15,000 kB. Your agent is 4.3–4.5 MB / 15.3 MB. Two unrelated processes landing on the
    same footprint is musl's arena behaviour showing through in both.

    Your ps warning is a good one and matches our experience — VmRSS from
    /proc/<pid>/status is the only number worth reading on these boards.

    On tc: confirmed absent, and here is the way around it

    Not an oversight you can work around on-box. CONFIG_TC is off in our busybox config,
    there is no iproute2 package in the tree, and the only thing shipped is busybox ip,
    which does addresses, links and routes but no qdiscs. There is nothing on the camera
    that can shape.

    The practical answer is to move the shaping off the camera: put it behind a Linux box
    acting as its AP or as a bridge, and run netem on that interface. You get real loss,
    delay and queue depth under your control, applied to the camera's actual uplink, and
    you can reproduce bufferbloat properly rather than borrowing airtime contention as a
    stand-in. Your contended-WiFi result is honest about being the wrong shape; a bridge
    in the middle gets you the right one for a few minutes of setup.

    And yes — instrumented cwnd/RTT/loss traces would be much more useful than the
    application's shadow of the controller
    , especially against a real modem. That is the
    gap between "my hysteresis is too optimistic" and "here is what Cubic does on a
    cellular link", and only the second one tells us anything we can act on.

    The benchmark: yes, and it is worth more than one row

    A per-SoC AEAD table is the most useful thing in your whole message, because it feeds a
    decision we currently make by hand. Hardware crypto is switched on for two boards
    because someone measured those two boards; every other SoC we ship is an assumption.
    A tool anyone can run, producing a row per chip, turns that into data.

    Probing and identification tooling lives in ipctool rather than here, and
    that is the natural home for something you run on a camera to find out what it is
    capable of — with the resulting table on the wiki. Worth opening the
    conversation there; point at this thread for the motivation.

    One finding of yours that is about us, not you

    19.5% → 21.9% for majestic to serve a loopback RTSP tap with no network client is
    the number I did not expect, and it is the one I am taking away from this. 2.4 points
    of a single core to move frames to a local consumer is worth understanding on its own
    terms, independently of anything QUIC.

    Thank you for measuring all of this properly. Very little of what reaches us has error
    bars on it.

  8. bneigher commented on Sep 1, 2026

    @bneigher
    ContributorAuthor

    Three of these I could act on immediately, and one of them turned into a correction
    to your ChaCha suggestion that I think is worth having on the record.

    The missing input: per-row AU rates

    You are right that the fit depends on these and I should have published them. Measured
    by the client, over the same 26 s windows:

    row wlan0 TX streams/s of which video audio
    + 1 LAN viewer, substream (video1, 12 fps cap) 192 kbit/s 23.8 8.2 ~15.6
    + 1 LAN viewer, 1080p (video0, 25 fps) 1292 kbit/s 40.0 24.6 ~15.4

    One correction: you inferred ~24/s for the substream and ~16/s for the 1080p viewer,
    and read the substream opening more streams than the main one as the thing that makes
    the system solvable. It is the other way round — 23.8 vs 40.0, the substream opens
    fewer. Solving with the measured rates gives α = 0.156 %core per stream/s and
    β = 0.0046 %core per kbit/s, which is what I quoted; your arithmetic check on the
    constants was right, the rates behind them just were not stated.

    Note the audio stream rate is nearly identical in both rows (AAC at 1024 samples /
    16 kHz ≈ 15.6/s regardless of video), so almost all of the stream-rate delta is video
    frames. That is convenient for the experiment you propose: it isolates cleanly.

    Your experiment is the right one and I have not run it yet

    I am not going to pretend otherwise. Your objection is correct and it is the one that
    matters: the 2×2 attributes everything that happens once per access unit to "stream
    count", and RTSP demux, the copy out of the tap, the classifier's per-frame path and
    allocator churn all scale with AU rate and survive multiplexing. If those dominate α,
    the rewrite buys framing I have to write myself and no CPU.

    One binary, two modes, same feed, N→N vs N→1 with my own length prefix, is exactly the
    right shape and I will run it. I would rather report that honestly than let the
    "per-stream" framing harden into a design decision on the strength of an experiment
    that cannot separate the two. I will post the result either way — including if it says
    my instinct was wrong, which is the outcome that would save me the rewrite.

    The ChaCha20 change does not work on QUIC, and the reason is structural

    This is the one I want to flag, because your reasoning is right for TLS and I
    implemented it verbatim before finding out it does not carry over.

    Restricting the rustls provider to TLS13_CHACHA20_POLY1305_SHA256 does not fail to
    negotiate — the endpoint refuses to construct at all:

    panicked at wtransport-0.7.2/src/config.rs:1085:
    CipherSuite::TLS13_AES_128_GCM_SHA256 missing: NoInitialCipherSuite { specific: false }
    

    QUIC derives Initial packet protection keys per RFC 9001 §5.2, and that is hardwired to
    AES-128-GCM — before any negotiation happens, and independently of what the handshake
    later selects. So a QUIC client cannot drop AES from its provider the way a TLS client
    can. The suite list has to keep AES-128-GCM in it for the transport to exist.

    What is available is preference: offer ChaCha20 first and AES-128-GCM second. rustls
    sends them in that order and most servers honour client order, but the server still
    picks, so it is a preference rather than the unilateral guarantee. I have shipped that
    ordering. I could not measure it: at the publish bitrate the crypto line is ~0.14 % of
    a core, an order of magnitude under my ±0.2 measurement noise, so agent CPU read
    8.4/8.2/8.2 % before and after. The 1.8 % vs 0.8 % figures only separate at 1.3 Mbit/s.

    Two consequences worth drawing out, since they cut against the optimistic reading of my
    own crypto table:

    • The chip has no AES hardware, no crypto engine you build for, and cannot decline
      AES for QUIC Initials. Some AES on this silicon is unavoidable.
    • It is a small fraction of the packets — Initials only — so the practical cost is
      negligible. But "just force ChaCha20" is not available to a QUIC endpoint, and I
      would not want that repeated as advice without the caveat.

    Accepted without reservation

    SRTP. You are right and I was speculating past my evidence. DTLS-SRTP negotiates
    from a fixed profile list with no deployed ChaCha20 option, and unlike TLS there is no
    list to trim. I should not have extended a TLS-shaped conclusion to it. Your engine
    numbers (86.4 → 38.4 µs, 87.6 → 42.0 µs per 1100 B) landing near my 131.9 µs/1200 B
    software AES figure on different silicon is a good independent check, and the gen-3
    part being slower through its engine is exactly the sort of result that only shows up
    if someone measures instead of assuming.

    The osmem asterisk is important and I should have caught it. Confirmed on my
    board — bootargs=mem=40M …, which is where 35,268 kB comes from. Anyone budgeting
    from my numbers on a stock 32M image has ~8 MB less than I did, and my 4.3 MB resident
    is a materially larger slice of that box. The reassuring conclusion survives, but it
    deserves the caveat and the caveat is yours, not mine.

    Your GK7205V200 reporting 35,232 kB against my 35,268 is a nice confirmation that the
    two benches are configured the same.

    tc. Thank you for confirming it is absent by construction rather than by
    oversight. Shaping from a Linux box acting as AP or bridge is obviously the right
    answer and I will set that up — it gets me real loss, delay and queue depth on the
    camera's actual uplink instead of borrowing airtime contention as a stand-in, which I
    knew was the wrong shape and said so mainly because I had nothing better.

    That plus instrumented cwnd/RTT/loss is the combination worth doing properly, and I
    agree it is the only version that says anything actionable about Cubic on a cellular
    link. Currently I can only report what the application feels.

    The benchmark, and ipctool

    Agreed that a per-SoC table beats one row, and that "hardware crypto is on for two
    boards because someone measured those two boards" is exactly the kind of thing a
    runnable tool fixes. I will open the conversation on ipctool and point at this thread.

    The benchmark is ~40 lines against ring, sealing packet-sized buffers rather than bulk,
    which matters — bulk numbers flatter AES on chips where the per-call overhead dominates.
    Happy to shape it to whatever ipctool wants for probe output.

    On majestic's 2.4 points

    Glad that one is useful, and I would not over-read it from a single board yet. It is
    19.5 % → 21.9 % of one core measured as majestic's own utime+stime across a 25 s
    window, with the agent stopped versus running and no network client attached — so
    it is the cost of serving one loopback RTSP consumer, nothing more. n=1 board, one
    encoder config (H.265 640×360@12 + AAC). If it reproduces on your bench I would trust
    it; if it does not, mine is the one to doubt.

  9. bneigher commented on Sep 1, 2026

    @bneigher
    ContributorAuthor

    Ran it. You were right and my headline was wrong: multiplexing recovers about a
    tenth of what I attributed to stream count.
    Posting the number that kills my own
    conclusion, as promised.

    Method

    One binary, two modes, selected by the connect URL so nothing else differs:

    • A — one uni stream per access unit (what ships)
    • B — ?mux=1, everything on one long-lived uni stream, [u32 BE len][payload]

    Same feed, same encoder, same CONFIG, same replay ring, same broadcast receiver, same
    frames in the same order. Smoke test before measuring, to confirm the arms were
    actually comparable rather than assuming it:

    A: held 10s: 406 streams (249 video)
    B: held 10s MUXED: 406 frames (249 video) on 1 stream
    

    Identical counts. Then 11 interleaved 22-second windows on the GK7202V300, video0 at
    25 fps, agent CPU from utime+stime in /proc/<pid>/stat.

    Result

    A (stream per AU)   23.62% +/- 0.76 of core    2079 kbit/s    597 video frames/window
    B (one stream)      23.55% +/- 0.94 of core    2182 kbit/s    605 video frames/window
    

    Scene motion moved the bitrate between runs, so regressing CPU on bitrate and arm
    rather than comparing means directly:

    multiplexing removes 0.70 points (SE 0.25, t = −2.75, 95% CI [−1.19, −0.20]) out of
    ~23.6. Small, but not zero.

    Against that, the 2×2 attributed 6.95 points to "per stream" at ~45 streams/s
    (α = 0.156 × 45). So:

    • QUIC stream lifecycle is ~10% of what I called per-stream cost
    • ~90% is per-access-unit work that survives multiplexing — RTSP demux, the copy
      out of the tap, allocator churn, the broadcast hop
    • even the optimistic end of the CI only recovers 17%

    Your framing was exactly right: the experiment could not separate the two, and it turns
    out to be almost entirely the half that multiplexing does not touch.

    So I am not rewriting the wire format. It would buy ~3% of one core in exchange for
    hand-rolled framing and losing the drop-on-saturation policy that keeps a slow reader
    from stalling the session. One-stream-per-frame stays.

    Two honest caveats about arm B

    It is not a clean isolation of QUIC stream machinery. Removing the per-AU stream
    also removes the per-frame tokio::spawn and the bounded-inflight drop policy — a
    length-framed stream desyncs if you skip an item, so B blocks where A drops. Those three
    travel together; that IS "one stream" as anyone would build it. So the measurement
    answers the design question, not "what does open_uni cost in isolation". Since the
    answer came back negative, the distinction does not change the decision.

    One A run is excluded. It delivered 797 streams against ~974 for every other run,
    with two gaps >500 ms — it dropped frames, so it was doing less work and its lower CPU
    was not comparable. Worth flagging that excluding it cuts against my conclusion: it
    was the cheapest A run, and dropping it makes A look more expensive, not less.

    The AU rates, since they were the missing input

    Both arms, per 22 s window: ~600 video frames plus AAC at ~15.6/s, so ~45 streams/s in
    arm A. Arm B opens one stream for the whole session.

    What I take from this

    A two-point solve names a correlate, not a mechanism. The constants were arithmetically
    fine — you checked them and so did I — and the strong form held too: 12.6× the bytes for
    1.85× the cost really does mean bytes are not what dominates at these bitrates. What was
    wrong was the leap from "not bytes" to "therefore stream lifecycle", when "per access
    unit, wherever that work lives" was the honest statement of the same data.

    That is the second time in a week I have published a number that was real and an
    attribution that was not — the other was gating our build on sizes.<plat>.json's
    uimage_bytes (1,917,812) when the uImage in the nor-lite tarball is 1,883,172. That one
    nearly bought a fleet-wide repartition of boards whose recovery path I know to be
    unreliable. Same shape: a proxy that nearly works produces a confident wrong answer
    instead of an obvious failure.

    The harness stays in the tree behind ?mux=1, off by default, so this is re-runnable in
    minutes against any future claim about per-stream cost — including on other SoCs, if that
    is ever useful to you.

    Thank you for pushing on it rather than taking the table at face value. This is the most
    useful correction I have had in a while, and it saved a rewrite.

  10. bneigher commented on Sep 2, 2026

    @bneigher
    ContributorAuthor

    Instrumented the controller as you suggested. It corrects what I told you: the
    oscillation I reported was not the application's shadow of Cubic — Cubic was doing
    nothing at all.
    And it turned out to be a real bug in our QoS, now fixed.

    What I added

    wtransport exposes the inner quinn connection behind a quinn feature, so
    Connection::quic_connection().stats().path gives real PathStats: cwnd, smoothed
    rtt, congestion_events, lost_packets, current_mtu, black_holes_detected.
    Sampled once a second, logged as per-interval deltas — cumulative totals integrate
    away exactly the back-off/recovery behaviour worth seeing.

    Steady state: Cubic never leaves its initial window

    Cloud publish, H.265 640×360@12 + AAC, ~74 kbit/s of media, ~44 pkt/s:

    rtt=27-35ms  cwnd=12000  mtu=1452  lost=0  loss=0.00%  cong_events=0  blackholes=0
    

    cwnd is pinned at 12000 = 10 × 1200 = quinn's initial window, and never moves in
    either direction. At this bitrate the connection is application-limited, never
    cwnd-limited
    , so Cubic has no reason to grow the window and no loss to shrink it.

    That is worth stating plainly, because it undercuts the premise of my own offer:
    in normal operation on this hardware, the congestion controller is untested. Any
    claim about Cubic's behaviour here — mine included — is currently unsupported.

    The saturation run, and the bug it exposed

    I then saturated the radio (three cameras streaming 1080p on one 40 MHz channel) so our
    QoS would fire, expecting to catch the controller reacting:

    QoS: uplink pressure -> video0 600kbps
      ... 10 consecutive samples, ALL: cwnd=12000 lost=0 cong_events=0, rtt flat 28-35ms
    QoS: pressure cleared -> 1200kbps   23:46:20.641
    QoS: uplink pressure -> 600kbps     23:46:21.651   <- one second later
    

    The path was completely healthy the entire time. No loss, no congestion event, no
    RTT inflation, cwnd untouched.

    The cause was ours. That pressure signal fired on sem.available_permits() == 0 — our
    sender's own 8-permit concurrency semaphore — not on anything the network reported. So
    we were halving the local video bitrate because our own bounded window filled, most
    plausibly from CPU contention holding permits longer while the agent was at ~23% of the
    single core. I have not proven that causation; what is proven is that the link was not
    congested when the signal fired.

    So my "my hysteresis is too optimistic" was wrong twice over: the hysteresis was fine,
    and it was watching a signal with no relationship to congestion.

    Fixed

    Pressure now comes only from evidence the transport actually reports — any of:

    • congestion_events increased (Cubic itself reacted)
    • packets lost in the interval
    • smoothed RTT > 2× the minimum observed RTT — queue building ahead of us, which on
      a buffered link shows up before loss. min-RTT is learned from observation rather than
      configured, so it adapts to whatever path a given camera has.

    Re-ran the identical saturation test after the change: zero pressure events, zero
    throttling
    , RTT 26.5–30.5 ms against a tracked minimum of 26.0. The permit-count
    trigger is deleted, with a comment explaining what the trace showed so nobody restores it.

    What this means for the question you actually asked

    I still cannot tell you what Cubic does on a bad link, and now I can show why rather
    than assert it: we never stress the controller, so there is nothing to observe. The
    contended-WiFi stand-in does not stress QUIC at all — it stresses our CPU.

    That needs a genuinely degraded path, which means your suggestion: netem on a Linux box
    acting as AP or bridge, since tc is absent by construction on these images. The
    instrumentation is in place and ships dark behind a flag file, so when I have that
    bridge the traces come for free — cwnd, RTT, loss and congestion events per second
    through a controlled loss/delay/queue profile, and against a real modem after that.

    Two smaller things from the same trace, in case they are useful: PLPMTUD settles at
    1452 on this path, and black_holes_detected stayed 0 across every run.

  11. bneigher commented on Sep 4, 2026

    @bneigher
    ContributorAuthor

    Got a shaper in front of a camera and ran the traces. Cubic numbers from this hardware
    below
    , plus the bug the experiment found in my own code — which is the more useful half.

    Setup, and its honest limits

    No Linux bridge to hand, so I used what was: macOS dnctl/pfctl dummynet on the
    receiving machine, shaping the camera's LAN WebTransport path. The camera is the
    sender there, so its quinn/Cubic drives a link whose loss and delay I control.

    Three caveats, all of which narrow what this proves:

    • It is the LAN path, not the cloud uplink. Same quinn, same Cubic, same silicon —
      different connection.
    • dummynet loss is uniform-random. Real wireless and cellular loss is bursty, and
      loss-based controllers behave differently under bursts.
    • Still not a cellular modem. Nothing here reproduces RAN scheduling or handover.

    Instrumentation is quinn's own PathStats sampled every 2 s, counters as per-interval
    deltas.

    Baseline, and why the cloud path told us nothing

    cloud publish (~74 kbit/s) LAN viewer (~1.5–2 Mbit/s), unshaped
    cwnd 12,000, never moves 80,564 → 237,380
    RTT 27–35 ms 10.8–40.8 ms
    congestion events 0 0 (in a clean 68 s run)

    The cloud figure is the one I reported before: cwnd pinned at quinn's initial window
    because at 74 kbit/s we are application-limited and never fill it. On the LAN path the
    same stack grows to 237,380 and parks — still no loss, but now demonstrably cwnd-active.

    2% loss + 50 ms

    cwnd   min 6,339   max 22,768   mean 13,616      (vs ~237,000 unshaped)
    rtt    63.4 – 96.9 ms, mean 76.4
    loss   34/34 intervals, mean 2.59%, peak 6.20%
    cong   166 events over 68 s  (~2.4/s)
    

    cwnd collapses ~17× and stays in a continuous sawtooth: 20,776 → 17,347 → 9,334 →
    9,614 → 13,011 → 15,591 → 14,269 → 9,825 → 12,179 → 15,245. Multiplicative decrease on
    loss, cubic regrowth, repeatedly, with no sign of collapse to a floor or of oscillating
    pathologically. Measured loss runs above the configured 2% because retransmissions are
    dropped at the same rate.

    Application-level outcome: the viewer still received 1,236 video frames in 70 s (~17.6
    fps against 25 nominal) with 5 stalls >500 ms. Degraded, session intact.

    200 ms delay, plr 0 — the bufferbloat case

    I expected Cubic to ignore the latency and fill the buffer. It half did:

    cwnd   min 26,974  max 85,896  mean 49,923
    rtt    217.5 – 264.2 ms, mean 231.5
    loss   11/34 intervals, mean 0.61%   <-- with plr 0
    cong   17 events
    

    cwnd settled ~3.7× higher than the lossy case but well below the unshaped 237,000 — and
    not because it saw the delay. With plr 0 there was still 0.61% loss, because the
    pipe's finite queue overflows once 200 ms of BDP is in flight. Every reaction traces to
    that loss. Cubic did not respond to latency; it responded to the drops that latency
    eventually caused.

    Which is the expected result for a loss-based controller, but worth having measured: on
    a path that buffers deeply enough to delay without dropping, this stack has no signal at
    all until the buffer finally overflows. That is the argument for delay-sensitivity on
    cellular, made with numbers rather than assertion.

    The bug this found in my own code

    Worth more than the traces. I had added a bufferbloat trigger to our QoS: raise pressure
    when smoothed RTT exceeds 2× the minimum observed. The minimum was all-time,
    per-connection.

    A session that connects through an already-bloated path learns its minimum from that
    path
    :

    lan-cc rtt=249.3ms min=249.3ms    <- connected through 200ms of shaping
    lan-cc rtt=226.1ms min=219.2ms    <- ratio 1.03x; the 2x test can never fire
    

    So the trigger was dead precisely on the paths it existed for — a camera on a
    permanently bloated uplink would report nothing wrong, forever. Fixed by making the
    minimum a 300 s rolling window (gated on ≥30 samples, so a session's first seconds
    do not make every path look clean against itself). Same connection, shaper removed
    mid-session:

    lan-cc id=5 rtt=226.1ms min=219.2ms   <- before
            ── shaping removed, no reconnect ──
    lan-cc id=5 rtt=16.5ms  min=16.5ms    <- re-baselined
    

    Threshold goes from a meaningless 438 ms back to 33 ms.

    The residual limit cannot be fixed this way and I am not going to pretend otherwise:
    a path bloated for the entire window is indistinguishable from a genuinely distant relay
    using RTT alone. This detects bloat that develops, not bloat that was always there. I
    considered an absolute ceiling and rejected it — 250 ms is pathological on my LAN and
    normal on satellite, so a fixed threshold relocates the wrong answer rather than removing
    it. If there is a standard trick for this I would rather hear it than invent one.

    Also worth stating: across every run here, shaped or not, every congestion event came
    from loss
    . The RTT arm of that trigger is correctly baselined now but has still never
    fired in anger on a real path.

    What would make this actually answer your question

    A real modem. What I have says "here is Cubic on this silicon under synthetic uniform
    loss on a LAN". Cellular adds burst loss, RAN scheduling delay and handovers, and I would
    not extrapolate from these numbers to that. If a cellular-attached board is something
    OpenIPC has access to, the instrumentation is a flag file away from producing the same
    traces there, and I am happy to run whatever profile is useful.

  12. widgetii commented on Oct 4, 2026

    @widgetii
    Member

    Closing: the question of whether WebRTC belongs in lite was answered (yes, by design), with measured costs. The QUIC/WebTransport work is your own project. Thanks for publishing the data.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions