This level assumes you finished 101. It uses the words stream, datagram, track, relay, and publisher/subscriber without re-defining them.
Level 101 taught you four general-purpose ideas: QUIC, MoQ, WebTransport, and datagrams. This level shows you what TeleQuick actually builds with them.

1. The relay, for real

In level 101, “the relay” was a diagram box. Here is what it actually is. Our relay is a native QUIC server. It speaks MoQT to publishers and subscribers, running as a native client and server process rather than a browser-only stack. A relay holds open QUIC connections from every publisher and subscriber currently using it, and for every track it forwards, it copies each object from the publisher’s connection onto every subscriber’s connection. One relay cannot hold every connection in the world, and one that could would still be in the wrong place for most of its subscribers. We run many relays, spread across regions (“PoPs” — points of presence), and steer each connection to a nearby one, so two people talking to each other from the same city do not pay round-trip latency to a data center on a different continent. When a publisher and its subscribers land on different regions, the relays forward objects between themselves — a subscriber never has to know or care which region its publisher connected to. relays in multiple regions forwarding objects between themselves None of this changes anything about how you use a track — you still publish and subscribe the same way. It explains why the relay layer exists as its own piece of infrastructure, instead of every publisher connecting directly to every subscriber (which section 5 of level 101 already told you does not scale). Congestion control, in practice. Level 101 section 4 told you RFC 9002 leaves the sending-rate algorithm pluggable, and that the MoQ working group has written down real caveats about which behavior suits live media (bufferbloat, application-limited sending, throughput-optimized algorithms trading smoothness for speed). This is not a footnote for us — it is a live decision our relay’s transport layer has to make on every connection, and it is why “just turn on the newest congestion control algorithm everywhere” is not automatically the right call for a relay carrying latency-sensitive tracks. A relay tuned for maximum throughput and a relay tuned to avoid bufferbloat on a live call are not the same relay.

2. What a subscribe actually looks like on the wire

Level 101 gave you MoQ’s vocabulary — track, group, subgroup, object, publisher, subscriber, relay. Here is how those words turn into actual messages, using the real names from the MoQ Transport specification (as of draft-18; the protocol is still evolving, so treat exact message names as “currently this,” not “forever this”).
  1. SETUP. The very first thing either side sends, on a dedicated control stream, right after the QUIC/WebTransport connection opens. It negotiates the MoQT version both sides will speak. Nothing else is valid before this exchange completes.
  2. PUBLISH_NAMESPACE. A publisher tells the relay it has one or more tracks available under a namespace, before anyone has asked for them. Think of it as “here’s what I have,” separate from “someone actually wants it.”
  3. SUBSCRIBE. A subscriber asks the relay for a specific track by name.
  4. SUBSCRIBE_OK. The relay’s (or publisher’s) acceptance, at which point the subscription is live and objects start flowing.
  5. Objects, on subgroup streams. Once a subscription is live, objects arrive on one or more QUIC streams, each one opened with a small subgroup header naming the track, group, and subgroup it belongs to. A relay that needs to shed load under congestion can drop or deprioritize one subgroup’s stream — say, a high-detail temporal layer — without touching the others, because MoQT deliberately keeps a group’s subgroups on independent streams for exactly this reason.
A subscriber never has to construct any of this by hand — your SDK’s subscribe() call does all five steps for you. Knowing the real shape underneath is what turns “the relay is doing something weird” into “I can tell you which message exchange to look at.” SETUP, PUBLISH_NAMESPACE, SUBSCRIBE, SUBSCRIBE_OK, then objects, between publisher, relay, and subscriber

3. Streaming

The Streams modality is the most direct use of everything in level 101: one publisher’s audio and video, fanned out live to any number of subscribers, through the relay. The metric that matters most for streaming is glass-to-glass latency: the time from a photon hitting the publisher’s camera lens to a pixel lighting up on a subscriber’s screen. Every choice in the pipeline — encoding, the network hop to the relay, the relay’s own forwarding delay, the network hop to the subscriber, decoding, rendering — adds to that number. Traditional CDN-based live streaming (think: video segmented into a few seconds of chunks, fetched over HTTP) accepts several seconds of glass-to-glass latency in exchange for using ordinary web caches. A MoQT relay forwards each object as it arrives, so the achievable latency is much closer to the raw network round trip — typically well under a second, often under 100 ms within one region. This is also where the stream/datagram choice from level 101 shows up concretely: a video track’s frames are usually carried as short QUIC streams (one per frame or per small group of frames), because a frame is either fully useful or not worth having at all, while very time-sensitive, high-frequency signals elsewhere in the stack (see robotics, next) lean toward datagrams. It is also where MoQ’s Publisher Priority (level 101, section 5) earns its keep. A publisher sending multiple quality layers of the same content marks its lower layers with a less urgent priority. Under real congestion, the relay favors the high-priority layer and lets the lower one fall behind first — a viewer’s stream degrades gracefully, in a controlled order the publisher chose, instead of every layer degrading at once and unpredictably.

4. Robotics and teleoperation

Teleoperation means a human, or another computer, controls a robot from somewhere else — sometimes across the room, sometimes across the world. It is the application area where the difference between a stream and a datagram stops being academic and starts being a safety property. We split a teleoperated robot’s traffic into two planes: The control plane carries commands: “move joint 3 to this angle,” “stop,” “switch mode.” These commands can require reliable delivery and application acknowledgments. Reliable QUIC streams help recover lost bytes, but cannot guarantee a delivery deadline. A safe control design also needs command expiry and defined behavior when communication fails. The data plane carries continuous telemetry: joint positions, force sensor readings, video from an onboard camera. A single dropped joint reading, correctly followed by the NEXT reading a few milliseconds later, is harmless — resending the old one would only make the control loop react to stale information. Data-plane traffic favors datagrams, for the exact reason level 101 section 5 explained. This split also shapes how we think about failure. A control connection going quiet is treated very differently from a data connection going quiet: the robot needs a locally enforced response when control input expires. The appropriate safe state depends on the machine and its operating conditions. A network-delivered stop command alone does not establish that safety behavior. We will not go deeper into the exact safety mechanisms here — that is a topic for the team actually building robotics products, not this course — but you now have the vocabulary (control plane vs. data plane, streams vs. datagrams) to understand why the split exists when you see it in real code or diagrams.

5. Voice inferencing

A voice AI agent takes a caller’s speech, decides what to say, and speaks back — often with well under a second of turnaround, in the middle of a live phone call. Two architectures do this: Cascaded. Speech-to-text (ASR) turns audio into text. A language model (LLM) decides what to say, as text. Text-to-speech (TTS) turns that back into audio. Three separate stages, three separate services, chained together. This is easier to build, easier to swap one stage for a better model later, and easier to debug (you can look at the text in the middle). Its downside is that each stage adds its own latency, and none of the stages can hear things text cannot represent well — tone of voice, a laugh, someone talking over the caller. Realtime speech-to-speech. One model takes audio in and produces audio out directly, without a text bottleneck in the middle. This can react faster and can preserve more of the nuance in someone’s voice, at the cost of being newer, harder to steer precisely, and harder to debug. Either architecture needs somewhere to run the model — usually a GPU, sometimes far from the caller. This is exactly the WAN-latency problem level 101 introduced: the transport carrying audio to and from the inference server has to be fast and to tolerate loss gracefully, which is why this leg of the system runs over the same QUIC/MoQT foundation as streaming and robotics, not a separate ad hoc protocol.

6. What ties all three together

Streaming, robotics, and voice inferencing look like different products. Underneath, they are the same three ingredients, in different proportions: You now have the vocabulary to read a system diagram for any modality we ship and know which parts are relay, which are transport-level choices, and why.

Where to read more

This course does not repeat API reference. Once you have the concepts, go to the docs for exact method names and working code:

Exercise (30 minutes)

  1. Pick one modality from the table in section 6 that you have not worked on before. Open its docs page above and find one concrete detail (a method name, a message type, a config option) that maps to something you read in this level. Write one sentence connecting the two.
  2. A teammate asks: “Why don’t we just use one relay in one data center for everyone?” Answer using section 1: one reason is about how many connections a single relay can hold at all, and a separate reason is about physical distance — explain both, and why either constraint can independently justify additional relays.
  3. Walk through the five-message exchange from section 2 for a viewer joining an already-live stream. Which message accepts the subscription, and how does actual object delivery differ from that acceptance?
  4. Explain, in two sentences, why a robot’s emergency-stop command and its continuous joint-position telemetry should not travel the same way over the network.
Compare your reasoning with the worked answers. Complete the streaming walkthrough before moving to level 103.