This level assumes you finished 101 and 102. It is the most advanced level in this course.
This level has three parts. Part 1 is security. Part 2 is programming against our SDK. Part 3 is a method for solving problems when something does not work.

Part 1 — Network security for real-time systems

Session identity and operation permissions

A persistent session can establish an application identity during setup. The transport handshake, application authentication, and operation authorization are separate responsibilities. TLS protects a connection; it does not automatically grant permission to publish every track. The relay must check whether that identity may perform the requested operation. Session setup errors and denied publish or subscribe operations can therefore occur at different stages. Use the authentication documentation for the deployed mechanism.

TLS is not optional, and certificates are per-tenant

Every QUIC connection in our stack negotiates TLS 1.3 as part of the QUIC handshake itself — there is no “insecure QUIC” mode to fall back to. We serve many tenants (customers) and, for the two brands, two public domains, off the same relay fleet. That works because the relay presents the right certificate for whichever hostname the client asked for, using TLS’s SNI (Server Name Indication) extension — the same mechanism your browser uses to let one IP address serve many HTTPS websites, each with its own certificate.

Credentials are scoped, not universal

We do not issue one all-powerful API key per customer. Credential prefixes identify intended credential categories. They do not prove the permissions that the server grants. Check the credential’s actual scope and the authorization result when an operation fails. Two habits follow directly from this table. First, never treat a publishable credential (cwk_) as secret, and never treat any other prefix as safe to publish — a widget key embedded in a web page is expected and fine; a voice key (cvk_) in a web page is an incident. Second, when you add a new capability, give it its own prefix and its own scope rather than folding it into an existing broad one — a grab-bag credential that covers five unrelated capabilities means a bug or a leak in any one of the five compromises all five.

Failure modes worth knowing, in general

These are not specific to us — they show up in any real-time, multi-tenant system, and knowing them will make you a better reviewer of any code you touch here:
  • Open redirect. Any URL parameter that decides where a user’s browser goes next (returnTo, next, redirect_uri, …) must be checked against an allow-list before you use it. Otherwise, an attacker crafts a link that looks like it points at us but actually sends a signed-in user’s browser somewhere else, right after what looked like a legitimate sign-in. If you ever add a redirect parameter, validate it against your own domain first.
  • Replay. A message that was valid once (a command, a signed request) must not be valid forever just because it is well-formed. Time-bound it, or make it single-use, or both.
  • Unbounded resource creation. QUIC lets one connection open many streams cheaply. “Cheaply” for a legitimate client is also “cheaply” for an attacker — a connection that opens streams without limit can exhaust a relay’s memory. Any place that creates a resource per request (a stream, a subscription, a session) needs a cap. QUIC itself does not leave this entirely to application code. MAX_STREAMS limits the number of streams a peer can open. MAX_DATA and MAX_STREAM_DATA provide separate byte-based flow control. Application resources, such as subscriptions and tenant state, still need their own limits.
  • Trusting the client’s view of state. A client-side check (“this button is disabled because you are not entitled to this feature”) is a UX nicety, not a security control. The server must enforce every entitlement and permission check again, independently, because a client can always be modified or bypassed.
  • Congestion-signal manipulation. An unauthenticated peer influences your sender’s behavior more than it might seem: RFC 9002’s own security considerations (§8.1, §8.3) note that a receiver can misreport loss, delay, or ECN markings to force a sender to slow down — a denial-of- service angle the spec acknowledges as a residual risk rather than something it fully closes. Worth knowing if you are ever debugging “why did this connection throttle itself for no visible reason” — the answer is sometimes hostile, not just a bad network.
None of these are unique to TeleQuick — they are worth knowing on any real-time or multi-tenant system you ever work on. Two of the specific defenses in the packet and handshake lesson (the 1200-byte Initial padding requirement and the three-times anti-amplification limit) are a template worth copying whenever you design a new “someone can talk to us before we know who they are” surface: assume the address is unverified, and cap what you are willing to spend on it until it proves otherwise.

Part 2 — Programming against the SDK

A typical client opens a session, publishes or subscribes, exchanges data, and closes its resources. Exact methods differ by SDK and modality. For the Python data modality, the actual API uses Data.publish and Data.subscribe. Work through the publisher/subscriber lab. It provides a complete program, a local encoding check, and a two-terminal relay exercise. The program also demonstrates resource cleanup and explains which output proves delivery. Use the Python installation guide to prepare the supported SDK and native library. The data lab uses telequick-sdk, imported as telequick. For complete signatures and supported configurations, use the SDK reference and realtime tracks documentation.

Part 3 — A method for solving problems

New engineers often debug real-time systems by guessing: change something, run it again, see if it helps. This wastes time, because a real-time bug often has a delay between cause and symptom. Use this order instead. Each step rules out a whole category before you look inside it.

Step 1 — Is the session even connecting?

Before anything else, confirm the QUIC/MoQT session itself is up. A connect failure and a “connected but nothing arrives” failure have completely different causes, and treating them as the same problem wastes your first ten minutes.
  • Check for a connect error in your SDK’s logs first — a credential problem, a network path problem (firewalled UDP/443, common on locked-down corporate networks — QUIC needs UDP, not just TCP), or a TLS/certificate problem will surface here.
  • If connect succeeds, confirm which relay/PoP you landed on. “It works from my laptop but not from the office” is very often a network path problem, not a code problem.

Step 2 — Is the credential scoped for what you are trying to do?

Use the prefix table to identify the intended credential category. Then check the actual tenant, scope, expiry, and authorization result for the failing operation. A prefix alone does not prove permission. Check this before investigating forwarding behavior.

Step 3 — Isolate publisher, relay, and subscriber

A track has three parties. A symptom at the subscriber (“no video”) can originate at any of the three:
  • Publisher problem: is it actually sending objects? Log at the publish call site, not just at the network layer.
  • Relay problem: is the relay actually forwarding? This is where cross-region fan-out (level 102) or an over-scoped/incorrectly-scoped subscription can silently produce zero delivery even though the publisher is sending correctly.
  • Subscriber problem: is it subscribed to the right track name and namespace? A typo in a track name looks identical to “nothing is being published” from the subscriber’s side — you get silence either way, so do not assume the more interesting-sounding cause.
Change one variable at a time here. If you swap in a different client and the problem disappears, you have isolated it to the client you swapped out — not “fixed” anything yet, but you now know where to keep looking.

Step 4 — Check the observability surface before you add your own logging

Before you sprinkle print statements everywhere, check whether the platform already has a dashboard for this. Traffic accounting, tenant logs, and QUIC-transport metrics all exist specifically to answer “is data moving, and how much”: Platform observability · Tenant logs · QUIC transport metrics

Step 5 — Write down what you ruled out

Once you find the cause, write one line for whoever hits this next: what the symptom was, and which step actually found it. Most repeated debugging time is spent re-discovering something a teammate already ruled out last month.

Capstone exercise (45 minutes)

A teammate reports: “My subscriber connects fine, and the docs example worked yesterday. Today I get zero objects on my track, no error.” Using the five steps in Part 3, in order, write down:
  1. What you would check first, and why that step comes before the others.
  2. Two different root causes that would both produce exactly this symptom (“connects fine, zero objects, no error”) — one at the publisher side and one at the subscriber side.
  3. Which dashboard from Step 4 you would open to tell those two causes apart, without needing either engineer to add new logging.
Compare your answer with the worked solution. Then complete the debugging lab to observe failures and recovery.

You finished the course

Check that you can explain a transport choice, run the local SDK exercise, and distinguish the three debugging scenarios. Record live relay completion separately from written or local exercises. From here, the docs site is your day-to-day reference. Come back to this course whenever you need to re-explain one of these ideas to the next new hire.