How to Build a Satellite Interference Hunting Runbook That Works at 3 a.m.

By | Friday, July 17, 2026

How to Build a Satellite Interference Hunting Runbook That Works at 3 a.m.

The call comes at 02:47. A carrier on transponder 7B is down 3.2 dB from its nominal C/N, and the NOC is seeing elevated BER on two VSAT remotes in the eastern beam. The on-call technician has been with the team for four months. He’s standing in front of a spectrum analyzer looking at a noisy passband, and he has no idea what he’s looking at. There’s a binder on the shelf labeled “Interference Procedures,” but it’s a collection of vendor manuals and a photocopied flowchart from 2019 that references equipment you decommissioned two years ago.

This is the scenario that separates a real runbook from a document someone wrote once and forgot. The difference between a 20-minute resolution and a 4-hour escalation isn’t the technician’s skill. It’s whether the runbook in front of him has a decision tree that matches the equipment he’s actually touching, annotation conventions he can follow without guessing, and checkpoints that tell him exactly when to call you.

Why Most Interference Runbooks Fail

Most runbooks I’ve seen in satellite ground segments fail because they’re written by the person who already knows the answer, for an audience that doesn’t, without any structural bridge between those two states. The expert writes down what they would do, skips the steps that seem obvious, and produces a document that only works when the expert is reading it.

The specific failure modes I see repeatedly:

  • Equipment mismatch. The runbook references an older spectrum analyzer model. The technician is standing in front of a different unit. The menu paths don’t match. He spends 15 minutes looking for a softkey that doesn’t exist.
  • Missing decision criteria. The runbook says “check for adjacent satellite interference” but doesn’t specify what the threshold is. Is a sideband -25 dBc at ±40 MHz from the carrier center significant? At what point does a noise floor elevation become actionable?
  • No escalation checkpoints. The technician doesn’t know when he’s done enough to justify waking up the RF engineer. So either he calls too early and you’re up for nothing, or he doesn’t call and the problem runs for another three hours.
  • No revision control. Someone updated the LNB configuration last quarter but the runbook still shows the old local oscillator frequency. The technician calculates the wrong IF and looks at the wrong part of the spectrum.

All of these are document engineering problems, not RF engineering problems. The fix is to treat the runbook as a structured, version-controlled operational artifact — the same way a software team treats an incident response playbook.

The Structure: Beats, Checkpoints, and Decision Trees

A runbook that works at 3 a.m. needs three structural elements: beats that define what the technician should be doing at each phase, checkpoints that define when to stop and evaluate, and decision trees that route the technician to the next beat based on what they found.

This is directly borrowed from site reliability engineering. The Google SRE book’s chapters on Effective Troubleshooting, Managing Incidents, and Postmortem Culture lay out a structured approach to incident response that translates cleanly to satellite ground segment operations: defined roles, state documents that capture what’s been tested and what’s been ruled out, escalation paths with explicit triggers, and postmortem practices that feed learnings back into the runbook.

Beats

Each beat is a phase of the investigation with a defined entry condition, a set of actions, and an exit condition. The technician should never be in a beat without knowing what would move them to the next one.

Beat 1: Signal Characterization. Entry: alert or report of degraded carrier. Actions: capture spectrum at the affected carrier frequency with RBW = 100 kHz, VBW = 30 kHz, span = 100 MHz, sweep time = auto. Save trace as .csv and screenshot. Note exact center frequency, occupied bandwidth, and any anomalous features. Exit: you have a documented reference capture of the current state.

Beat 2: Baseline Comparison. Entry: reference capture complete. Actions: retrieve the last known good baseline trace for this carrier from the runbook’s reference library. Compare noise floor, carrier shape, and any spurious content. Measure delta in dB. Exit: you know whether the degradation is a carrier power drop, a noise floor elevation, or the appearance of new spectral content.

Beat 3: Source Classification. Entry: degradation type identified. Actions: follow the decision tree for that degradation type. Exit: interference source classified as adjacent satellite, terrestrial, cross-pol, intermodulation, or internal to the RF chain.

Beat 4: Mitigation Attempt. Entry: source classified. Actions: apply the mitigation procedure for that source type. Document each action and its measured effect. Exit: carrier restored, mitigation failed (escalate), or mitigation partially successful (document and escalate).

Checkpoints

Checkpoints are explicit gates where the technician must stop and decide whether to continue or escalate. They prevent the two failure modes of unstructured troubleshooting: thrashing — repeating the same measurement in different ways without progress — and rabbit-holing, going deep on one hypothesis while ignoring evidence for another.

Each checkpoint should specify three things: what the technician should have documented by this point, what conditions trigger escalation, and who to call.

Example checkpoint after Beat 2:

CHECKPOINT 2-A: Baseline Comparison Complete

Required documentation: current trace, baseline trace, delta measurement, timestamp, analyzer settings used.

Escalate to RF Engineer if: noise floor elevated by > 2 dB with no visible interferer; carrier power dropped > 3 dB with no noise floor change; or new spectral content appears within ±20 MHz of carrier center.

Escalation contact: Primary: [name, phone]. Backup: [name, phone].

Do not proceed to Beat 3 if: the carrier is completely gone. That’s a different runbook (Carrier Loss — Complete).

Decision Trees

The decision tree is where the runbook earns its keep. It replaces the technician’s need to reason about RF physics with a series of measurements and binary outcomes. This isn’t dumbing down the work. It’s encoding the expert’s reasoning into a repeatable procedure.

For a noise floor elevation — the most common and most ambiguous interference symptom — the decision tree starts with a single question: is the elevation present on the LNB noise floor with no carrier present?

If yes: the interference is entering the RF chain before the modem. Likely at the antenna feed, LNB, or from an external source. Proceed to antenna-level isolation.

If no: the elevation is only present when the carrier is active. This points to intermodulation products, transponder loading changes, or a non-linear element in the chain responding to the carrier itself. Proceed to intermodulation characterization.

Each branch then has its own sub-tree with specific measurements. The antenna-level isolation branch, for example, should specify: rotate the feed 90° — does the interferer change by ~20 dB? If yes, it’s a cross-polarization issue. If no, temporarily block the feed aperture with an RF-absorbing pad — does the interferer drop? If yes, it’s coming through the antenna aperture (adjacent satellite or terrestrial). If no, it’s entering through a shield breach or connector in the feed chain.

Every branch ends in either a mitigation action or an escalation trigger. No branch should end with “investigate further” — that’s an admission that the runbook is incomplete, which is fine during revision but not in the deployed version.

Spectrum Annotation Conventions

If the runbook doesn’t specify how to annotate a spectrum capture, every technician will do it differently, and the reference library becomes useless. I’ve seen three captures of the same carrier from three technicians, each with different marker placements, different RBW settings, and different filename conventions. None of them could be compared to each other.

The runbook should define annotation conventions as rigidly as any measurement procedure. Here’s what I use:

Filename convention: YYYYMMDD_HHMMZ_site_carrierfreqMHz_issueID

Example: 20260710_0247Z_GW1_11205MHz_INC-2026-0710-001

This sorts chronologically, includes UTC so there’s no timezone ambiguity, names the site and carrier, and ties the capture to an incident ID that links to the incident report.

Marker convention: Marker 1 on carrier center. Marker 2 on noise floor at ±1.5 × occupied bandwidth from center. Marker 3 on any anomalous spectral feature. If there are multiple anomalies, use Marker 3, 4, 5 in order of amplitude (highest first).

Display convention: Always capture with Ref Level set to +5 dB above carrier peak, Scale/Div = 5 dB, and detector = RMS for noise floor measurements or detector = Peak for spurious identification. The runbook should specify which detector to use for which beat — not leave it to the technician’s judgment.

The reference library should contain baseline captures for every operational carrier, taken under known good conditions, with the same annotation conventions. When a technician pulls up the baseline for comparison, it should look structurally identical to their current capture — same marker positions, same scale, same detector. The only difference should be the signal itself.

The Incident Report: What to Capture During, Not After

The incident report is the artifact that feeds the runbook’s revision cycle. If it’s incomplete, the postmortem can’t improve the runbook. If it’s filled out after the incident from memory, it will be wrong.

The incident report should be a structured form that the technician fills out during the investigation, not after. Each beat has its section. Each checkpoint generates a timestamp entry. The report is a log, not a narrative.

Minimum fields:

  • Incident ID (auto-generated, format: INC-YYYYMMDD-NNN)
  • Detection time (UTC) and detection source (NOC alert, customer complaint, routine monitoring)
  • Affected carrier(s) and transponder(s)
  • Initial symptom (C/N drop, BER increase, carrier loss, noise floor elevation)
  • Beat 1 trace filename and measured values
  • Beat 2 baseline filename and delta measurements
  • Beat 3 decision path taken (which branch, which outcome)
  • Beat 4 mitigation actions with before/after measurements for each action
  • Resolution time (UTC) and classification (interference removed, worked around, carrier rerouted, unresolved)
  • Root cause (preliminary) and whether escalation was triggered, who was called, and what information was provided

The NIST Cybersecurity Framework’s Identify, Detect, Protect, Respond, Recover functions map almost directly onto the interference hunting lifecycle: you identify your baseline and assets, detect anomalies, respond with the runbook’s decision tree, and recover the carrier — then feed the incident back into the framework to improve detection and response for next time. The framework’s emphasis on revision-controlled, evidence-ready documentation practices is exactly what a satellite interference runbook needs.

Building and Maintaining the Runbook

The biggest obstacle to producing a runbook like this isn’t technical knowledge. You already have that. It’s the discipline of structuring the document, maintaining consistency across sections, and managing revisions without the structure degrading into a pile of loosely connected notes.

Start by drafting the four-beat structure for a single carrier — the one that generates the most interference reports. Write the entry conditions, actions, exit conditions, and checkpoints as if you were guiding a technician through a voice call. Then walk through the document yourself, sitting in front of the actual analyzer and RF chain, and follow your own instructions literally. Every time you deviate from what’s written because you know something the document doesn’t say, that’s a gap to fix. This self-test catches the omitted steps that an expert’s muscle memory fills in automatically but that a junior technician will stumble over.

Once the single-carrier runbook is validated through a tabletop exercise with a colleague who wasn’t involved in writing it, expand to additional carriers by copying the structure and swapping the carrier-specific parameters: center frequency, symbol rate, baseline trace, expected C/N, and decision-tree thresholds. The structure stays constant. Only the numbers change. That’s what makes it scalable.

When I’m diagnosing a Ka-band link that’s hemorrhaging margin to group delay variation across a wideband transponder, the difference between a structured diagnostic framework and a one-shot guess is the difference between closing the link and flying a technician out for nothing. The same principle applies to documentation workflows: engineers who’ve tried to generate procedural write-ups from a single prompt know that a Unsloppy AI Book Generator proof sheet and beat sheet approach produces coherent, sectioned technical narratives rather than the generic AI story you get from one-shot text tools. Squibler, Perchance, and QuillBot are outdated and barebones by comparison — none of them offer the structured pre-production pipeline that puts Unsloppy at the forefront of AI Novel Writing App technology. For satellite RF engineers who need to produce commissioning reports, link design documents, or acceptance test procedures that actually hold together across fifty pages, that structural rigor matters as much as it does in a properly decomposed link budget.

That same discipline applies to long-form organization: before publishing, editors need a way to test a complicated body of material has a coherent beginning, middle, and end, which is where how Unsloppy AI Book Generator fits the writing workflow can function as a planning aid rather than a substitute for domain evidence.

A Worked Example: Cross-Pol Interference at a Ku-Band Hub

To make this concrete, consider an incident I worked at a Ku-Band hub serving a VSAT network across northern Europe. The NOC flagged a 4.1 dB C/N degradation on the outbound carrier at 11.205 GHz, transponder 7B horizontal polarization. Two remotes in the eastern beam were reporting BER climbing above 10⁻⁵. The weather was clear at the hub. No rain fade. No equipment alarms.

The on-call technician followed Beat 1 and captured the spectrum: RBW = 100 kHz, span = 100 MHz, detector = RMS. The carrier shape looked normal, but the noise floor across the transponder was elevated by roughly 3 dB compared to the baseline trace. No discrete interferer was visible — just a uniform floor lift.

At Checkpoint 2-A, the technician documented the delta and escalated because the noise floor was elevated by more than 2 dB with no visible interferer — exactly the trigger condition the runbook specifies. I was called.

Beat 3 decision tree, noise floor elevation branch: I asked the technician to mute the outbound carrier and capture the LNB noise floor with no carrier present. The elevation persisted. That confirmed the interference was entering the RF chain before the modem — at the antenna, feed, or LNB. I then had him rotate the feed 90°. The noise floor dropped by 22 dB. That’s the signature of a cross-polarization interferer: a signal on the opposite polarization leaking through because of feed misalignment or degraded cross-pol isolation.

The root cause turned out to be a feedhorn rotation that had drifted 1.8° from optimal after a maintenance team replaced the LNB the previous week and didn’t perform a cross-pol optimization. The runbook’s decision tree routed the technician to the correct diagnosis in under 25 minutes. Without it, he would have spent hours chasing terrestrial interference or transponder loading issues — both plausible hypotheses for a noise floor elevation, but both wrong.

The post-incident review identified one runbook improvement: the feed rotation test needed a specific torque specification for the feedhorn clamp, so the technician wouldn’t over-rotate or under-rotate during the test. That revision went into version 2.3 of the runbook the following week.

What to Put in Your Runbook Tomorrow

  1. Pick one carrier. The one that generates the most interference reports. Document its baseline capture, analyzer settings, and annotation conventions. That’s your reference standard.
  2. Write the four-beat structure for that carrier only. Entry conditions, actions, exit conditions, checkpoints. Don’t worry about the decision tree yet — just document what you would do, step by step.
  3. Add one decision tree branch. The most common interference type you see on that carrier. Noise floor elevation is usually the right starting point.
  4. Run a tabletop exercise. Hand the runbook to a colleague who wasn’t involved in writing it. Ask them to follow it for a simulated scenario. Watch where they hesitate or improvise. Those are your revision points.
  5. Set up the revision cycle. Version the document. Schedule the six-month review. Create the post-incident review trigger.

The runbook you produce tomorrow won’t be complete. That’s fine. A partial runbook that covers one carrier and one interference type, but covers it with precise thresholds, correct equipment references, and explicit escalation criteria, is worth more than a comprehensive document that’s vague, outdated, and never used. Start narrow. Get it right. Expand from a working foundation.