Coordinator / NOC Operator

Coordinator and NOC Operator are the same role in this app — one person, one login, both jobs. This is your daily routine end to end: the shift schedule, the triage flow, the ticket pipeline, and the rules that make a day count as successful. Source: NOC Daily Playbook v2.

Your shift, in six steps

Open the NOC Dashboard (sidebar → Network Operation Center → your project → NOC Dashboard) and keep it on screen all shift. These six steps repeat every day, in this order:

1

Shift start · first 30 min

The Sweep

Check the NMS is polling — broken monitoring makes everything below fiction.

Count: N down out of N total. Real numbers.

Cross-check every down site against its provider cloud. This is the job.

Cluster by area, provider, or utility before opening anything.

Check weather and utility advisories.

Clear the "needs a ticket" counter — every site past 24h gets one today.

Cause-code everything, even if provisional.

Nobody watched the network overnight. This is the most important half hour of your day.

2

+0:30 · every day

Morning Report

Sites up / down out of total.

NMS health, and any NMS-vs-provider discrepancies.

Down sites grouped by cause code, not listed one by one.

Sites past 24h and whether each has a ticket.

Any cluster: area named, external reference, map attached.

Carried over from yesterday, plus next action.

Send it on quiet days too. "All up, nothing carried over" is a complete report.

3

All shift

The Loop

Always: NMS on screen. New offline site → run the triage flow.

Hourly: rescan the down list, advance the 24h clocks, chase what has not moved.

Every 2h: status line on any open ticket for a 24/7 site or a cluster.

Chase any ticket sitting in one pipeline stage without activity.

Acking alarms then going quiet until 5pm is watching, not monitoring.

4

Midday · 15 minutes

Reconcile

Past 24h with no ticket? Raise it now.

Ticket open but site recovered in both sources? Close it with a cause.

Still tagged UNK? Investigate or escalate.

Any ticket untouched 48h? Chase or escalate.

Spreadsheet disagrees with the ticket? The ticket wins.

Fifteen minutes here saves an hour at close.

5

1 hour before close

Pre-Close

Resolve anything an hour of work would close.

Escalate every 24/7 site still down, and every ticket that will breach, now.

Tell affected sites what's down overnight.

Chase carriers and coordinators one last time; log their commitment.

Write tomorrow's first action on every open ticket.

Nothing enters the unattended overnight window without a superadmin knowing.

6

Last 15 minutes

Handover

Lead with what stays down overnight and whether the site was told.

Sites whose 24h clock expires tomorrow — flag them for the next operator.

Open tickets, their stage, and the first action tomorrow.

Say "nothing outstanding" explicitly if true.

Post it and get it acknowledged before logging off (Helpdesk → Handover tab).

No handover = incomplete shift, however good the network was.

The triage flow: verify before you act

We are a network monitoring center, not an end-of-day monitoring center. A site showing down in the NMS is not yet an outage — it only becomes one after it survives verification and its hold window. Every offline site goes through this, in order:

  1. Site shows OFFLINE in the NMS. Note the site code, group, provider, and offline timer. Don't act on the NMS reading alone.
  2. Cross-check the provider cloud. Open the matching portal — Ruijie, Huawei, Omada, or Starlink — and check the same device. This is the verification step and it is never skipped.

Down on both

Real outage → go to step 3

Both sources agree. The site is genuinely unreachable. Continue to the hold rule.

Down on NMS, up on provider

Discrepancy → not a site outage

Our monitoring is wrong, not the site. Log it as NMS-DISC, report it in the morning report, and escalate to a superadmin if more than two appear in a day.

Provider portal unreachable

Treat as down on both

Don't stall waiting for a portal. Proceed on the NMS reading and note that verification wasn't possible.
  1. Check the site's schedule tag. The hold rule depends entirely on this — check it before starting any clock.

Runs 24/7

Ticket immediately. No hold.

There is no legitimate reason for this site to be off. Waiting 24 hours on a 24/7 site is lost time, not caution.

Scheduled

Inside its window? No ticket.

Note it in the morning report. If it's down outside its scheduled window, treat it as a standard site.

Standard site

Start the 24-hour hold.

Schools switch off after hours and on weekends. Watch it. Recovered inside 24h means no ticket, but the event is still logged.
  1. Still down at 24 hours → raise the ticket. The NMS already flags these for you: "Offline > 24h, needs a ticket." That counter should be zero every time you finish a sweep — it's the single clearest measure of whether the NOC is keeping up.

Why we still ticket intentional shutdowns

A ticket is how we obtain acknowledgement that a site was switched off on purpose. Closing it as SCHED-OFF with the site's confirmation is a successful outcome, not a wasted ticket — it turns an unexplained gap into a documented one, and protects us when uptime is questioned.

The color bar: which down sites need you

In the Sites tab of the NOC Dashboard, each site row can carry a colored bar down its left edge — the triage flow above, made visible at a glance. Hover the site code to see the exact reason.

  • Amber — needs a ticket. The site has been offline more than 24 hours and has no open ticket. This is your cue to open one.
  • Blue — tracked. There's an open ticket and it has had activity in the last 48 hours. It's covered and moving; no action needed.
  • Red — neglected. There's an open ticket but nobody has touched it in 48 hours (napapabayaan). Chase it — add an update, move it forward, or escalate.

A down site with no bar is either offline under 24 hours (too early to require a ticket) or its ticket is already in Review or Done. Don't open a second ticket for a site that already shows blue or red.

Where the ticket goes: the helpdesk pipeline

Every ticket moves through five stages. In this app, the New and Coordinator stages are both worked by you — Coordinator and NOC Operator are the same role, not a handoff between two people. There's also no separate "NOC Lead" title: Review and escalation authority sit with a coordinator or superadmin.

StageOwnerMoves on when
NewYou (NOC Operator)Ticket created with site code, offline duration, both-source verification result, and cause code (or UNK).
CoordinatorYou (Coordinator)Site or school contacted. Outcome recorded: intentional shutdown, power issue, or fault requiring intervention.
ITIT SupportRemote diagnosis and remediation attempted and logged. Escalates to Installer if hands are needed.
InstallerFieldOnsite visit completed, findings recorded, service restored or blocker documented.
ReviewCoordinator / SuperadminReal cause code, usable resolution note, site confirmed back online in both sources.

Watch this

Tickets stall in Coordinator. Sat there 48h with no activity? It's neglected — the NMS flags it (the red bar above). Chase or escalate.

Cause codes: every ticket gets one

Provisional is fine while a ticket is open — but it needs a real code, not UNK, before it closes.

SCHED-OFFConfirmed intentional shutdown, acknowledged by the site
NMS-DISCDown on NMS, up on provider cloud
PWR-COMCommercial power out at site
PWR-SITEUnplugged, UPS, battery, breaker
WX-RAINRain or flooding
WX-STORMTyphoon or severe weather
SAT-OBSStarlink obstruction or rain fade
FIB-CUTFiber or cable cut
CAR-FAULTCarrier fault, needs their ticket ref
EQP-FAILHardware failure
CFG-CHGCaused by a change, ours or theirs
SITE-ACCSite access refused or unavailable
CLI-SIDEClient equipment or client action
MNT-PLNPlanned maintenance
UNKOpen only — never a closing code

The rules, in one place

RuleMeaning
Verify twiceNMS reading alone is never enough. Provider cloud confirms or contradicts it.
24-hour holdStandard sites only. Recovered inside 24h → log it, no ticket.
24/7 sitesNo hold. Ticket the same day it drops.
Scheduled sitesInside window → no ticket. Outside window → treat as standard.
Ticket anywayEven confirmed intentional shutdowns get a ticket, closed with acknowledgement.
48-hour ruleNo ticket sits in one stage without activity for 48h.
Ticket is truthIf it's not in the ticket, it did not happen. The spreadsheet is output, not a record.

Never

Ticket straight off the NMS without checking the provider cloud
Apply the 24h hold to a 24/7 site
Let a site pass 24h with no ticket
Skip the ticket because "it's just the school turning it off" — get the acknowledgement
Close a ticket before the site is back in both sources
Let a ticket sit 48h untouched in any stage
Silently fix a discrepancy in the spreadsheet instead of the ticket
Batch reporting to end of day
Wait for the site to call us
Close a ticket as UNK
Leave without a written handover

What every ticket must contain

FieldWhat
Site code + namee.g. PP-BATS-2026-011
Offline sinceTimestamp, not "yesterday"
VerifiedNMS + which provider portal, and what each showed
Schedule tag24/7, scheduled, or standard
Cause codeProvisional is fine on open
Contact attemptsWho, when, outcome
Next action + ownerAlways populated

One line to remember

The NMS tells you a site is unreachable. Only the provider cloud plus a phone call tells you it is broken. Never report the first as if it were the second.

Diagnose, then escalate

For each down site, work the cheapest action first:

1

Check the provider cloud / remote reset

Coordinator → IT Support

Look at the provider cloud/controller for the device's status and logs. If it is reachable or flapping, attempt a remote reset/reboot. Most outages clear here.

2

Call the site contact

Coordinator

If the remote reset does not bring it back, call the beneficiary/site contact to ask what is happening on the ground (power brownout, weather, unplugged router) and have them power-cycle the router.

3

Recommend a site visit

Dispatch Installer

ONLY if the remote reset failed AND an on-site restart also did not work, or the contact reports physical/hardware damage. Only then send a technician.

Escalation here means the response ladder above — remote reset, then a call, then (only as a last resort) dispatch. Separately, the table below is about visibility: when a coordinator or superadmin needs to know something is happening, regardless of whether it's resolved yet. When you escalate, hand off the full picture — project, site/group, device, how long it's been down, what you already tried, and anything the beneficiary reported.

Escalate whenTiming
24/7 site downSame day
3+ sites down in one areaImmediately
More than 2 NMS discrepancies in a daySame day
NMS itself is down or not pollingPhone, now
Ticket neglected 48hSame shift
Site unreachable / won't acknowledgeAfter 2 attempts
Site told us before we told themSame shift, logged as a miss
Anything open at close of business1 hour before close

Area-wide outages behave differently

When three or more sites within about 10 km go down together, the Outage Triage flags it as an area-wide cluster. These almost always come from a shared cause — a power utility outage or a backhaul provider issue — not the sites themselves.

For clusters, hold dispatch and coordinate with the power/backhaul provider. Sending an installer to one site won't fix a regional outage. Escalate immediately — see the table above.

Before you log off

You had a successful day if

Sweep done and report out within 30 min of start
Every down site was cross-checked against its provider cloud
The "offline > 24h, needs a ticket" counter is at zero
Every 24/7 site that dropped got a ticket the same day
No ticket sat in one stage for 48h without activity
Any NMS-vs-provider discrepancy was reported, not silently ignored
No site told us before we told them
Every closed ticket has a real cause code, none left as UNK
Overnight items escalated and sites informed
Handover posted and acknowledged

What success is not

A successful day is not a day with no outages. You don't control the weather, the grid, or a school's light switch. A successful day is one where everything that happened was verified against two sources, held or ticketed by the rule, explained, and written down.