Network Engineering Portfolio

Multi-Site Network Lab

This is a real network, running right now. Not a diagram, not a screenshot — the topology and device states below are read live from the lab itself. Three sites carry real traffic across WireGuard tunnels, behind OPNsense firewalls enforcing real policy, watched by a full Prometheus, Grafana and alerting stack. It is built to be broken on purpose: link failures and outages are injected on a schedule to prove the alerting actually fires.

Devices
36
Running
36
Links
35
Sites
3

Live — polled from the running lab every 60 seconds · last update moments ago

Topology

Every node below is a running system. Select any device to trace its connections. Squares are switches; circles are hosts, firewalls and appliances.

15 of 36 devices are checked directly against the container runtime rather than trusting the emulator's reported state.

In simulate mode, click devices to take them offline — the cascade is computed in your browser and changes nothing in the lab.
InternetPrimary SiteEdge FirewallPrimaryCore SwitchWorkstationHypervisorInternetExchangeSecondary SiteEdge FirewallSecondaryCore SwitchSwarmNode 01SwarmNode 02ISP RouterPrimaryWLAN SwitchPrimary SiteWLAN ClientPrimary AccessSwitch 01Primary AccessSwitch 03NASSecondaryAccess Switch 01Secondary SiteApplication HostSecondaryAccess Switch 02Monitoring &Automation HostSwarmNode 03IngressProxyGrafanaPrimary AccessSwitch 02PrimarySite IngressAlertmanagerBlackboxExporterNotificationServiceLogAggregatorSecondaryWLAN SwitchSecondary SiteWLAN ClientCloud VPS 01Cloud VPS 02Remote SiteEdge FirewallRemoteSite SwitchRemote SiteProbe Host
Fig. 1 Physical and logical topology, drawn from the lab's own layout. Colour encodes live run state.

Monitoring

Prometheus scrapes every device in the lab; Alertmanager routes what fires. These figures are queried live — no dashboard is exposed to reach them.

Scrape targets healthy

16 / 16

24 hours · now 16

Mean probe round-trip

1.71 ms

24 hours · spikes are injected failures

Monitoring availability

94.3 %

Currently observing · checked from outside the lab every 60s

Last blind spot 6 min ago — prometheus unreachable (URLError)

Scrape jobs: alertmanager 1 · blackbox_gated 1 · blackbox_icmp 11 · node 2 · prometheus 1

Alert rules

InstanceDownA monitored host stopped answering scrapes.60sinactive
ProbeDownAn ICMP or HTTP probe target became unreachable.60sinactive
SSOGateMissingThe SSO gate stopped refusing anonymous requests — catches a security control failing open, not just failing.60sinactive
TunnelDownA site-to-site WireGuard tunnel stopped passing traffic.120sinactive
WatchdogAlways-firing heartbeat. If this stops arriving, the alert pipeline itself has failed — the classic dead-man's switch.instantfiring

Continuous verification

Alerting that has never fired is a guess. On a schedule, a real failure is injected into the running lab, the time to detection is measured, and the lab is healed automatically. Fastest detection: 71s. Every run below actually happened.

Severed the primary site's WAN uplinkdetected by ProbeDown1 h ago122s
Severed the primary site's WAN uplinkdetected by ProbeDown1 h ago71s
Severed the primary site's WAN uplinkdetected by ProbeDown2 h ago81s
Severed the primary site's WAN uplinkdetected by ProbeDown7 h ago71s
Severed the primary site's WAN uplinkdetected by ProbeDown7 h ago81s
Severed the primary site's WAN uplinkdetected by ProbeDown7 h ago91s

Inventory

Device names are generalised to their operational role. Platform and tooling are named as-is — that part is the point.

Primary Site17

Alertmanageralert routingup
Blackbox Exporterprobesup
Grafanadashboardsup
HypervisorProxmox VEup
Ingress ProxyTraefikup
Log AggregatorVictoriaLogsup
Monitoring & Automation HostPrometheus / n8nup
NASOpenMediaVaultup
Primary Access Switch 01up
Primary Access Switch 02up
Primary Access Switch 03up
Primary Core Switchup
Primary Site Edge FirewallOPNsenseup
Primary Site IngressTraefikup
Primary Site WLAN Clientup
Primary WLAN Switchup
Workstationup

Secondary Site11

Notification Servicentfyup
Secondary Access Switch 01up
Secondary Access Switch 02up
Secondary Core Switchup
Secondary Site Application Hostup
Secondary Site Edge FirewallOPNsenseup
Secondary Site WLAN Clientup
Secondary WLAN Switchup
Swarm Node 01Docker Swarmup
Swarm Node 02Docker Swarmup
Swarm Node 03Docker Swarmup

Remote Site3

Remote Site Edge FirewallOPNsenseup
Remote Site Probe Hostsmokepingup
Remote Site Switchup

Cloud2

Cloud VPS 01Automationup
Cloud VPS 02Monitoringup

Edge3

ISP Routerdouble-NATup
Internetup
Internet Exchangeup