Network Engineering Portfolio
This is a real network, running right now. Not a diagram, not a screenshot — the topology and device states below are read live from the lab itself. Three sites carry real traffic across WireGuard tunnels, behind OPNsense firewalls enforcing real policy, watched by a full Prometheus, Grafana and alerting stack. It is built to be broken on purpose: link failures and outages are injected on a schedule to prove the alerting actually fires.
Live — polled from the running lab every 60 seconds · last update moments ago
Every node below is a running system. Select any device to trace its connections. Squares are switches; circles are hosts, firewalls and appliances.
15 of 36 devices are checked directly against the container runtime rather than trusting the emulator's reported state.
Prometheus scrapes every device in the lab; Alertmanager routes what fires. These figures are queried live — no dashboard is exposed to reach them.
Scrape targets healthy
16 / 16
24 hours · now 16
Mean probe round-trip
1.71 ms
24 hours · spikes are injected failures
Monitoring availability
94.3 %
Currently observing · checked from outside the lab every 60s
Last blind spot 6 min ago — prometheus unreachable (URLError)
Scrape jobs: alertmanager 1 · blackbox_gated 1 · blackbox_icmp 11 · node 2 · prometheus 1
| InstanceDown | A monitored host stopped answering scrapes. | 60s | inactive |
| ProbeDown | An ICMP or HTTP probe target became unreachable. | 60s | inactive |
| SSOGateMissing | The SSO gate stopped refusing anonymous requests — catches a security control failing open, not just failing. | 60s | inactive |
| TunnelDown | A site-to-site WireGuard tunnel stopped passing traffic. | 120s | inactive |
| Watchdog | Always-firing heartbeat. If this stops arriving, the alert pipeline itself has failed — the classic dead-man's switch. | instant | firing |
Alerting that has never fired is a guess. On a schedule, a real failure is injected into the running lab, the time to detection is measured, and the lab is healed automatically. Fastest detection: 71s. Every run below actually happened.
| Severed the primary site's WAN uplink | detected by ProbeDown | 1 h ago | 122s |
| Severed the primary site's WAN uplink | detected by ProbeDown | 1 h ago | 71s |
| Severed the primary site's WAN uplink | detected by ProbeDown | 2 h ago | 81s |
| Severed the primary site's WAN uplink | detected by ProbeDown | 7 h ago | 71s |
| Severed the primary site's WAN uplink | detected by ProbeDown | 7 h ago | 81s |
| Severed the primary site's WAN uplink | detected by ProbeDown | 7 h ago | 91s |
Device names are generalised to their operational role. Platform and tooling are named as-is — that part is the point.
| Alertmanager | alert routing | up |
| Blackbox Exporter | probes | up |
| Grafana | dashboards | up |
| Hypervisor | Proxmox VE | up |
| Ingress Proxy | Traefik | up |
| Log Aggregator | VictoriaLogs | up |
| Monitoring & Automation Host | Prometheus / n8n | up |
| NAS | OpenMediaVault | up |
| Primary Access Switch 01 | — | up |
| Primary Access Switch 02 | — | up |
| Primary Access Switch 03 | — | up |
| Primary Core Switch | — | up |
| Primary Site Edge Firewall | OPNsense | up |
| Primary Site Ingress | Traefik | up |
| Primary Site WLAN Client | — | up |
| Primary WLAN Switch | — | up |
| Workstation | — | up |
| Notification Service | ntfy | up |
| Secondary Access Switch 01 | — | up |
| Secondary Access Switch 02 | — | up |
| Secondary Core Switch | — | up |
| Secondary Site Application Host | — | up |
| Secondary Site Edge Firewall | OPNsense | up |
| Secondary Site WLAN Client | — | up |
| Secondary WLAN Switch | — | up |
| Swarm Node 01 | Docker Swarm | up |
| Swarm Node 02 | Docker Swarm | up |
| Swarm Node 03 | Docker Swarm | up |
| Remote Site Edge Firewall | OPNsense | up |
| Remote Site Probe Host | smokeping | up |
| Remote Site Switch | — | up |
| Cloud VPS 01 | Automation | up |
| Cloud VPS 02 | Monitoring | up |
| ISP Router | double-NAT | up |
| Internet | — | up |
| Internet Exchange | — | up |