I'm taking a full service provider network from empty database to AI-monitored NOC, and writing down every step.
The series is 10 parts. By the end, we'll have 28 devices running ISIS, MPLS, BGP, and EVPN. Configs generated from a source of truth. Monitoring that catches failures across layers. And an AI agent that receives alerts, pulls route tables and interface state, and writes up what broke without anyone logging into a router.
All of it runs on one Proxmox box. No cloud. No physical gear beyond a single server.
Here's what the topology looks like: an MPLS L3VPN core connecting 3 datacenter customers. Each customer site runs a leaf-spine EVPN/VXLAN fabric. 13 Cisco IOS-XE routers and 15 Arista EOS switches. The monitoring stack is Prometheus, Loki, Grafana, and OTel Collector, with device inventory pulled from Nautobot. The AI piece is NetClaw, an agent I built that receives alerts from Alertmanager, queries all those systems for context, and produces root cause analysis.
What I'm showing across the series:
1. SoT-driven infrastructure where every device, IP, cable, and BGP session lives in Nautobot and everything else queries it
2. Config generation from templates, not manual CLI
3. Failure detection across network, transport, and application layers
4. AI root cause analysis pulling state from before and after the alert
5. Config compliance and drift detection
6. The whole thing reproducible on one box
I believe the next generation of network engineers won't get the apprenticeship I had. Nobody's going to hand them an expensive network and say "learn by doing" for two years. So I'm building the lab, the walkthroughs, and eventually a rentable environment where you can follow along hands-on.
Free subscribers get post previews and occasional full unlocks. Paid subscribers get every walkthrough, every config, every template. Founding Members get all of that plus 20% off lab rentals when they launch, and priority on what I cover next.
If you learn by building, you're in the right place.


