Daily operations runbook
This runbook condenses the most useful day-to-day commands for the lab. It does not replace the ATP: it summarizes it and links back to the detailed validations when deeper troubleshooting is needed.
1. Overall lab health
Quick review of Containerlab and container status:
sudo clab ins -t lab.yml
docker ps --format "table {{.Names}}\t{{.Status}}\t{{.Ports}}"
Expected result:
- relevant nodes in
running - published ports visible for Grafana, Prometheus, ONTs, and LEA
Main services:
- Grafana:
http://localhost:3030 - Prometheus:
http://localhost:9090 - LEA/LIG:
http://localhost:8092 - ONT1 Web:
http://localhost:8090 - ONT2 Web:
http://localhost:8091
ATP reference:
2. Quick base network review
BNG MASTER / SLAVE
show system alarms
show port
show router interface
show srrp
show router 9999 bgp summary
Expected:
- no critical alarms
- key ports and interfaces
up - SRRP operational
- BGP neighbors established
OLT
show interface all
show network-instance bd-50 interfaces
macsum bd-50
macsum bd-srrp
Expected:
- bridged subinterfaces
up - MAC-VRFs active
- MAC learning consistent with traffic
Carriers
show interface all
show network-instance route-table ipv4-unicast summary
show network-instance route-table ipv6-unicast summary
show network-instance protocols bgp summary
Expected:
- upstream interfaces
up - IPv4/IPv6 routes present
- BGP
establishedwith both BNGs
3. Check subscriber sessions
On the BNG:
show service active-subscribers
On ONT2 PPPoE:
docker exec ont2 sh -lc 'ip ad show ppp0; echo ---; ip -6 route'
On PC1 for delegated prefix:
docker exec pc1 sh -lc 'ip -6 ad show dev eth1'
Expected:
ONT-001and/orONT-002visible depending on the scenarioppp0operational onont2when PPPoE is active- delegated IPv6 prefix visible on
pc1
ATP reference:
4. Clear sessions and force reauthentication
On the BNG:
clear service id "9998" ipoe session all
Note:
- after this
clear, IPoE sessions may take several seconds to rebuild - it is normal for WAN IPv6, delegated prefix, and IPv4/IPv6 leases to come back with different values than before
- if
ONT-001does not reappear after convergence, use DHCP/DHCPv6 renew from the ONT or follow the ESM ATP reference
To suspend or re-enable the PPPoE subscriber:
docker exec containerbot sh -lc 'python3 /app/scripts/manage_authorize.py deactivate "test@test.com"'
docker exec containerbot sh -lc 'python3 /app/scripts/manage_authorize.py add "test@test.com" \
--title "ONT2-WAN1" \
--password "testlab123" \
--framed-pool "cgnat" \
--framed-ipv6-pool "IPv6" \
--delegated-ipv6-pool "IPv6" \
--subscriber-id "ONT-002" \
--subscriber-profile "subprofile" \
--msap-interface "ipv6-only" \
--sla-profile "100M"'
Expected:
- the subscriber disappears and then reappears in
show service active-subscribers ont2loses and recoversppp0
ATP reference:
5. Validate SRRP and BGP recovery
Base verification:
show srrp
show router 9999 bgp summary
show router 9999 bgp neighbor 172.16.1.1 detail | match "Export Policy"
show router 9999 bgp neighbor 172.16.2.1 detail | match "Export Policy"
Simulate role change:
docker exec containerbot sh -lc 'python3 /app/scripts/update-ports-sros.py --host 10.99.1.2 --state disable --port-id 1/1/c1/1'
docker exec containerbot sh -lc 'python3 /app/scripts/update-ports-sros.py --host 10.99.1.2 --state enable --port-id 1/1/c1/1'
This case tests the failure of the link between the BNGs. The SRRP role change also happens here because of policy 1, even though the SRRP message-path still runs over 1/1/c2/1 through the OLT.
Expected:
- SRRP transitions between
masterandbackup - BGP export policies change according to role
- state returns to normal after re-enabling the port
ATP reference:
6. Quick NAT64 validation
From ONT1:
docker exec ont1 sh -lc 'ping -6 -c 4 -I 2001:db8:cccc::1 64:ff9b::808:808'
On the BNG:
tools dump nat sessions
Optional:
pyexec "cf3:\scripts\nat64_portblocks.py"
Expected:
- successful ping to the NAT64 prefix
- NAT64 sessions visible on the BNG
Note:
- in the current validation, the working NAT64 path on
ont1used thewan2IPv6 source (2001:db8:cccc::1) - as a quick application-level alternative, this also works:
docker exec ont1 sh -lc 'curl -6 -I --max-time 15 http://example.com'
ATP reference:
7. Regenerate traffic for observability
Check orchestrator ports:
bash configs/cbot/scripts/observability-traffic.sh status
Generate traffic:
bash configs/cbot/scripts/observability-traffic.sh all
Show report:
bash configs/cbot/scripts/observability-traffic.sh report
Stop servers:
bash configs/cbot/scripts/observability-traffic.sh stop-servers
Expected:
- octet increase for
ONT-001andONT-002 - visibility in Prometheus and Grafana
ATP reference:
8. Quick LEA validation
Precondition:
- there must already be an active LI capture on the BNG
- the intercepted identity may be
ONT-001or a rebuiltMAC|SAPstylesubscriber-id, depending on AAA state - the WAN used to generate traffic must match that intercepted identity
- if you need to prepare or validate that part, use LEA and LI and LEA Console
Start the test server:
bash configs/cbot/scripts/dns-iperf-server.sh start
bash configs/cbot/scripts/dns-iperf-server.sh status
Generate traffic from ONT1:
ONT_USER=user ONT_PASS=test ONT_WAN=wan2 DURATION=12 PARALLEL=1 bash configs/cbot/scripts/ont1-subscriber-traffic.sh upload
If the active capture is tied to wan1, change this to ONT_WAN=wan1.
Query the API:
curl -s http://10.99.1.12:8080/api/stats | jq
curl -s 'http://10.99.1.12:8080/api/events?limit=20' | jq
Stop the server:
bash configs/cbot/scripts/dns-iperf-server.sh stop
Expected:
- events visible in LEA
- counters and flows present in the API
ATP reference:
9. Post-change checklist
After any operational change, validate at least:
sudo clab ins -t lab.yml
docker ps --format "table {{.Names}}\t{{.Status}}"
show srrp
show router 9999 bgp summary
show service active-subscribers
bash configs/cbot/scripts/observability-traffic.sh report
curl -s http://10.99.1.12:8080/api/stats | jq
If everything is healthy:
- lab in
running - SRRP and BGP healthy
- subscribers visible
- observability active
- LEA responsive