📝Diagnosing Network Problems | A Comprehensive Checklist✅
Discover & Fix Network Problems | A Detailed Checklist Guide

Driving SD-WAN Adoption in South Africa
Search for a command to run...
Discover & Fix Network Problems | A Detailed Checklist Guide

Driving SD-WAN Adoption in South Africa
No comments yet. Be the first to comment.
Enhancing Security and Compliance with ManageEngine ADAudit Plus

Why Pings Aren't Enough & NMS is Essential

What is SD-WAN (Software Defined Wide Area Networking)? | The Mechanics of this Groundbreaking New Network Technology

Embracing First Principles & the Scientific Method

Ever Wondered | "Is My Internet Really 99.9% Reliable?" 🤔

Network issues can bring business operations to a halt, making it vital to have a structured and thorough diagnostic approach. This checklist provides a step-by-step guide to identifying and resolving network issues. Each step is explained in detail, highlighting what to check, how to perform the check, and the potential symptoms or issues it addresses.
These checks focus on the core of the problem and aim to address the most likely and impactful issues.
What to Do: Document the symptoms observed (e.g., loss of connectivity, degraded services, or intermittent issues).
Why It’s Important: Understanding the problem's scope helps prioritize efforts and communicate effectively with stakeholders.
What to Do: Review whether Service Level Agreement (SLA) targets are being met for downtime or performance metrics.
Symptoms: SLA breaches may indicate systemic issues.
Fix: Escalate to higher-tier support or re-prioritize resources.
What to Do: Check change logs for recent updates to configurations or hardware.
Symptoms: New issues often correlate with recent changes.
Fix: Roll back changes if necessary.
What to Do: Review local weather conditions for storms or extreme temperatures.
Symptoms: Weather can cause signal loss, cable damage, or power outages.
Fix: Mitigate environmental effects or schedule repairs.
What to Do: Verify power supplies and UPS systems.
Symptoms: Unresponsive hardware or intermittent outages.
Fix: Address power disruptions and ensure redundancy.
What to Do: Inspect the hardware for faults or warnings.
Symptoms: Alerts from SNMP or CLI interfaces.
Fix: Replace faulty components or update firmware.
What to Do: Monitor CPU, memory, and bandwidth usage.
Symptoms: High utilization causes slow responses or dropped packets.
Fix: Optimize load or upgrade resources.
What to Do: Check CLI and SNMP access to affected devices.
Symptoms: Inaccessibility may indicate hardware or network issues.
Fix: Restore management access or replace the device.
What to Do: Verify port speed and duplex settings.
Symptoms: Mismatched settings cause collisions and slow speeds.
Fix: Align settings to match connected devices.
What to Do: Inspect ports for Cyclic Redundancy Check (CRC) errors.
Symptoms: Errors lead to packet retransmissions.
Fix: Replace faulty cables or clean fibre connectors.
What to Do: Physically inspect cables for damage or improper connections.
Symptoms: Frayed cables or loose connections impact signal quality.
Fix: Replace or secure cables.
What to Do: Validate rate limits against the customer order.
Symptoms: Misconfigured limits reduce throughput.
Fix: Adjust configurations.
What to Do: Check traffic patterns and load.
Symptoms: High latency or packet loss.
Fix: Re-route traffic or upgrade bandwidth.
What to Do: Look for STP logs or unusual broadcast storms.
Symptoms: Network instability and slowdowns.
Fix: Correct loop sources or enable spanning tree protocols.
These provide additional context and help isolate complex issues.
What to Do: Identify how the issue disrupts operations.
Symptoms: Lost revenue, reduced productivity.
Fix: Focus remediation on critical business processes.
What to Do: Record the start, detection, and resolution times.
Symptoms: Delayed responses exacerbate problems.
Fix: Improve monitoring and alerting.
What to Do: Check logs for hardware-related errors.
Symptoms: Device malfunctions or overheating.
Fix: Replace hardware.
What to Do: Run Ethernet OAM diagnostics.
Symptoms: Link issues on the last mile.
Fix: Coordinate with ISPs for repairs.
What to Do: Test fibre optic links for signal strength.
Symptoms: Intermittent or no connectivity.
Fix: Clean, replace, or re-splice fibre cables.
These help gather background information and provide clarity.
What to Do: Take photos and perform physical inspections.
Symptoms: Documentation aids in troubleshooting.
Fix: Identify overlooked physical issues.
What to Do: Reference network diagrams and configurations.
Symptoms: Misconfigurations can cause unexpected behaviour.
Fix: Update diagrams and verify configurations.
Recording event timings is critical for SLA analysis and post-mortem reviews:
Time of Incident Start: Establish the onset.
Time of Detection: Measure responsiveness.
Time of Repair and Recovery: Document resolution steps.
Downtime Duration: Understand the impact.
By systematically applying checklists, you can diagnose and resolve network problems efficiently, minimising downtime and ensuring reliable service for your business.
When troubleshooting network issues, identifying proximate causes is critical to resolving incidents quickly and effectively. The following detailed checks focus on physical, configuration, and transmission layer elements, which are often the root causes of network disruptions.
Visual Check of Cabling
Description: Inspect cables for visible damage, loose connections, or improper routing.
Symptoms: Frayed or tangled cables, improper bends, or exposed wires may lead to connectivity loss or degradation.
Resolution: Replace damaged cables and secure connections to avoid further disruptions.
Photos of Cabling
Description: Take photos for documentation and remote assessment.
Symptoms: Visual records help identify overlooked issues.
Resolution: Share images with team members or vendors for additional insights.
Patches and Fibre Optic Cables
Description: Examine patch cables and fibre optics for damage or improper terminations.
Symptoms: Bent, cracked, or improperly connected patches can cause high attenuation.
Resolution: Replace or reterminate damaged patches and clean fibre connectors.
SFP/XFP Ports
Description: Inspect transceiver modules for physical damage or malfunction.
Symptoms: Link flaps or complete link failures may result from damaged ports.
Resolution: Replace faulty transceivers.
Link Testing
Description: Perform diagnostic tests on the link, such as loopback or BER (Bit Error Rate) tests.
Symptoms: Errors indicate signal quality or transmission issues.
Resolution: Reconfigure or replace faulty transmission paths.
Maximum Link Lengths
Description: Verify that cable lengths adhere to standard limits (e.g., Ethernet maximum of 100m for copper).
Symptoms: Signal degradation and increased packet loss occur when limits are exceeded.
Resolution: Replace cables with appropriate lengths or install intermediate equipment.
Fibre Pigtails
Description: Ensure RX (Receive) and TX (Transmit) ends are correctly connected.
Symptoms: Incorrect connections cause no signal transmission.
Resolution: Swap connections as necessary and inspect for half-breaks.
Fibre Attenuation Limits
Description: Measure signal attenuation to ensure it is within allowable limits.
Symptoms: High attenuation leads to data loss and signal degradation.
Resolution: Clean connectors, replace damaged cables, or add signal boosters.
Link Frequency/Type Compatibility
Description: Confirm transceivers and fibres match in type and wavelength.
Symptoms: Mismatched components result in no link or poor performance.
Resolution: Use compatible equipment.
Cleaning Fibre
Description: Clean connectors to remove dust or debris.
Symptoms: Dirty connectors lead to signal loss and high error rates.
Resolution: Use fibre cleaning kits before reconnection.
Rate Limits
Description: Verify customer-ordered bandwidth limits are correctly applied.
Symptoms: Misconfigurations can result in slow or throttled connections.
Resolution: Adjust configurations to align with SLAs.
Management VLANs
Description: Ensure VLANs are correctly provisioned for management traffic.
Symptoms: Incorrect VLANs disrupt device access and monitoring.
Resolution: Update VLAN configurations.
Broadcast Traffic
Description: Monitor broadcast and unicast traffic ratios for anomalies.
Symptoms: Excessive broadcasts cause network congestion.
Resolution: Apply broadcast filtering and optimise traffic distribution.
IP/Subnet Configuration
Description: Verify the assigned IP, subnet mask, and gateway configurations.
Symptoms: Misconfigured addresses prevent connectivity or cause routing issues.
Resolution: Correct the network configurations as needed.
Ping and MTR Tests
Description: Conduct pings with varying packet sizes and MTR tests to assess connectivity and path quality.
Symptoms: Packet loss or high latency indicates transmission issues.
Resolution: Investigate faulty paths or congestion points.
Congestion Issues
Description: Check for bottlenecks on primary or backup transmission paths.
Symptoms: High latency and packet drops occur under heavy traffic.
Resolution: Re-route traffic or upgrade link capacity.
Layer 2 Loops
Description: Inspect for spanning tree issues or broadcast storms.
Symptoms: Network-wide slowdowns or outages.
Resolution: Enable STP or address the source of loops.
Path Protection
Description: Check path protection configurations and logs for flapping.
Symptoms: Intermittent disruptions and instability in redundancy mechanisms.
Resolution: Correct misconfigurations and stabilise paths.
Radio Interference
Description: Evaluate potential sources of interference affecting radio links.
Symptoms: Dropped signals or reduced throughput.
Resolution: Eliminate self-interference and external interference sources.
Line of Sight (LOS)
Description: Inspect for physical obstructions or misalignment.
Symptoms: Signal degradation or complete loss of radio links.
Resolution: Align antennas and clear obstructions.
Link Synchronisation
Description: Ensure links are synchronised for optimal performance.
Symptoms: Out-of-sync links cause jitter and packet loss.
Resolution: Resynchronise and optimise configuration.
MTU Alignment
Description: Verify MTU settings along the transport path.
Symptoms: Mismatched MTUs cause fragmentation and performance degradation.
Resolution: Align MTUs across devices and paths.
Throughput and Distance
Description: Measure throughput and verify it meets expected performance for the link’s distance.
Symptoms: Low throughput indicates potential signal loss or hardware limitations.
Resolution: Adjust configurations or upgrade components.
By addressing these detailed checks systematically, network teams can effectively identify and resolve root causes of connectivity or performance issues. Each step provides insights that contribute to faster resolution times and improved network reliability.
When investigating network outages or faults, conducting thorough checks on individual components is essential. The following checklist ensures a systematic approach to identifying and resolving issues, focusing on the physical, operational, and logical aspects of network components.
Issues Identified by Component Checks
Fan Speed Status
Temperature Status
Power Supply Status
Power-On Self-Tests (POSTs)
Unexplained Resets or Violations
CPU, Memory, and File System Status
VLAN Configuration
Firmware Status
Ensure the component runs the correct firmware version.
Review release notes of the latest firmware for fixes related to the observed problems.
MAC Address Learning
RFC2544 Test Results
SLA Measurements Using Y.1731
Utilization and Capacity
Separate Testing
Wireshark Analysis
Ethernet Port Settings
Ethernet Port Statistics
Assess port statistics for errors, drops, or anomalies.
Check for CRC errors, pause frames, or traffic inconsistencies.
Link LEDs
Cable and Radio Connections
Traffic Counters
Port Resets
Switch Port Configuration
Disabled Ports
Port and Cable Labelling
Physical Connections
Ethernet OAM on Last Mile Link
POE Functionality
By systematically working through this checklist, network teams can gain detailed insights into component-level issues, enabling more precise diagnoses and faster resolutions for outages or faults.
When diagnosing outages and faults in transport and network paths, a systematic approach to verifying physical, configuration, and operational elements can help identify the root cause. The checklist below outlines key checks to ensure all potential factors are examined.
Cabling Visual Inspection
Conduct a visual check of the cabling to ensure there are no physical defects or irregularities.
Verify that photos of the cabling are available for documentation and comparison.
Patch Cable Integrity
SFP/XFP Status
Fibre Pigtails
Ensure fibre pigtails are correctly connected, with RX linked to TX.
Check for half breaks or signs of wear in fibre pigtails.
Link Length and Attenuation
Verify that the link lengths do not exceed the specified maximums for the cable type.
Test fibre attenuation to ensure it falls within allowable limits.
Fibre Cleanliness
Signal Loss or Fibre Damage
Link Testing
Frequency and Type Compatibility
Rate Limit Compliance
Management VLANs
Broadcast and Unicast Ratios
Assess the ratio of broadcasts to unicasts to detect abnormal traffic patterns.
Confirm that a broadcast filter has been configured where necessary.
IP Address and Subnet Configuration
Ping and MTR Diagnostics
Test ping responses and traceroutes (MTRs) for packet loss, latency, and anomalies.
Check if pings with varying packet sizes fail, indicating potential MTU issues.
Congestion Issues
Layer 2 Loops
Path Protection
Verify if path protection mechanisms are active and configured correctly.
Review logs for any path protection flaps and adjust configurations as needed.
Asymmetrical Traffic
Link Synchronisation
Radio Interference
LOS (Line of Sight) Issues
Review photos and physical inspections for any visible LOS obstructions.
Ensure the expected throughput over the link is within budget.
Bit Error Rate (BER)
MTU Alignment
By following the above checklist, network teams can methodically examine the physical and logical aspects of transport and network paths. This ensures a comprehensive understanding of potential faults, aiding in faster resolution and improved network reliability.