Showing posts with label SD-WAN Troubleshooting. Show all posts
Showing posts with label SD-WAN Troubleshooting. Show all posts

Sunday, 9 August 2026

The Agent Loop in Cisco SD-WAN: How AI Can Automate Network Troubleshooting

 Artificial intelligence is changing the way network engineers approach monitoring, troubleshooting, and automation.

Traditionally, a network engineer might receive an alert, log in to the Cisco SD-WAN management platform, check device health, investigate WAN links, examine tunnel status, review application performance, and then decide what action should be taken.

Now imagine an AI agent performing many of these activities automatically.

It could collect SD-WAN information, analyze the results, decide what to investigate next, evaluate whether its findings make sense, and continue until it identifies the likely problem.

If a production configuration change is required, it could stop and request approval from the network engineer.

This is the idea behind the Agent Loop.

For network engineers, the concept can be summarized as:

Understand → Act → Observe → Reflect → Decide → Repeat or Stop

The source material describes the complete agent loop as an extension of the basic Think → Act → Observe cycle, adding reflection and a termination decision.

1. What Is an Agent Loop?

An Agent Loop is a continuous cycle in which an AI agent works toward a defined objective.

A simplified workflow looks like this:

                 Network Problem
                       ↓
                Understand Goal
                       ↓
                     Think
                       ↓
                      Act
                       ↓
                   Observe
                       ↓
                   Reflect
                       ↓
                Should I Stop?
                  ↙         ↘
                Yes          No
                 ↓            ↓
              Result       Next Action
                              ↓
                           Repeat

The important part is that the agent doesn't necessarily follow one fixed sequence.

The result of one action can influence the next action.

This makes the concept particularly interesting for SD-WAN troubleshooting.

For example, an agent might initially decide to check WAN connectivity. If the WAN links are healthy, it may decide that checking path performance or application routing is more useful than repeating the same connectivity check.

The source material describes reflection as an evaluation performed after an action, including questions about whether the action succeeded, whether the agent is closer to its goal, whether it should try something different, and whether it is time to stop.

2. Why Is Agentic AI Interesting for Cisco SD-WAN?

SD-WAN environments generate a large amount of operational information.

A network engineer may need to investigate:

  • WAN connectivity
  • Transport availability
  • Tunnel status
  • Latency
  • Jitter
  • Packet loss
  • Application performance
  • Routing
  • Device health
  • Policy behavior
  • Control-plane information
  • Security events

A human engineer can investigate these individually.

An AI agent could potentially correlate multiple pieces of information and determine what to investigate next.

For example:

Problem: Users report that an application is slow.

The agent could investigate:

WAN health → path performance → application traffic path → policy → routing → possible remediation

The important concept is that the next step can be influenced by the result of the previous step.

That is what makes an agent different from a simple fixed automation script.

3. The Basic Think → Act → Observe Model

A useful way to understand agent behavior is:

Think → Act → Observe

Let's apply this to a Cisco SD-WAN troubleshooting scenario.

Think

The agent receives:

"The ERP application is slow for users at the branch."

The agent decides that it needs to investigate whether WAN performance is affecting the application.

Act

The agent collects WAN performance information.

Observe

The agent discovers that one transport has elevated latency and packet loss.

Now the agent has new information.

It can decide what to investigate next.

For example:

"I found a degraded transport. I should determine whether ERP traffic is currently using this transport."

This is the beginning of an agent loop.

4. Reflection: The Step That Makes the Agent Adaptive

Reflection is one of the most important parts of the complete Agent Loop.

After an action, the agent evaluates the result.

It can ask:

  • Did my action succeed?
  • Did I receive useful information?
  • Am I closer to the objective?
  • Should I investigate something else?
  • Should I change my approach?
  • Is there enough evidence to continue?
  • Should I stop?

The source material identifies reflection as a self-assessment step after each action.

For SD-WAN troubleshooting, this can prevent the agent from blindly following a predefined checklist.

For example:

Check WAN transport

Result:

Both transports are operational.

The agent should not necessarily conclude:

"WAN is healthy, therefore WAN is not the problem."

Instead, it can reflect:

"Transport availability is normal. I should investigate latency, packet loss, application path selection, or policy behavior."

It then performs the next appropriate investigation.

5. Cisco SD-WAN Troubleshooting Example

Let's consider a practical example.

A branch office reports:

"The corporate ERP application is very slow."

The AI agent receives the objective:

Investigate the cause of ERP performance degradation.

The agent starts investigating.

Step 1 – Check Device Health

The agent checks the branch SD-WAN device.

Result: Device is operational.

The agent continues.

Step 2 – Check WAN Transports

The agent checks the available WAN transports.

Result: Both transports are available.

The agent continues.

Step 3 – Check Performance

The agent checks latency, jitter, and packet loss.

Result: Transport A has elevated latency and packet loss.

Step 4 – Reflect

The agent evaluates the result:

"Transport A is degraded. I need to determine whether the affected application is using this transport."

Step 5 – Check Application Path

The agent investigates the path being used by the application.

Result: ERP traffic is using Transport A.

Step 6 – Reflect Again

The agent now has stronger evidence:

"The application is using a transport that is experiencing performance degradation."

The agent can now investigate whether an alternate transport is available and whether the configured policy allows the traffic to use it.

This is a practical example of:

Think → Act → Observe → Reflect → Act Again

6. Agent Loop vs Traditional SD-WAN Automation

This distinction is important for network engineers.

Traditional AutomationAgentic Automation
Predefined workflowGoal-oriented workflow
Fixed sequenceDynamic next step
Engineer defines each stepAgent can select the next investigation
Usually predictableContext-dependent
Rule-based decisionsObservation-based decisions
Limited decision-makingReflection after actions
Fixed stopping pointDynamic termination

Traditional automation remains extremely useful.

For repetitive and predictable tasks, a script or predefined workflow may actually be preferable.

Agentic AI becomes more interesting when the problem requires investigation and the next step depends on what was discovered previously.

7. The Most Important Question: When Should the Agent Stop?

An autonomous system should not continue forever.

Imagine an agent repeatedly checking the same WAN path:

Check Path
   ↓
Timeout
   ↓
Check Path
   ↓
Timeout
   ↓
Check Path
   ↓
Timeout
   ↓
Repeat...

This wastes resources and does not solve the problem.

Every agent therefore needs termination conditions.

The source material identifies several common termination triggers, including task completion, maximum iterations, confidence thresholds, human handoff, and error states.

For network automation, these termination conditions are extremely important.

8. Termination Condition #1 – Task Completed

The simplest condition is:

The objective has been achieved.

For example, the agent was asked to investigate why ERP traffic is slow.

After investigation, it identifies the degraded transport and the application path.

If the problem has been safely resolved and verified, the agent can stop.

A final response could be:

"The application was using a WAN transport experiencing packet loss. Traffic was moved to the healthy path and application performance was verified."

The task is complete.

9. Termination Condition #2 – Maximum Iterations

Every autonomous troubleshooting workflow should have a maximum number of attempts.

For example:

Maximum iterations = 10

If the agent reaches ten meaningful actions without reaching a conclusion, it stops.

This prevents an uncontrolled troubleshooting loop.

The source material describes maximum iterations as a safety mechanism that prevents an agent from continuing indefinitely.

A good response could be:

"I could not determine the root cause within the configured investigation limit. Human investigation is recommended."

10. Termination Condition #3 – Confidence Threshold

An agent may also evaluate its confidence in a particular diagnosis.

For example:

"There is strong evidence that Transport A is responsible for the application performance problem."

However, network engineers should be careful with confidence scores.

An AI system can be confident and still be wrong.

Therefore:

AI confidence should support a decision, not replace engineering validation.

The source material specifically highlights the risk of agent overconfidence and recommends human validation for critical actions.

11. Termination Condition #4 – Human Handoff

Sometimes the agent identifies the likely solution but should not execute it automatically.

For example:

"Transport A is degraded and ERP traffic is using it. I recommend changing the traffic policy to prefer Transport B."

At this point, the agent can request approval.

The workflow becomes:

AI Investigation
       ↓
Diagnosis
       ↓
Recommended Change
       ↓
Human Approval
       ↓
Execute
       ↓
Verify

This is called Human-in-the-Loop, or HITL.

The source material describes HITL as a mechanism where the agent pauses and requests human input when the action is high-stakes or outside its authority.

12. Termination Condition #5 – Error State

Sometimes the problem isn't the network itself.

The management system or required tool may be unavailable.

For example:

Agent
 ↓
Query Management Platform
 ↓
API Unavailable
 ↓
Retry
 ↓
API Unavailable
 ↓
Retry Limit Reached
 ↓
Stop

The agent should report the issue instead of pretending that the investigation was completed.

The source material describes an error state as a situation where a required system is unreachable, a tool repeatedly fails, or a dependency does not respond.

13. What Happens When an AI Agent Gets Stuck?

An AI agent can get stuck even when the network itself is operational.

The source material identifies four important stuck states:

  1. Loop
  2. Oscillation
  3. Dead End
  4. Hallucination

Let's translate these concepts into SD-WAN troubleshooting.

14. Stuck State #1 – Loop

A loop occurs when the agent repeatedly performs the same action without making progress.

For example:

Check Tunnel
   ↓
Timeout
   ↓
Check Tunnel
   ↓
Timeout
   ↓
Check Tunnel
   ↓
Timeout

The agent is not learning anything new.

A properly designed system should detect the repeated operation and either try an approved alternative, escalate, or stop.

15. Stuck State #2 – Oscillation

Oscillation occurs when the agent keeps switching between approaches without making progress.

For example:

Check Transport A
       ↓
No conclusion
       ↓
Check Transport B
       ↓
No conclusion
       ↓
Check Transport A
       ↓
Check Transport B
       ↓
Repeat

The agent is changing tactics but not actually moving closer to the objective.

Action-history tracking can help identify this behavior.

16. Stuck State #3 – Dead End

A dead end occurs when all useful investigation paths fail.

For example:

  • Management API unavailable
  • Device access unavailable
  • Performance data unavailable
  • Monitoring data unavailable

The agent has no reliable path forward.

At this point, the safest action is to stop and escalate to a network engineer.

17. Stuck State #4 – Hallucination

Hallucination is particularly important when AI is connected to network automation tools.

An AI agent could potentially assume that a tool or API exists when it does not.

For example, it might attempt to call an invalid API endpoint or provide an incorrect input format.

The source material describes hallucination as attempting to call tools that do not exist or passing invalid inputs.

This is why tool validation and restricted permissions are important when building AI-powered network automation.

18. How Can We Prevent an Agent From Getting Stuck?

The source material recommends a layered approach that combines iteration limits, action-history tracking, fallback strategies, and human escalation.

For SD-WAN automation, we can apply the same principles.

18.1 Set Iteration Limits

Define the maximum number of actions the agent can perform.

18.2 Track Action History

Record what the agent has already checked.

For example:

Action 1: Check device health
Action 2: Check transport
Action 3: Check latency
Action 4: Check application path

The agent can use this information to avoid repeating the same investigation.

18.3 Detect Repeated Actions

If the same operation repeatedly produces the same result, the agent should recognize the pattern.

18.4 Provide Fallback Strategies

If the primary information source is unavailable, use an approved alternative.

For example:

Primary:
Management API

Fallback:
Approved telemetry source

Fallback:
Monitoring platform

Final:
Human escalation

18.5 Escalate to a Human

If all approved automated strategies are exhausted, the agent should involve a network engineer.

The objective isn't to keep the agent autonomous at all costs.

The objective is to solve the problem safely.

19. Human-in-the-Loop in Cisco SD-WAN

HITL becomes particularly important when an agent wants to modify production network behavior.

Imagine the agent determines:

"Transport A has high packet loss. ERP traffic would perform better over Transport B."

The agent can recommend a policy change.

Instead of immediately making the change:

AI Agent
   ↓
Analyze
   ↓
Recommend Change
   ↓
HITL Checkpoint
   ↓
Engineer Approval
   ↓
Execute
   ↓
Verify

The engineer can:

Approve → Agent executes the approved change.

Reject → Agent stops.

Modify → Engineer adjusts the proposed action.

The source material identifies these approval outcomes and emphasizes HITL for important production decisions.

20. When Should Human Approval Be Required?

For SD-WAN automation, human approval can be appropriate for actions such as:

  • Changing traffic policies
  • Modifying routing behavior
  • Changing path preferences
  • Modifying security policies
  • Changing production configuration
  • Actions affecting critical applications
  • Actions outside an approved change window

The source material identifies destructive actions, sensitive operations, uncertainty, policy requirements, and learning-mode deployments as situations where HITL may be required.

A good principle is:

Let AI investigate, but make production changes subject to appropriate controls.

21. Example: AI Agent Troubleshooting a Cisco SD-WAN Branch

Let's put everything together.

Imagine a branch office is experiencing poor ERP application performance.

Objective

"Investigate ERP performance from the branch."

Step 1 – Check Device Health

Result:

Device operational.

Step 2 – Check WAN Transports

Result:

Both transports are available.

Step 3 – Check Performance

Result:

Transport A has elevated packet loss and latency.

Step 4 – Reflect

The agent determines:

"Transport A is degraded. I need to determine whether ERP traffic is using this path."

Step 5 – Check Application Path

Result:

ERP traffic is using Transport A.

Step 6 – Reflect

The agent concludes:

"The degraded transport may be contributing to the application performance problem."

Step 7 – Check Alternate Path

Result:

Transport B is healthy and available.

Step 8 – Generate Recommendation

"Recommend moving ERP traffic to Transport B."

Step 9 – Human Approval

The agent waits for the network engineer.

Step 10 – Execute

The engineer approves the change.

The agent performs the approved action.

Step 11 – Verify

The agent checks the application performance again.

Step 12 – Terminate

The issue is resolved.

This is a complete Agent Loop:

Think → Act → Observe → Reflect → Act → Verify → Stop

22. What Should the Agent Log?

For production SD-WAN environments, logging is extremely important.

A useful agent trace might look like:

Incident ID: INC-10234

09:10 - Task received
09:11 - Branch health checked
09:12 - WAN transports checked
09:13 - Transport A degradation detected
09:14 - Application path analyzed
09:15 - ERP traffic found on Transport A
09:16 - Alternate transport validated
09:17 - Recommendation generated
09:20 - Human approval received
09:21 - Approved action executed
09:22 - Application performance verified
09:23 - Task completed

This creates an audit trail.

If something goes wrong, the engineer can understand:

  • What the agent attempted
  • What it observed
  • Why it made the decision
  • What action was performed
  • Why the agent stopped

The source material emphasizes that termination reasons should be recorded because silent failures make diagnosis and accountability difficult.

23. Why Network Engineers Should Care About Agentic AI

SD-WAN already provides significant automation.

Agentic AI potentially adds another layer.

Traditional SD-WAN

Provides automated connectivity, routing, policy enforcement, and path selection.

AI-Assisted SD-WAN

Helps engineers understand network behavior and investigate problems.

Agentic SD-WAN

Could potentially investigate a problem, collect evidence, determine the next troubleshooting step, recommend remediation, and execute approved actions.

This is why network engineers should understand Agentic AI.

You don't necessarily need to become an AI researcher.

But understanding concepts such as:

  • AI agents
  • Agent loops
  • APIs
  • Tool calling
  • Reflection
  • Termination
  • Human-in-the-Loop
  • Guardrails
  • Automation
  • Observability

can become increasingly valuable.

24. Automation vs Agentic Automation

A simple way to remember the difference is:

Traditional Automation

IF condition
     ↓
Perform predefined action

Agentic Automation

Goal
 ↓
Observe
 ↓
Reason
 ↓
Act
 ↓
Evaluate
 ↓
Select Next Action
 ↓
Repeat

Safe Agentic Automation

Goal
 ↓
Observe
 ↓
Reason
 ↓
Act
 ↓
Evaluate
 ↓
Safety Check
 ↓
Human Approval if Required
 ↓
Verify
 ↓
Terminate

For production networking, the third model is the most useful mental model.

25. Cisco SD-WAN + AI: A Possible Architecture

A conceptual architecture could look like this:

                 Network Engineer
                       │
                       ↓
                  AI Agent
                       │
                       ↓
              Reasoning / Planning
                       │
                       ↓
                Safety Layer
                       │
             ┌─────────┴─────────┐
             ↓                   ↓
       Read Operations      Change Operations
             │                   │
             ↓                   ↓
     SD-WAN Information     HITL Approval
             │                   │
             └─────────┬─────────┘
                       ↓
                  SD-WAN Network
                       │
                       ↓
                Telemetry / Data
                       │
                       └────→ AI Agent

The key component is the Safety Layer.

An AI agent should not have unrestricted access to a production network.

26. Read-Only vs Change-Enabled AI Agents

Network teams can approach AI adoption gradually.

Level 1 – AI Assistant

The engineer asks:

"Why is this SD-WAN path being selected?"

The AI provides an explanation.

Level 2 – Read-Only Agent

The agent can collect operational information such as:

  • Device health
  • Transport status
  • Performance
  • Routing information

But it cannot modify the network.

Level 3 – Investigation Agent

The agent investigates an incident and provides a likely root cause.

For example:

"The application is using a transport experiencing high packet loss."

Level 4 – Recommendation Agent

The agent proposes remediation.

For example:

"Recommend moving the application to the alternate transport."

Level 5 – Human-Approved Automation

The agent proposes the change.

The engineer approves it.

The agent executes it.

Level 6 – Controlled Autonomous Operations

Selected low-risk actions can potentially be automated under strict policies and guardrails.

This should only be considered after extensive testing, monitoring, validation, and governance.

27. Safety: Never Trust an AI Agent Blindly

One of the most important lessons from Agentic AI is that confidence does not automatically mean correctness.

An AI agent can produce an incorrect conclusion with high confidence.

The source material highlights several safety considerations, including silent failures, overconfidence, context loss during human handoffs, and automation bias.

For network engineers, remember:

AI recommendation ≠ Network Truth

Critical conclusions should be validated.

For production changes, consider:

  • Change approval
  • RBAC
  • Audit logging
  • Rollback capability
  • Iteration limits
  • Tool restrictions
  • Human approval
  • Post-change verification

28. The Importance of Verification

The agent should not simply perform an action and assume that the problem is solved.

For example:

Agent: Move traffic to Transport B.

That is not the end of the workflow.

The agent should verify:

  • Is traffic now using the expected path?
  • Has packet loss improved?
  • Has latency improved?
  • Is the application responding normally?
  • Did the change introduce another problem?

The complete workflow should therefore be:

Diagnose → Change → Verify → Terminate

This is a familiar concept to network engineers.

After all, we rarely make a production change without checking whether the expected result occurred.

29. The Golden Rules for AI-Based SD-WAN Automation

If you are designing an AI agent for SD-WAN operations, remember these principles:

Rule 1 – Define the Goal

The agent needs a clear objective.

Rule 2 – Limit Iterations

Never allow unlimited execution.

Rule 3 – Track History

Know what the agent has already attempted.

Rule 4 – Detect Loops

Repeated actions should trigger investigation or escalation.

Rule 5 – Provide Fallback Paths

Don't depend on one data source.

Rule 6 – Use Human Approval

High-impact production changes should have appropriate approval.

Rule 7 – Log Important Actions

Maintain a useful audit trail.

Rule 8 – Don't Trust AI Confidence Blindly

Validate important conclusions.

Rule 9 – Verify the Result

After an action, confirm that the expected outcome occurred.

Rule 10 – Define When to Stop

Every agent needs a clear termination strategy.

30. How This Applies to FortiGate SD-WAN

The same Agent Loop methodology can also be applied to FortiGate SD-WAN environments.

The vendor terminology, management interfaces, APIs, and available operational data will be different, but the basic methodology remains the same.

For example:

Application Performance Problem
            ↓
       Understand Goal
            ↓
      Check SD-WAN Health
            ↓
       Observe SLA Status
            ↓
       Analyze Path Choice
            ↓
          Reflect
            ↓
     Identify Likely Cause
            ↓
    Recommend Remediation
            ↓
       Human Approval
            ↓
     Apply Approved Change
            ↓
        Verify Result
            ↓
          Stop

This demonstrates an important point:

The Agent Loop is a methodology rather than a vendor-specific feature.

It can be applied to Cisco SD-WAN, FortiGate SD-WAN, Cisco ACI, and other network platforms when appropriate data and controlled automation interfaces are available.

31. A Simple Mental Model for Network Engineers

If Agentic AI terminology seems complicated, think about how you already troubleshoot a network.

When an application is slow, you don't normally execute commands randomly.

You:

  1. Understand the problem.
  2. Form a hypothesis.
  3. Collect information.
  4. Analyze the result.
  5. Change the hypothesis if necessary.
  6. Investigate another area.
  7. Identify the likely cause.
  8. Apply a controlled fix.
  9. Verify the result.
  10. Close the incident.

This is remarkably similar to an Agent Loop.

The major difference is that an AI agent attempts to automate portions of this reasoning and execution process.

32. Final Thoughts

SD-WAN has already changed the way organizations design and operate WAN infrastructure.

The next evolution could be the combination of:

SD-WAN + APIs + Automation + AI Agents

For network engineers, the most important concept is not simply learning how to build an AI agent.

It is learning how to build an agent that operates safely.

Remember the complete loop:

Think → Act → Observe → Reflect → Decide → Repeat or Stop

Then add the production safeguards:

Iteration Limits + Action History + Fallback Strategies + Logging + Guardrails + Human-in-the-Loop

This combination can turn an AI agent from an interesting experiment into a potentially useful network operations assistant.

The future network engineer may not be the person who manually performs every troubleshooting step.

Instead, it may be the engineer who understands:

What should be automated, what should be delegated to AI, what must be validated, and what should always require human approval.

That is where Agentic AI becomes particularly interesting for Cisco SD-WAN and modern network operations.

33. Frequently Asked Questions

What is an Agent Loop in networking?

An Agent Loop is a process where an AI agent investigates a network objective, performs actions, observes results, reflects on the findings, and decides whether to continue or stop.

Can AI troubleshoot Cisco SD-WAN?

AI agents can potentially assist with SD-WAN troubleshooting by collecting operational information, analyzing performance data, identifying patterns, and recommending the next troubleshooting steps.

Can an AI agent automatically change SD-WAN policies?

An agent can potentially be connected to automation interfaces capable of making configuration changes. However, production changes should be protected by authorization, policy controls, validation, and appropriate human approval.

What is Human-in-the-Loop in SD-WAN?

Human-in-the-Loop means that the AI agent pauses and requests an engineer's approval before performing a decision or action that requires human judgment or carries significant operational risk.

What happens if an AI network agent gets stuck?

The agent should detect repeated actions, oscillation, dead ends, or tool failures and either use an approved fallback strategy, escalate to a human, or terminate.

Is Agentic AI the same as network automation?

No. Traditional automation generally follows predefined workflows, while an AI agent can evaluate observations and dynamically determine the next step within its defined boundaries.

34. Related Articles From Netterrene

Generative AI Fundamentals for Beginners
https://netterrene.blogspot.com/2026/06/generative-ai-fundamentals-for-beginners.html

Agentic AI for Network Engineers
https://netterrene.blogspot.com/2026/07/agentic-ai-for-network-engineers-guide.html

Cisco ACI Explained
https://netterrene.blogspot.com/2026/04/cisco-aci-explained-concepts-learning.html

Cisco ACI vPC Explained
https://netterrene.blogspot.com/2026/06/cisco-aci-vpc-explained-architecture_01563772518.html

Cisco ACI Service Graphs
https://netterrene.blogspot.com/2026/05/why-service-graphs-matter-in-cisco-aci.html

Understanding L3Out Subnet Scope Options in Cisco ACI
https://netterrene.blogspot.com/2025/08/l3out-subnet-scope-options-in-cisco-aci.html