Showing posts with label Artificial Intelligence. Show all posts
Showing posts with label Artificial Intelligence. Show all posts

Friday, 14 August 2026

Top 9 Agentic AI Multiple Choice Questions (MCQs) with Answers and Explanations

 Agentic AI is transforming how organizations automate operations, perform reasoning, and execute actions autonomously. The following questions cover key concepts including agent architecture, governance, frameworks, automation platforms, and AI risks.


Question 1

A security team is deploying an agent that can quarantine potentially infected hosts. Which human-in-the-loop pattern is MOST appropriate?

Options

A. Notification only - the agent acts and informs the team afterward

B. Full automation - the agent acts without any human involvement

C. No agent involvement - humans should handle all security decisions manually

D. Approval required - a human must confirm before the agent quarantines a host

Correct Answer

✅ D. Approval required - a human must confirm before the agent quarantines a host

Explanation

Quarantining a host can disrupt business services and impact users. Because of the potentially significant consequences, human review should occur before the action is executed.

Human approval provides:

  • Risk mitigation
  • Better decision accuracy
  • Reduced operational disruption
  • Governance and accountability

Question 2

A network operations team needs to build an agent that monitors alerts, queries a knowledge base, and creates tickets. The team includes NOC analysts with limited Python experience, and they need the solution within a week. Which approach is MOST appropriate?

Options

A. n8n or similar low-code platform for rapid development

B. Custom C++ implementation for performance

C. Wait to hire a Python developer

D. Python with LangChain for maximum flexibility

Correct Answer

✅ A. n8n or similar low-code platform for rapid development

Explanation

The requirements emphasize:

  • Rapid implementation
  • Low coding complexity
  • Limited Python expertise
  • Workflow automation

Low-code tools such as n8n offer visual workflow designers and pre-built integrations, making them ideal for fast delivery.


Question 3

Match each ecosystem component to its role in agent architecture.

Components

  • Vector Database
  • Model Provider (LLM)
  • Code Sandbox
  • LangChain

Correct Matching

ComponentRole
Vector DatabaseProvides memory
Model Provider (LLM)Provides the reasoning core
Code SandboxEnables safe code execution
LangChainOrchestrates

Explanation

Vector Database Stores embeddings and supports semantic retrieval, acting as long-term memory.

Model Provider (LLM) Performs reasoning, understanding, and response generation.

Code Sandbox Allows secure execution of generated code.

LangChain Coordinates interactions among models, tools, databases, and workflows.


Question 4

Which of the following BEST describes an AI agent?

Options

A. A machine learning model that generates text responses

B. A script that automates repetitive tasks based on schedules

C. An autonomous system that perceives, reasons, acts, and learns from outcomes

D. A chatbot that responds to user queries using a knowledge base

Correct Answer

✅ C. An autonomous system that perceives, reasons, acts, and learns from outcomes

Explanation

An AI agent typically:

  • Perceives information
  • Reasons about data
  • Takes actions
  • Learns from results

Unlike basic chatbots or scripts, agents pursue goals with varying degrees of autonomy.


Question 5

Which risk category is BEST described by the following scenario?

"An agent confidently recommends a network configuration change based on incorrect information it generated."

Options

A. Operational risk - cost overrun

B. Reliability risk - hallucination

C. Governance risk - unexplainability

D. Security risk - unauthorized access

Correct Answer

✅ B. Reliability risk - hallucination

Explanation

This represents an AI hallucination where the model generates incorrect information while appearing confident.

Reliability risks include:

  • Hallucinations
  • Incorrect recommendations
  • Inaccurate outputs
  • Poor decision quality

Organizations commonly reduce this risk through validation, retrieval systems, and human review.


Question 6

Which characteristic distinguishes agentic AI from generative AI like ChatGPT?

Options

A. The ability to process natural language input

B. The capacity to maintain conversation context

C. The use of large language models for reasoning

D. The capability to take autonomous actions that affect the environment

Correct Answer

✅ D. The capability to take autonomous actions that affect the environment

Explanation

Generative AI primarily creates content.

Agentic AI goes further by:

  • Planning
  • Making decisions
  • Using tools
  • Executing actions
  • Pursuing goals autonomously

A useful distinction is:

Generative AI creates. Agentic AI creates and acts.


Question 7

A system receives customer support tickets, analyzes their content, creates draft responses, routes tickets to appropriate teams, monitors resolution time, and escalates overdue tickets automatically. This system is BEST classified as:

Options

A. Generative AI

B. Agentic AI

C. Reactive automation

D. Rule-based workflow

Correct Answer

✅ B. Agentic AI

Explanation

The system:

  • Analyzes information
  • Generates responses
  • Makes routing decisions
  • Monitors progress
  • Escalates issues automatically

These behaviors demonstrate goal-oriented autonomy, making it an example of Agentic AI.


Question 8

Which of the following is a key advantage of n8n for building agentic workflows?

Options

A. It provides 400+ pre-built integrations and self-hosting capability

B. It offers better performance than any coded solution

C. It requires extensive Python programming knowledge

D. It must be deployed only in public cloud environments

Correct Answer

✅ A. It provides 400+ pre-built integrations and self-hosting capability

Explanation

n8n is widely adopted because it offers:

  • Large integration ecosystem
  • Visual workflow design
  • Self-hosting support
  • Fast deployment
  • Low-code development

These features make it attractive for rapid AI automation projects.


Question 9

DevOps team needs to build an agent that maintains conversation state across sessions, supports human-in-the-loop interruption for approval steps, and can checkpoint and resume workflows. Which framework should they choose?

Options

A. CrewAI - for role-based team coordination

B. LangChain - for its large ecosystem

C. LangGraph - for explicit state management and checkpointing

D. AutoGen - for multi-agent conversations

Correct Answer

✅ C. LangGraph - for explicit state management and checkpointing

Explanation

LangGraph is specifically designed for:

  • Stateful workflows
  • Durable execution
  • Human-in-the-loop operations
  • Checkpointing
  • Workflow recovery
  • Long-running agent processes

These capabilities make it ideal for enterprise DevOps scenarios.


Key Takeaways

Agent Architecture

  • LLM = Reasoning
  • Vector Database = Memory
  • Sandbox = Safe Execution
  • LangChain = Orchestration

Governance

  • High-risk actions should use Human-in-the-Loop approvals.
  • Hallucinations are Reliability Risks.

Platforms and Frameworks

  • n8n = Rapid low-code automation
  • LangChain = Ecosystem and integrations
  • LangGraph = Stateful workflows and checkpointing
  • CrewAI = Agent teams
  • AutoGen = Multi-agent conversations

Agentic AI

Agentic AI extends generative AI by enabling autonomous decision-making and action execution to achieve objectives.

Conclusion

Understanding AI agents, agent architecture, governance models, and frameworks such as LangGraph and LangChain is essential for designing reliable enterprise AI solutions. These MCQs provide a strong foundation for certification preparation, interviews, and practical implementation of agentic AI systems.

Tuesday, 11 August 2026

Multi-Agent AI Systems Explained for Cisco SD-WAN Engineers: Supervisor, Peer-to-Peer, and Hierarchical Patterns

Multi-Agent AI Systems Explained for Cisco SD-WAN Engineers: Supervisor, Peer-to-Peer, and Hierarchical Patterns

One AI agent can only be so good at everything. Ask it to inventory branches in a region, check tunnel health, correlate an SLA breach, and draft a fix — and you're asking one generalist to do the job of four specialists. On a large SD-WAN deployment spanning dozens of branches and multiple regions, that generalist starts to strain.

That's the problem multi-agent systems solve: instead of one agent doing everything, you split the work across several agents, each good at one thing, coordinated in a specific pattern. This post walks through the three orchestration patterns you'll see in agentic AI tooling — Supervisor, Peer-to-Peer, and Hierarchical — each mapped to a real SD-WAN operational scenario.


When Does an SD-WAN Team Actually Need Multiple Agents?

Before reaching for a multi-agent design, it's worth being honest about when it earns its complexity:

  • Complex workflows. Investigating an overlay issue genuinely needs different skills — checking tunnel/BFD state, reading SLA performance, interpreting centralized policy. One agent trying to be great at all three ends up mediocre at each.
  • Parallel work. Checking three branches at once, or querying both the vManage API and a third-party ISP status page simultaneously, is naturally parallel — a single sequential agent just makes you wait longer.
  • Specialization. A "policy analysis" agent can be tuned with SD-WAN-specific prompts and tools that would just be noise for a "tunnel status" agent.
  • Separation of concerns. Smaller, focused agents are easier to test and debug than one do-everything agent whose failures could be coming from anywhere.

The catch: more agents means more coordination overhead, more places for something to go wrong, and a harder system to debug. If a single agent with a couple of tools can already answer the question, that's the better answer. Start simple. Add agents only when a single agent hits its limits.


Pattern 1: Supervisor — One Coordinator, Specialized Workers

This is the most common pattern, and usually the right starting point. One supervisor agent receives the request, decides what needs to happen, and delegates to specialized worker agents — then assembles their results into a final answer.



SD-WAN scenario: "Why is voice choppy at the Chicago branch?"

  1. The Supervisor decides this needs tunnel status, SLA performance, and a policy check — in that order.
  2. → Tunnel Agent: checks BFD and transport state for the Chicago branch's tunnels. ← Returns: MPLS tunnel is up; Internet tunnel flapped twice in the last hour.
  3. → SLA Agent: pulls loss/latency/jitter for the voice traffic class on both tunnels. ← Returns: Internet tunnel jitter spiked to 45ms during the flaps.
  4. → Policy Agent: checks whether application-aware routing actually steered voice traffic off the degraded tunnel. ← Returns: policy didn't fail over — the SLA class threshold was set too loose to trigger a switch.
  5. Supervisor assembles the answer: "Voice quality at Chicago traces to Internet tunnel jitter during BFD flaps — the SLA class threshold needs tightening so app-aware routing fails over sooner."

When this works well: You want central control, a clean audit trail, and predictable delegation — which matters a lot in change-managed SD-WAN environments where you need to show exactly which check ran and in what order.

The trade-off: the Supervisor is a single point of failure and can become a bottleneck if too much routing logic gets crammed into it.

Other SD-WAN Supervisor use cases:

  • A single "overlay health assistant" that routes questions to a tunnel-status agent, an SLA-summary agent, or a capacity-planning agent depending on what's asked
  • Coordinating a fleet-wide health check across a Tunnel Agent, Control-Connection Agent, and Software-Version Agent before a maintenance window

Pattern 2: Peer-to-Peer — No Coordinator, Agents Consult Each Other

Here there's no central agent directing traffic. Any agent can talk to any other agent directly, and the answer emerges from their back-and-forth — closer to a group of specialists in a room than a chain of command.



SD-WAN scenario: App performance complaints from three branches in the same region.

[SLA Agent]: "Seeing loss spikes on the Internet transport class across three branches — anyone have context?"

[Tunnel Agent] responds: "All three branches are homed to the same regional hub. BFD sessions are stable, no flaps."

[Transport Agent] chimes in: "Checking underlying ISP health… this looks like a regional ISP brownout, not a device issue."

[Policy Agent] adds: "App-aware routing is trying to fail over, but the MPLS backup path is also showing elevated latency at the same hub."

[Transport Agent] concludes: "Confirmed — both paths converge at the same regional hub, so the brownout is affecting both transports simultaneously."

Agents collectively surface: "Regional ISP brownout at the hub is degrading both transport paths for all three branches — this isn't a config issue."

When this works well: genuine collaboration where no agent has the full picture alone, and where resilience matters — if the Policy Agent is unavailable, the other three can still reach a conclusion.

The trade-off: harder to debug. There's no single transcript to read top-to-bottom — the reasoning is scattered across several conversations.

Other SD-WAN Peer-to-Peer use cases:

  • Multi-region troubleshooting where a Latency Agent, a Routing Agent, and a Hub Agent need to jointly rule causes in or out
  • A design-review "roundtable" where a Security Agent, Capacity Agent, and Topology Agent debate a proposed transport redesign before it's finalized

Pattern 3: Hierarchical — Layered Control for Large Deployments

This mirrors an org chart. An executive agent sets overlay-wide strategy and delegates to leads — one per region, or one per hub — and each lead manages its own workers. Results roll back up the chain.


SD-WAN scenario: "Is this a fleet-wide slowdown, or just one region?"

  • [Executive Agent]: "Delegate a health check to each regional lead."
    • → [Region 1 Lead]
      • → Research Worker: inventories Region 1 branches → 12 branches, 2 hubs
      • → Diagnostics Worker: checks tunnel health → 3 branches showing elevated jitter
      • ← Region 1 Lead reports: "Localized degradation on 3 branches, all homed to Hub-East."
    • → [Region 2 Lead]
      • → Research Worker: inventories Region 2 branches → 9 branches, 1 hub
      • → Diagnostics Worker: checks tunnel health → all normal
      • ← Region 2 Lead reports: "No issues."
  • [Executive Agent] aggregates: "Fleet-wide slowdown is actually isolated to Region 1, specifically branches behind Hub-East. Region 2 is healthy."

When this works well: your SD-WAN environment is already organized this way — multiple regions, multiple hubs, regional teams each owning their own branches — so the agent hierarchy just mirrors structure you already have.

The trade-off: more layers means more latency. A question has to travel down through leads to workers and back up again before you get an answer.

Other SD-WAN Hierarchical use cases:

  • A global deployment where each region has its own lead agent managing local hub-level workers, rolling up to a global orchestrator
  • Change-approval workflows where a centralized policy push has to pass through a Regional lead and then a Global lead before being approved — naturally mapping to an approval chain that mirrors your org structure

Choosing a Pattern for Your SD-WAN Use Case

If you need…ChooseWhy
Simple delegation with clear, well-defined tasksSupervisorMost common pattern — straightforward and predictable
A workflow spanning multiple regions or hubsHierarchicalMirrors how large deployments and teams are already organized
Flexible collaboration where agents genuinely need to consult each otherPeer-to-PeerBest for ambiguous, multi-cause investigations — but harder to debug
A clear audit trail for change management or complianceSupervisorCentral control point makes the decision path traceable
High resilience with no single point of failurePeer-to-PeerThe investigation continues even if one agent is unavailable

Complexity Trade-offs: Single Agent vs. Multi-Agent

FactorSingle agentMulti-agent
SimplicitySimpleComplex
DebuggingEasyHarder
LatencyLowerHigher
CostLowerHigher

The rule that matters most: start simple, and only add agents once a single agent actually hits its limits. Most day-to-day SD-WAN questions — "what's this tunnel's BFD state," "what transport is this app using right now" — don't need a multi-agent system at all. Reach for one when the task genuinely spans multiple domains of expertise, the way a real fleet-wide incident does.


Final Thoughts

None of these patterns are about the AI being "smarter" — they're about matching the coordination structure to the shape of the problem. A quick lookup doesn't need a supervisor and three workers. A fleet-wide incident spanning multiple regions might genuinely benefit from one. The skill isn't picking the most sophisticated pattern — it's picking the one that fits the blast radius and complexity of what you're actually investigating, the same instinct that makes a good SD-WAN engineer good at escalation and delegation in the first place.


FAQ

Q: Should every SD-WAN troubleshooting AI system use multiple agents? A: No. Most single-tunnel or single-branch questions are handled better and faster by one agent with the right tools. Multi-agent systems earn their overhead on genuinely multi-domain problems — multi-region investigations, fleet-wide incidents, or workflows that already mirror an organizational hierarchy.

Q: Which pattern gives the clearest audit trail for change management? A: Supervisor. Because one agent owns delegation and assembles the final answer, there's a single, traceable decision path — useful when you need to show exactly what was checked and in what order before a centralized policy change.

Q: Is Peer-to-Peer riskier to run against production vManage? A: Not inherently riskier in terms of what it does, but harder to review before the fact, since there's no single plan to inspect — the reasoning is distributed across several agent-to-agent exchanges. Pair it with the same execution guardrails (preview mode, confirmation, scoped tools) you'd use for any other agent design.


Related Reading on Networklearner:


Need help with SD-WAN, Cisco ACI, Nexus, data center networking, or network automation?

I am a CCIE Data Center engineer with 18+ years of enterprise networking experience, working hands-on with production SD-WAN and ACI environments.

Contact me for consulting, troubleshooting, design reviews, and project support: rockingoa@gmail.com

Multi-Agent AI Systems Explained for Cisco ACI Engineers: Supervisor, Peer-to-Peer, and Hierarchical Patterns

 

Multi-Agent AI Systems Explained for Cisco ACI Engineers: Supervisor, Peer-to-Peer, and Hierarchical Patterns

One AI agent can only be so good at everything. Ask it to inventory devices in a Pod, check health scores, correlate a traffic anomaly, and draft a fix — and you're asking one generalist to do the job of four specialists. On a large ACI environment with multiple Pods, tenants, and Sites, that generalist starts to strain.

That's the problem multi-agent systems solve: instead of one agent doing everything, you split the work across several agents, each good at one thing, coordinated in a specific pattern. This post walks through the three orchestration patterns you'll see in agentic AI tooling — Supervisor, Peer-to-Peer, and Hierarchical — each mapped to a real ACI operational scenario.


When Does an ACI Team Actually Need Multiple Agents?

Before reaching for a multi-agent design, it's worth being honest about when it earns its complexity:

  • Complex workflows. Investigating a fabric issue genuinely needs different skills — pulling topology, reading health scores, interpreting zoning-rule programming. One agent trying to be great at all three ends up mediocre at each.
  • Parallel work. Checking three leaf switches at once, or querying both the APIC fault database and the Nexus Dashboard Insights API simultaneously, is naturally parallel — a single sequential agent just makes you wait longer.
  • Specialization. A "contract analysis" agent can be tuned with ACI-specific prompts and tools that would just be noise for a "topology lookup" agent.
  • Separation of concerns. Smaller, focused agents are easier to test and debug than one do-everything agent whose failures could be coming from anywhere.

The catch: more agents means more coordination overhead, more places for something to go wrong, and a harder system to debug. If a single agent with a couple of tools can already answer the question, that's the better answer. Start simple. Add agents only when a single agent hits its limits.


Pattern 1: Supervisor — One Coordinator, Specialized Workers

This is the most common pattern, and usually the right starting point. One supervisor agent receives the request, decides what needs to happen, and delegates to specialized worker agents — then assembles their results into a final answer.



ACI scenario: "Why is EPG WEB slow?"

  1. The Supervisor decides this needs topology context, current fault data, and a contract check — in that order.
  2. → Topology Agent: maps EPG WEB to its leaf switches and ports. ← Returns: EPG WEB spans Leaf-101, Leaf-102, and Leaf-104.
  3. → Fault Agent: pulls current faults and health scores for those three leaves. ← Returns: Leaf-104 health score is 68; two contract-related faults are active.
  4. → Contract Agent: checks whether the contract between EPG WEB and EPG DB actually rendered correctly on Leaf-104. ← Returns: zoning-rule programming shows a stale entry from a contract update two days ago.
  5. Supervisor assembles the answer: "EPG WEB's slowness traces to a stale zoning-rule entry on Leaf-104 from the last contract push — a policy resolve/redeploy should clear it."

When this works well: You want central control, a clean audit trail, and predictable delegation — which matters a lot in change-managed ACI environments where you need to show exactly which check ran and in what order.

The trade-off: the Supervisor is a single point of failure and can become a bottleneck if too much routing logic gets crammed into it.

Other ACI Supervisor use cases:

  • A single "fabric health assistant" that routes questions to a health-score agent, a fault-summary agent, or a capacity-planning agent depending on what's asked
  • Coordinating a Multi-Site health check across a Topology Agent, Latency Agent, and Schema-Sync Agent

Pattern 2: Peer-to-Peer — No Coordinator, Agents Consult Each Other

Here there's no central agent directing traffic. Any agent can talk to any other agent directly, and the answer emerges from their back-and-forth — closer to a group of specialists in a room than a chain of command.



ACI scenario: Voice quality complaints from Building A.

[Health Agent]: "Seeing 95% CPU on Leaf-103 in Building A — anyone have context?"

[Topology Agent] responds: "Leaf-103 serves EPG VOICE and EPG WEB. A contract update landed on it about an hour ago."

[Contract Agent] chimes in: "That update added a new filter to the VOICE-to-WAN contract. Checking whether it's over-matching traffic…"

[Security Agent] adds: "No threat signatures on this leaf — this looks operational, not malicious."

[Contract Agent] concludes: "Confirmed — the new filter is unintentionally catching broadcast traffic, consistent with the CPU spike."

Agents collectively surface: "CPU spike on Leaf-103 traced to an over-broad contract filter added in the last update."

When this works well: genuine collaboration where no agent has the full picture alone, and where resilience matters — if the Security Agent is unavailable, the other three can still reach a conclusion.

The trade-off: harder to debug. There's no single transcript to read top-to-bottom — the reasoning is scattered across several conversations.

Other ACI Peer-to-Peer use cases:

  • Multi-Site troubleshooting where a Latency Agent, a Schema Agent, and a Connectivity Agent need to jointly rule causes in or out
  • A design-review "roundtable" where a Security Agent, Capacity Agent, and Topology Agent debate a proposed contract change before it's finalized

Pattern 3: Hierarchical — Layered Control for Large Fabrics

This mirrors an org chart. An executive agent sets fabric-wide strategy and delegates to leads — one per Pod, Site, or domain — and each lead manages its own workers. Results roll back up the chain.



ACI scenario: "Is the whole campus experiencing slowness, or just one Pod?"

  • [Executive Agent]: "Delegate a health check to each Pod lead."
    • → [Pod 1 Lead]
      • → Research Worker: inventories Pod 1 leaves → Leaf-101 through Leaf-104
      • → Diagnostics Worker: checks health scores → Leaf-104 at 68, one active fault
      • ← Pod 1 Lead reports: "Localized issue on Leaf-104."
    • → [Pod 2 Lead]
      • → Research Worker: inventories Pod 2 leaves → Leaf-201, Leaf-202
      • → Diagnostics Worker: checks health scores → all normal
      • ← Pod 2 Lead reports: "No issues."
  • [Executive Agent] aggregates: "Campus slowness is isolated to Pod 1 — specifically Leaf-104. Pod 2 is healthy."

When this works well: your ACI environment is already organized this way — multiple Pods, multiple Sites, regional teams each owning their domain — so the agent hierarchy just mirrors structure you already have.

The trade-off: more layers means more latency. A question has to travel down through leads to workers and back up again before you get an answer.

Other ACI Hierarchical use cases:

  • A Multi-Site environment where each Site has its own lead agent managing local Pod-level workers, rolling up to a global orchestrator
  • Change-approval workflows where a request has to pass through a Tenant-level lead and then a Fabric-level executive before being approved — naturally mapping to an approval chain that mirrors your org structure

Choosing a Pattern for Your ACI Use Case

If you need…ChooseWhy
Simple delegation with clear, well-defined tasksSupervisorMost common pattern — straightforward and predictable
A workflow spanning multiple Pods, Sites, or tenant domainsHierarchicalMirrors how large fabrics and teams are already organized
Flexible collaboration where agents genuinely need to consult each otherPeer-to-PeerBest for ambiguous, multi-cause investigations — but harder to debug
A clear audit trail for change management or complianceSupervisorCentral control point makes the decision path traceable
High resilience with no single point of failurePeer-to-PeerThe investigation continues even if one agent is unavailable

Complexity Trade-offs: Single Agent vs. Multi-Agent

FactorSingle agentMulti-agent
SimplicitySimpleComplex
DebuggingEasyHarder
LatencyLowerHigher
CostLowerHigher

The rule that matters most: start simple, and only add agents once a single agent actually hits its limits. Most day-to-day ACI questions — "what's this leaf's health score," "list active faults on this EPG" — don't need a multi-agent system at all. Reach for one when the task genuinely spans multiple domains of expertise, the way a real fabric-wide incident does.


Final Thoughts

None of these patterns are about the AI being "smarter" — they're about matching the coordination structure to the shape of the problem. A quick lookup doesn't need a supervisor and three workers. A campus-wide incident spanning multiple Pods might genuinely benefit from one. The skill isn't picking the most sophisticated pattern — it's picking the one that fits the blast radius and complexity of what you're actually investigating, the same instinct that makes a good ACI engineer good at escalation and delegation in the first place.


FAQ

Q: Should every ACI troubleshooting AI system use multiple agents? A: No. Most single-EPG or single-device questions are handled better and faster by one agent with the right tools. Multi-agent systems earn their overhead on genuinely multi-domain problems — Multi-Site investigations, campus-wide incidents, or workflows that already mirror an organizational hierarchy.

Q: Which pattern gives the clearest audit trail for change management? A: Supervisor. Because one agent owns delegation and assembles the final answer, there's a single, traceable decision path — useful when you need to show exactly what was checked and in what order.

Q: Is Peer-to-Peer riskier to run against production APIC? A: Not inherently riskier in terms of what it does, but harder to review before the fact, since there's no single plan to inspect — the reasoning is distributed across several agent-to-agent exchanges. Pair it with the same execution guardrails (preview mode, confirmation, scoped tools) you'd use for any other agent design.


Related Reading on Networklearner:


Need help with Cisco ACI, Nexus, data center networking, or network automation?

I am a CCIE Data Center engineer with 18+ years of enterprise networking experience, working hands-on with production ACI fabrics.

Contact me for consulting, troubleshooting, design reviews, and project support: rockingoa@gmail.com

Tuesday, 4 August 2026

How AI Agents Actually "Touch" Your SD-WAN Overlay: Tools, Schemas, and Guardrails Explained

 An AI agent that can only talk isn't much use to an SD-WAN operations team. It can summarize a wall of syslog you paste in, sure — but it can't tell you a branch's current BFD state right now, and it definitely shouldn't be pushing a centralized policy change just because it "reasoned" its way there. What separates a chatbot from something you'd actually let near vManage is tools — the functions an agent is allowed to call to read from, act on, or talk about your overlay.

This post breaks down how tools work in an AI agent, using SD-WAN as the running example throughout, so the concepts map directly onto things you already manage — tunnels, SLA classes, device templates, and centralized policy.


Why "Just Talking" Isn't Enough on SD-WAN

Picture an agent with no tools at all:

Agent: "I think Branch-42's MPLS tunnel might be down, but I have no way to check."

Not useful. Now give it a single tool that can query vManage's device API:

Agent: [calls get_tunnel_status(branch="Branch-42", transport="MPLS")] → "Branch-42's MPLS tunnel is down. BFD lost sync 4 minutes ago; the branch has failed over to Internet transport."

Same question, completely different value. That's the entire point of tools — they turn an agent from something that speculates into something that can actually verify against the overlay's real state.


The Four Tool Categories, Mapped to SD-WAN

Every tool an SD-WAN-aware agent might use falls into one of four buckets. Knowing which bucket a task belongs to tells you immediately how much oversight it needs.

1. Retrieval Tools — Read-Only Lookups

These pull information without changing anything: tunnel status, SLA performance, control-connection state, transport in use.

SD-WAN examples:

  • get_tunnel_status(branch, transport) — pull current BFD/session state for a specific tunnel
  • get_sla_performance(tunnel, app_class) — loss/latency/jitter for a given app class
  • get_control_connections(device) — check a device's control-plane state to vSmart
  • get_active_transport(branch, app) — which underlay path an app is currently steered over

Retrieval tools are the safest category — an agent can call these freely without much risk, which is exactly why they're the easiest place to start trusting AI in SD-WAN operations.

2. Execution Tools — Tools That Change the Overlay

These make actual changes: pushing a centralized data policy, updating a device template, restarting a tunnel, triggering a software upgrade. Because SD-WAN policy is orchestrated centrally, execution tools deserve the most design care of any category — a bad push doesn't stay local, it propagates to every site attached to that policy or template.

SD-WAN examples:

  • deploy_data_policy(site_list, app_class, sla_class, mode)
  • update_device_template(device, template, mode)
  • trigger_software_upgrade(device, target_version, mode)

Notice the repeated mode parameter — more on that below. It's the single most important detail in an SD-WAN execution tool's design.

3. Communication Tools — Looping in Humans

These don't touch the overlay at all — they notify people. Sending a Slack alert about a degraded transport, opening a ServiceNow ticket for a recurring BFD flap, paging the on-call engineer when a regional hub loses reachability.

SD-WAN examples:

  • send_slack_alert(channel, message) — e.g., posting when an SLA class breaches threshold on a tunnel
  • create_servicenow_ticket(summary, severity, affected_site)
  • page_oncall(team, reason) — for something like a vSmart cluster losing quorum

These tools are how an agent stays useful even when it shouldn't act on its own — escalating to a human is often the correct behavior, not a fallback.

4. Perception Tools — Making Sense of Raw Data

These interpret information rather than fetch or change it: parsing a flood of BFD flap events into a plain-English summary, correlating a voice-quality complaint with a recent policy push, summarizing a week of vManage audit logs.

SD-WAN examples:

  • summarize_alarms(site, time_range) — turn 150 raw alarms into three actionable findings
  • correlate_app_degradation(app, time_range) — check if a slowdown lines up with a recent policy or template change
  • parse_audit_log(controller, time_range) — extract what actually changed and who changed it

Perception tools are what let an agent reason well before it decides whether a retrieval or execution tool is even needed.


How the Agent Actually Picks a Tool

The agent doesn't understand SD-WAN the way you do — it reads tool descriptions and matches them against the question. This is why the wording of a tool's description matters as much as the code behind it.

Say someone asks: "Is the voice traffic class healthy on the Chicago-to-DC tunnel right now?"

The agent scans its available tools and finds get_sla_performance described as "Retrieve current loss, latency, and jitter for a given tunnel and application class." That's a strong match — it extracts the tunnel and the voice class as inputs and calls it.

A vague description like "Gets SD-WAN stuff" would leave the agent guessing between three different tools that all sound plausible. A precise description — naming exactly what the tool returns and for what object type — is what makes tool selection reliable instead of a coin flip.


Anatomy of an SD-WAN Tool Schema

Take deploy_data_policy as a worked example of what a well-designed execution tool schema looks like:

  • Name: deploy_data_policy — the unique identifier the agent calls
  • Description: "Push a centralized data policy affecting application routing for a site list. Warning: makes real overlay-wide changes." — the explicit warning matters; it tells the agent (and anyone reviewing its plan) that this isn't a harmless lookup
  • Input schema:
    • site_list — which sites this policy scope applies to
    • app_class — the application or traffic match criteria
    • sla_class — which SLA class the traffic should be pinned to
    • mode — constrained to an enum of "preview" or "apply", defaulting to "preview"

That last field is the load-bearing detail. A default of preview means the agent's first call only shows what would happen — the policy doesn't actually get activated and pushed to vSmart until a human (or a separate, explicit step) chooses apply.

Schema design principles worth carrying into any SD-WAN tool you build:

  • Write descriptions specific enough that two tools never sound interchangeable
  • Type every input (don't let a site list or device ID be passed as a free-text string with no validation)
  • Use enums to constrain choices like mode, severity, or scope
  • Default to the safe option, never the destructive one
  • Mark required fields so the agent can't fire a call with the site list or app class left blank

Safety Considerations for SD-WAN-Facing Tools

A tool that can act on a centrally orchestrated overlay needs guardrails baked in from the start — not bolted on after the first incident.

Destructive actions. A tool like deploy_data_policy can affect application routing across every branch in the site list. Mitigation: default to preview mode, require explicit confirmation before applying, and never let an agent's very first call against a tool be a live push.

Credential exposure. If a tool logs its inputs and one of those inputs happens to include a vManage API token or a RADIUS credential passed through for auth, that's a real exposure. Mitigation: never log sensitive fields — scrub credentials before anything gets written to a log or a transcript.

Ineffective guardrails. A tool scoped to "manage the entire overlay" is too broad — it hands an agent far more blast radius than any single task requires. Mitigation: scope each tool tightly — a policy-deployment tool shouldn't also be able to touch device templates or controller certificates.

Cascading failures. SD-WAN workflows chain tool calls — the output of a tunnel-health check might feed into whether a firmware upgrade proceeds. If one tool returns bad or stale data, a downstream tool can act on it. Mitigation: build rollback into any tool that changes state, and don't let a single failed check silently get treated as a pass.


The Validation Pipeline Before Anything Executes

Before an execution tool is allowed to actually touch vManage or vSmart, three checks should pass, in order:

  1. Schema compliance — Are all required fields present and correctly typed? If site_list is missing or mode isn't one of the allowed enum values, reject the call before it goes anywhere near the overlay.
  2. Authorization — Does this user or agent identity actually have permission to call this tool? An agent scoped to read-only monitoring shouldn't be able to invoke deploy_data_policy at all, regardless of what it "decides" to do.
  3. Safety checks — Is this action allowed right now? A change-freeze window, an active P1 incident, or an in-progress controller upgrade are all reasons to block an otherwise-valid call.

Only after all three pass should the tool actually run against the controllers.


Quick Reference: Tool Category vs. Oversight Needed

CategorySD-WAN ExampleOversight Level
Retrievalget_tunnel_status, get_sla_performanceMinimal — safe to run freely
Executiondeploy_data_policy, update_device_templateHigh — preview mode, confirmation, rollback
Communicationsend_slack_alert, page_oncallLow — but should avoid alert fatigue
Perceptionsummarize_alarms, correlate_app_degradationLow — but accuracy matters, since downstream decisions rely on it

Final Thoughts

Tools are what make an AI agent useful on a real SD-WAN overlay instead of just a chatbot that can describe what an SLA class is. The four categories — retrieval, execution, communication, perception — map cleanly onto operations work you already do every day. The schema design determines whether tool selection is reliable or a guessing game. And the safety layer — preview-by-default execution, tight scoping, credential hygiene, and a validation pipeline before anything runs — is what determines whether you'd actually trust an agent near production vManage.

None of this replaces your judgment. It's what lets an agent earn a little bit of it, one well-scoped tool at a time.


FAQ

Q: Should an AI agent ever have direct, unsupervised write access to vManage? A: Generally no. Execution tools should default to preview mode and require explicit confirmation before applying, with authorization and safety checks run before every call — the same discipline you'd want from any junior engineer making overlay-wide changes.

Q: What's the biggest mistake in designing SD-WAN tool schemas for an agent? A: Vague tool descriptions. If two tools' descriptions sound interchangeable, the agent will eventually pick the wrong one — and on SD-WAN, the wrong tool call can mean a fleet-wide policy change instead of a simple status check.

Q: Are perception tools (like alarm summarization) risky the same way execution tools are? A: Not in the same way — they don't change the overlay — but their accuracy still matters, because a bad summary can lead a human or a downstream tool call to the wrong conclusion.


Related Reading on Networklearner:


Need help with SD-WAN, Cisco ACI, Nexus, data center networking, or network automation?

I am a CCIE Data Center engineer with 18+ years of enterprise networking experience, working hands-on with production SD-WAN and ACI environments.

Contact me for consulting, troubleshooting, design reviews, and project support: rockingoa@gmail.com

How AI Agents Actually "Touch" Your Cisco ACI Fabric: Tools, Schemas, and Guardrails Explained

 An AI agent that can only talk isn't much use to an ACI operations team. It can summarize a fault log you paste in, sure — but it can't tell you Leaf-104's health score right now, and it definitely shouldn't be pushing a contract change just because it "reasoned" its way there. What separates a chatbot from something you'd actually let near APIC is tools — the functions an agent is allowed to call to read from, act on, or talk about your fabric.

This post breaks down how tools work in an AI agent, using Cisco ACI as the running example throughout, so the concepts map directly onto things you already manage — health scores, contracts, EPGs, and change windows.

Why "Just Talking" Isn't Enough on ACI

Picture an agent with no tools at all:

Agent: "I think Leaf-104 might be unhealthy, but I have no way to check."

Not useful. Now give it a single tool that can query APIC's health score API:

Agent: [calls get_node_health(node="Leaf-104")] → "Leaf-104's health score is 62. Two contract-related faults are currently active on this leaf."

Same question, completely different value. That's the entire point of tools — they turn an agent from something that speculates into something that can actually verify against the fabric's real state.

The Four Tool Categories, Mapped to ACI

Every tool an ACI-aware agent might use falls into one of four buckets. Knowing which bucket a task belongs to tells you immediately how much oversight it needs.

1. Retrieval Tools — Read-Only Lookups

These pull information without changing anything: health scores, fault counts, contract relationships, endpoint locations.

ACI examples:

  • get_node_health(node) — pull a leaf or spine's current health score
  • get_epg_faults(epg) — list active faults on a given EPG
  • get_endpoint_location(ip_or_mac) — find which leaf/port an endpoint is learned on
  • get_contract_relationships(tenant, epg) — list which contracts an EPG consumes/provides

Retrieval tools are the safest category — an agent can call these freely without much risk, which is exactly why they're the easiest place to start trusting AI in ACI operations.

2. Execution Tools — Tools That Change the Fabric

These make actual changes: pushing a contract, modifying a Bridge Domain setting, restarting a service, triggering a firmware upgrade. Because ACI's policy model propagates changes fabric-wide, execution tools deserve the most design care of any category.

ACI examples:

  • deploy_contract(tenant, epg_consumer, epg_provider, filter, mode)
  • update_bridge_domain(tenant, bd, setting, value, mode)
  • trigger_firmware_upgrade(node, target_version, mode)

Notice the repeated mode parameter — more on that below. It's the single most important detail in an ACI execution tool's design.

3. Communication Tools — Looping in Humans

These don't touch the fabric at all — they notify people. Sending a Slack alert about a degraded leaf, opening a ServiceNow ticket for a recurring fault, paging the on-call engineer when a Multi-Site link drops.

ACI examples:

  • send_slack_alert(channel, message) — e.g., posting when a leaf's health score drops below threshold
  • create_servicenow_ticket(summary, severity, affected_tenant)
  • page_oncall(team, reason) — for something like an APIC cluster losing quorum

These tools are how an agent stays useful even when it shouldn't act on its own — escalating to a human is often the correct behavior, not a fallback.

4. Perception Tools — Making Sense of Raw Data

These interpret information rather than fetch or change it: parsing a wall of fault codes into a plain-English summary, correlating a traffic spike with a recent contract change, summarizing a week of APIC audit logs.

ACI examples:

  • summarize_faults(tenant, time_range) — turn 200 raw fault codes into three actionable findings
  • correlate_traffic_anomaly(epg, time_range) — check if a spike lines up with a recent policy push
  • parse_audit_log(node, time_range) — extract what actually changed and who changed it

Perception tools are what let an agent reason well before it decides whether a retrieval or execution tool is even needed.

How the Agent Actually Picks a Tool

The agent doesn't understand ACI the way you do — it reads tool descriptions and matches them against the question. This is why the wording of a tool's description matters as much as the code behind it.

Say someone asks: "Is EPG WEB-EPG healthy right now?"

The agent scans its available tools and finds get_epg_faults described as "Retrieve current fault count and severity for a given EPG." That's a strong match — it extracts WEB-EPG as the input and calls it.

A vague description like "Gets ACI stuff" would leave the agent guessing between three different tools that all sound plausible. A precise description — naming exactly what the tool returns and for what object type — is what makes tool selection reliable instead of a coin flip.

Anatomy of an ACI Tool Schema

Take deploy_contract as a worked example of what a well-designed execution tool schema looks like:

  • Name: deploy_contract — the unique identifier the agent calls
  • Description: "Push a contract between two EPGs in a tenant. Warning: makes real fabric changes." — the explicit warning matters; it tells the agent (and anyone reviewing its plan) that this isn't a harmless lookup
  • Input schema:
    • tenant — which tenant this applies to
    • epg_consumer / epg_provider — the two EPGs the contract connects
    • filter — the port/protocol filter being applied
    • mode — constrained to an enum of "preview" or "apply", defaulting to "preview"

That last field is the load-bearing detail. A default of preview means the agent's first call only shows what would happen — the actual zoning-rule programming on the leaves doesn't happen until a human (or a separate, explicit step) chooses apply.

Schema design principles worth carrying into any ACI tool you build:

  • Write descriptions specific enough that two tools never sound interchangeable
  • Type every input (don't let a device ID be passed as a free-text string with no validation)
  • Use enums to constrain choices like mode, severity, or scope
  • Default to the safe option, never the destructive one
  • Mark required fields so the agent can't fire a call with the tenant or EPG left blank

Safety Considerations for ACI-Facing Tools

A tool that can act on a shared fabric needs guardrails baked in from the start — not bolted on after the first incident.

Destructive actions. A tool like update_bridge_domain can affect flooding or ARP behavior fabric-wide. Mitigation: default to preview mode, require explicit confirmation before applying, and never let an agent's very first call against a tool be a live push.

Credential exposure. If a tool logs its inputs and one of those inputs happens to include an APIC admin token or a TACACS credential passed through for auth, that's a real exposure. Mitigation: never log sensitive fields — scrub credentials before anything gets written to a log or a transcript.

Ineffective guardrails. A tool scoped to "manage the entire fabric" is too broad — it hands an agent far more blast radius than any single task requires. Mitigation: scope each tool tightly — a contract-deployment tool shouldn't also be able to touch fabric access policies.

Cascading failures. ACI workflows chain tool calls — the output of a health check might feed into whether a firmware upgrade proceeds. If one tool returns bad or stale data, a downstream tool can act on it. Mitigation: build rollback into any tool that changes state, and don't let a single failed check silently get treated as a pass.

The Validation Pipeline Before Anything Executes

Before an execution tool is allowed to actually touch APIC, three checks should pass, in order:

  1. Schema compliance — Are all required fields present and correctly typed? If tenant is missing or mode isn't one of the allowed enum values, reject the call before it goes anywhere near the fabric.
  2. Authorization — Does this user or agent identity actually have permission to call this tool? An agent scoped to read-only monitoring shouldn't be able to invoke deploy_contract at all, regardless of what it "decides" to do.
  3. Safety checks — Is this action allowed right now? A change-freeze window, an active P1 incident, or an in-progress firmware upgrade are all reasons to block an otherwise-valid call.

Only after all three pass should the tool actually run against APIC.

Quick Reference: Tool Category vs. Oversight Needed

CategoryACI ExampleOversight Level
Retrievalget_node_health, get_epg_faultsMinimal — safe to run freely
Executiondeploy_contract, update_bridge_domainHigh — preview mode, confirmation, rollback
Communicationsend_slack_alert, page_oncallLow — but should avoid alert fatigue
Perceptionsummarize_faults, correlate_traffic_anomalyLow — but accuracy matters, since downstream decisions rely on it

Final Thoughts

Tools are what make an AI agent useful on a real ACI fabric instead of just a chatbot that can describe what a health score is. The four categories — retrieval, execution, communication, perception — map cleanly onto operations work you already do every day. The schema design determines whether tool selection is reliable or a guessing game. And the safety layer — preview-by-default execution, tight scoping, credential hygiene, and a validation pipeline before anything runs — is what determines whether you'd actually trust an agent near production APIC.

None of this replaces your judgment. It's what lets an agent earn a little bit of it, one well-scoped tool at a time.

FAQ

Q: Should an AI agent ever have direct, unsupervised write access to APIC? A: Generally no. Execution tools should default to preview mode and require explicit confirmation before applying, with authorization and safety checks run before every call — the same discipline you'd want from any junior engineer making fabric changes.

Q: What's the biggest mistake in designing ACI tool schemas for an agent? A: Vague tool descriptions. If two tools' descriptions sound interchangeable, the agent will eventually pick the wrong one — and on ACI, the wrong tool call can mean a fabric-wide policy change instead of a simple status check.

Q: Are perception tools (like fault summarization) risky the same way execution tools are? A: Not in the same way — they don't change the fabric — but their accuracy still matters, because a bad summary can lead a human or a downstream tool call to the wrong conclusion.

Related Reading on Networklearner:

Need help with Cisco ACI, Nexus, data center networking, or network automation?

I am a CCIE Data Center engineer with 18+ years of enterprise networking experience, working hands-on with production ACI fabrics.

Contact me for consulting, troubleshooting, design reviews, and project support: rockingoa@gmail.com