Wednesday, 5 August 2026

How AI Memory Can Transform Cisco SD-WAN Operations: Working Memory, Long-Term Memory & RAG Explained with Real Examples

 Enterprise WAN networks are becoming increasingly complex. A modern Cisco SD-WAN deployment may include hundreds of branch locations, multiple transport circuits, cloud connectivity, SaaS applications, and centralized policy management.

While Cisco SD-WAN simplifies operations through centralized control, troubleshooting issues across multiple sites can still consume significant time.

Imagine receiving alerts that twenty branch offices have simultaneously lost connectivity to Microsoft Azure. Instead of manually checking control connections, OMP routes, TLOC status, application-aware routing policies, and tunnel health, an AI-powered assistant could analyze the entire environment in seconds.

But how does an AI assistant remember previous troubleshooting steps, understand your SD-WAN architecture, and retrieve your organization's deployment standards?

The answer lies in AI Memory.

Just as an experienced SD-WAN engineer relies on operational knowledge, design documentation, and previous incidents, AI agents use different types of memory to provide intelligent, context-aware assistance.

In this article, we'll explore how AI memory works using practical Cisco SD-WAN examples.


Why Memory Matters in Cisco SD-WAN

Consider a large enterprise with:

  • Two vManage controllers
  • Two vBond orchestrators
  • Three vSmart controllers
  • 400 WAN Edge routers
  • MPLS, Internet, and 5G transports
  • Hundreds of VPNs
  • Thousands of OMP routes

Now imagine that users across multiple branches report poor Microsoft Teams performance.

Without memory, an AI assistant would repeatedly ask:

  • Which sites are affected?
  • Which transport is failing?
  • What policies are configured?
  • Are control connections established?
  • Has any template changed recently?

An experienced engineer already understands much of this context. AI memory enables an intelligent assistant to retain relevant information, retrieve documentation, and build on previous investigations instead of starting from scratch.


Understanding AI Memory

AI memory can be compared to how an experienced Cisco SD-WAN administrator manages information.

AI MemoryCisco SD-WAN Example
Working MemoryCurrent WAN outage investigation
Long-Term MemorySD-WAN design guides and runbooks
Episodic MemoryPrevious outages and administrator preferences

Together, these memory types help AI provide faster and more accurate recommendations.


Working Memory – Understanding the Current Incident

Working memory contains the information the AI agent is actively using during the current troubleshooting session. It is temporary and focused on the ongoing investigation.

Cisco SD-WAN Example

An engineer reports:

"Branch-105 cannot access Azure."

The AI assistant remembers:

  • Branch name
  • WAN Edge router
  • VPN ID
  • Transport circuits
  • TLOC status
  • OMP advertisements
  • BFD session state
  • Application-aware routing policy
  • Results from previous commands

Instead of requesting the same details repeatedly, it builds on the existing context to accelerate troubleshooting.


Long-Term Memory – Your Enterprise SD-WAN Knowledge Base

Working memory disappears after the session ends.

Long-term memory stores persistent information that the AI can retrieve whenever needed. This may include architecture documents, operational procedures, and deployment standards.

For Cisco SD-WAN, this could include:

  • WAN architecture diagrams
  • Controller deployment standards
  • VPN segmentation guidelines
  • Application-aware routing policies
  • Security policies
  • Template standards
  • Branch deployment procedures
  • Change management documentation

Cisco SD-WAN Example

An engineer asks:

"What is our standard QoS policy for Microsoft Teams?"

Instead of generating a generic answer, the AI retrieves the organization's approved policy documentation and provides guidance that aligns with internal standards.

This ensures consistency across every branch deployment.


What is RAG?

RAG (Retrieval-Augmented Generation) allows AI to retrieve trusted documentation before generating an answer. Rather than relying only on its training data, the AI searches your organization's SD-WAN knowledge base, such as deployment guides, runbooks, and design documents, to produce responses grounded in your own environment.

Cisco SD-WAN Example

An engineer asks:

"How do we normally configure Direct Internet Access for branch offices?"

The AI retrieves the organization's approved deployment guide and recommends the documented configuration instead of offering generic internet advice.


Episodic Memory – Learning from Previous Outages

Episodic memory stores previous interactions and user preferences, allowing the AI to personalize future assistance.

Cisco SD-WAN Example

Suppose that three months ago you resolved an issue where unstable BFD sessions over the broadband circuit caused intermittent application failures.

When a similar pattern appears again, the AI highlights the previous incident and suggests checking BFD stability before exploring more complex causes.

It may also remember that you prefer CLI commands alongside vManage workflows, tailoring its recommendations to your working style.


How AI Memory Helps During a WAN Outage

Imagine a global outage affecting dozens of branch offices.

An AI assistant can:

  1. Use Working Memory to track the current investigation, including affected sites and diagnostic results.
  2. Use Long-Term Memory to retrieve SD-WAN design standards, routing policies, and operational runbooks.
  3. Use Episodic Memory to compare the current symptoms with previous incidents and recommend proven remediation steps.

This combination reduces Mean Time to Resolution (MTTR) and helps engineers resolve issues more efficiently.


Benefits for Cisco SD-WAN Engineers

AI memory can significantly improve daily operations by:

  • Reducing repetitive troubleshooting tasks
  • Accelerating root cause analysis
  • Retrieving internal documentation instantly
  • Preserving operational knowledge
  • Supporting junior engineers with guided diagnostics
  • Standardizing troubleshooting procedures
  • Improving consistency across large-scale SD-WAN deployments
  • Enabling AI-assisted network operations

Real-World Use Cases

AI memory can support many common Cisco SD-WAN tasks, including:

  • Diagnosing OMP route advertisement issues
  • Troubleshooting BFD session flaps
  • Investigating TLOC extension failures
  • Validating centralized and localized policies
  • Checking application-aware routing decisions
  • Reviewing controller certificate and control connection issues
  • Verifying software upgrade procedures
  • Assisting with Zero-Touch Provisioning (ZTP)
  • Troubleshooting SaaS connectivity
  • Supporting cloud on-ramp deployments

Frequently Asked Questions

Can AI remember my Cisco SD-WAN deployment?

Yes, if it is connected to persistent knowledge sources and designed to retain relevant operational context.

How does RAG improve SD-WAN troubleshooting?

It enables AI to retrieve approved runbooks, design guides, and operational documentation before generating recommendations.

Can AI replace Cisco SD-WAN engineers?

No. AI enhances productivity by accelerating troubleshooting and surfacing relevant information, while engineers remain responsible for design, validation, and operational decisions.


How AI Memory Can Revolutionize Cisco ACI Operations

Cisco ACI has transformed data center networking by introducing policy-based automation, centralized management, and application-centric networking. However, as enterprise fabrics grow, troubleshooting becomes increasingly challenging.

Imagine receiving multiple faults from different leaf switches after a maintenance window. Instead of manually checking APIC health, interface counters, contracts, endpoint learning, and event logs one by one, an AI assistant could analyze everything in seconds.

But how can an AI assistant remember previous troubleshooting steps, understand your fabric, and retrieve your organization's deployment standards?

The answer lies in AI Memory.

Just like an experienced Cisco ACI engineer relies on operational knowledge, design documents, and previous troubleshooting experience, AI agents use different types of memory to provide intelligent assistance.

In this article, we'll explore how AI memory works using real Cisco ACI scenarios.


Why Memory Matters in Cisco ACI

Suppose your production ACI fabric consists of:

  • Three APIC controllers
  • Two Spine switches
  • Twenty Leaf switches
  • Hundreds of EPGs
  • Multiple VRFs
  • External L3Out connections
  • VMware VMM integration

Now imagine that an application suddenly loses connectivity.

Without memory, an AI assistant would ask the same questions every time:

  • Which tenant is affected?
  • Which EPGs are involved?
  • What contracts exist?
  • Which leaf switch hosts the endpoint?
  • Was any policy recently modified?

An experienced engineer already remembers much of this context. Likewise, AI memory allows an intelligent assistant to retain relevant information, retrieve documentation, and personalize future troubleshooting.


The Three Types of AI Memory

Think of AI memory as the way a senior Cisco ACI architect organizes information.

AI MemoryCisco ACI Example
Working MemoryCurrent APIC fault investigation
Long-Term MemoryACI design documents and runbooks
Episodic MemoryPrevious incidents and administrator preferences

Together, these memory types help AI solve problems faster and more accurately.


Working Memory – The Current Troubleshooting Session

Working memory contains information that the AI agent is actively processing during the current investigation. It is temporary and limited to the ongoing session.

Cisco ACI Example

An engineer asks:

"Why are endpoints in EPG-Web unable to communicate with EPG-App?"

The AI assistant keeps track of:

  • Tenant name
  • VRF
  • Bridge Domain
  • EPG names
  • Applied contracts
  • Recent APIC fault messages
  • Results of previous checks

Instead of asking for the same information repeatedly, it builds on the conversation until the issue is resolved.


Long-Term Memory – Your Cisco ACI Knowledge Base

Long-term memory stores persistent information that does not disappear after the conversation ends. In AI systems, this often includes documentation and runbooks that can be retrieved when needed.

For Cisco ACI, this could include:

  • Fabric architecture diagrams
  • Tenant standards
  • Naming conventions
  • Interface policies
  • L3Out design guides
  • APIC backup procedures
  • Security policies
  • Change management documents

Cisco ACI Example

An engineer asks:

"What is our standard configuration for external routed networks?"

Instead of relying on generic knowledge, the AI searches your organization's approved ACI design guide and provides recommendations based on your own standards.

This improves consistency and reduces the risk of configuration drift.


What is RAG?

RAG (Retrieval-Augmented Generation) allows an AI assistant to retrieve relevant documentation before generating a response. Rather than guessing, it searches trusted sources such as your ACI runbooks, design guides, and operational procedures.

Cisco ACI Example

You ask:

"How do we normally configure BGP authentication on our L3Outs?"

The AI retrieves the organization's approved implementation guide and answers using that document rather than generic internet advice.


Episodic Memory – Learning from Previous Incidents

Episodic memory records previous interactions and user preferences so future assistance can be more personalized.

Cisco ACI Example

Suppose that last month you resolved a fault caused by a missing contract between two EPGs.

Months later, a similar fault occurs.

The AI recognizes the similarity and suggests checking contracts early in the troubleshooting process, saving valuable time.

It may also remember that you prefer CLI outputs alongside APIC GUI navigation, allowing responses to match your working style.


Bringing It All Together

Imagine a production outage affecting application connectivity.

An AI assistant could:

  1. Use Working Memory to keep track of the current troubleshooting session.
  2. Use Long-Term Memory to retrieve your organization's ACI standards and runbooks.
  3. Use Episodic Memory to recognize similar past incidents and apply successful troubleshooting patterns.

This combination provides faster diagnostics, more consistent recommendations, and reduced troubleshooting time.


Benefits for Cisco ACI Engineers

By combining AI memory with Cisco ACI, organizations can:

  • Accelerate root cause analysis.
  • Reduce repetitive troubleshooting.
  • Retrieve design documentation instantly.
  • Improve adherence to operational standards.
  • Preserve knowledge from experienced engineers.
  • Shorten onboarding time for new team members.
  • Enable more intelligent AI-driven network operations.

Key Takeaways

  • Working Memory manages the current troubleshooting context.
  • Long-Term Memory stores ACI documentation, standards, and runbooks.
  • Episodic Memory captures previous incidents and user preferences.
  • RAG connects AI to trusted enterprise documentation instead of relying only on model training.
  • Together, these capabilities can significantly improve Cisco ACI operations and troubleshooting efficiency.

Frequently Asked Questions

Can AI remember my Cisco ACI fabric permanently?
Only if the AI platform is designed to store and retrieve persistent knowledge such as documentation and previous interactions.

How does RAG help Cisco ACI administrators?
It allows AI to search approved ACI documentation and generate answers based on your organization's standards instead of generic information.

Can AI replace Cisco ACI engineers?
No. AI augments engineers by reducing repetitive tasks and surfacing relevant information, while design decisions and operational oversight remain with experienced professionals.


Related Articles to add to "How AI Memory Can Revolutionize Cisco ACI Operations":

Need help with Cisco ACI, Nexus, data center networking, or network automation?

I am a CCIE Data Center engineer with 18+ years of enterprise networking experience.

Contact me for consulting, troubleshooting, design reviews, and project support.


Tuesday, 4 August 2026

How AI Agents Actually "Touch" Your SD-WAN Overlay: Tools, Schemas, and Guardrails Explained

 An AI agent that can only talk isn't much use to an SD-WAN operations team. It can summarize a wall of syslog you paste in, sure — but it can't tell you a branch's current BFD state right now, and it definitely shouldn't be pushing a centralized policy change just because it "reasoned" its way there. What separates a chatbot from something you'd actually let near vManage is tools — the functions an agent is allowed to call to read from, act on, or talk about your overlay.

This post breaks down how tools work in an AI agent, using SD-WAN as the running example throughout, so the concepts map directly onto things you already manage — tunnels, SLA classes, device templates, and centralized policy.


Why "Just Talking" Isn't Enough on SD-WAN

Picture an agent with no tools at all:

Agent: "I think Branch-42's MPLS tunnel might be down, but I have no way to check."

Not useful. Now give it a single tool that can query vManage's device API:

Agent: [calls get_tunnel_status(branch="Branch-42", transport="MPLS")] → "Branch-42's MPLS tunnel is down. BFD lost sync 4 minutes ago; the branch has failed over to Internet transport."

Same question, completely different value. That's the entire point of tools — they turn an agent from something that speculates into something that can actually verify against the overlay's real state.


The Four Tool Categories, Mapped to SD-WAN

Every tool an SD-WAN-aware agent might use falls into one of four buckets. Knowing which bucket a task belongs to tells you immediately how much oversight it needs.

1. Retrieval Tools — Read-Only Lookups

These pull information without changing anything: tunnel status, SLA performance, control-connection state, transport in use.

SD-WAN examples:

  • get_tunnel_status(branch, transport) — pull current BFD/session state for a specific tunnel
  • get_sla_performance(tunnel, app_class) — loss/latency/jitter for a given app class
  • get_control_connections(device) — check a device's control-plane state to vSmart
  • get_active_transport(branch, app) — which underlay path an app is currently steered over

Retrieval tools are the safest category — an agent can call these freely without much risk, which is exactly why they're the easiest place to start trusting AI in SD-WAN operations.

2. Execution Tools — Tools That Change the Overlay

These make actual changes: pushing a centralized data policy, updating a device template, restarting a tunnel, triggering a software upgrade. Because SD-WAN policy is orchestrated centrally, execution tools deserve the most design care of any category — a bad push doesn't stay local, it propagates to every site attached to that policy or template.

SD-WAN examples:

  • deploy_data_policy(site_list, app_class, sla_class, mode)
  • update_device_template(device, template, mode)
  • trigger_software_upgrade(device, target_version, mode)

Notice the repeated mode parameter — more on that below. It's the single most important detail in an SD-WAN execution tool's design.

3. Communication Tools — Looping in Humans

These don't touch the overlay at all — they notify people. Sending a Slack alert about a degraded transport, opening a ServiceNow ticket for a recurring BFD flap, paging the on-call engineer when a regional hub loses reachability.

SD-WAN examples:

  • send_slack_alert(channel, message) — e.g., posting when an SLA class breaches threshold on a tunnel
  • create_servicenow_ticket(summary, severity, affected_site)
  • page_oncall(team, reason) — for something like a vSmart cluster losing quorum

These tools are how an agent stays useful even when it shouldn't act on its own — escalating to a human is often the correct behavior, not a fallback.

4. Perception Tools — Making Sense of Raw Data

These interpret information rather than fetch or change it: parsing a flood of BFD flap events into a plain-English summary, correlating a voice-quality complaint with a recent policy push, summarizing a week of vManage audit logs.

SD-WAN examples:

  • summarize_alarms(site, time_range) — turn 150 raw alarms into three actionable findings
  • correlate_app_degradation(app, time_range) — check if a slowdown lines up with a recent policy or template change
  • parse_audit_log(controller, time_range) — extract what actually changed and who changed it

Perception tools are what let an agent reason well before it decides whether a retrieval or execution tool is even needed.


How the Agent Actually Picks a Tool

The agent doesn't understand SD-WAN the way you do — it reads tool descriptions and matches them against the question. This is why the wording of a tool's description matters as much as the code behind it.

Say someone asks: "Is the voice traffic class healthy on the Chicago-to-DC tunnel right now?"

The agent scans its available tools and finds get_sla_performance described as "Retrieve current loss, latency, and jitter for a given tunnel and application class." That's a strong match — it extracts the tunnel and the voice class as inputs and calls it.

A vague description like "Gets SD-WAN stuff" would leave the agent guessing between three different tools that all sound plausible. A precise description — naming exactly what the tool returns and for what object type — is what makes tool selection reliable instead of a coin flip.


Anatomy of an SD-WAN Tool Schema

Take deploy_data_policy as a worked example of what a well-designed execution tool schema looks like:

  • Name: deploy_data_policy — the unique identifier the agent calls
  • Description: "Push a centralized data policy affecting application routing for a site list. Warning: makes real overlay-wide changes." — the explicit warning matters; it tells the agent (and anyone reviewing its plan) that this isn't a harmless lookup
  • Input schema:
    • site_list — which sites this policy scope applies to
    • app_class — the application or traffic match criteria
    • sla_class — which SLA class the traffic should be pinned to
    • mode — constrained to an enum of "preview" or "apply", defaulting to "preview"

That last field is the load-bearing detail. A default of preview means the agent's first call only shows what would happen — the policy doesn't actually get activated and pushed to vSmart until a human (or a separate, explicit step) chooses apply.

Schema design principles worth carrying into any SD-WAN tool you build:

  • Write descriptions specific enough that two tools never sound interchangeable
  • Type every input (don't let a site list or device ID be passed as a free-text string with no validation)
  • Use enums to constrain choices like mode, severity, or scope
  • Default to the safe option, never the destructive one
  • Mark required fields so the agent can't fire a call with the site list or app class left blank

Safety Considerations for SD-WAN-Facing Tools

A tool that can act on a centrally orchestrated overlay needs guardrails baked in from the start — not bolted on after the first incident.

Destructive actions. A tool like deploy_data_policy can affect application routing across every branch in the site list. Mitigation: default to preview mode, require explicit confirmation before applying, and never let an agent's very first call against a tool be a live push.

Credential exposure. If a tool logs its inputs and one of those inputs happens to include a vManage API token or a RADIUS credential passed through for auth, that's a real exposure. Mitigation: never log sensitive fields — scrub credentials before anything gets written to a log or a transcript.

Ineffective guardrails. A tool scoped to "manage the entire overlay" is too broad — it hands an agent far more blast radius than any single task requires. Mitigation: scope each tool tightly — a policy-deployment tool shouldn't also be able to touch device templates or controller certificates.

Cascading failures. SD-WAN workflows chain tool calls — the output of a tunnel-health check might feed into whether a firmware upgrade proceeds. If one tool returns bad or stale data, a downstream tool can act on it. Mitigation: build rollback into any tool that changes state, and don't let a single failed check silently get treated as a pass.


The Validation Pipeline Before Anything Executes

Before an execution tool is allowed to actually touch vManage or vSmart, three checks should pass, in order:

  1. Schema compliance — Are all required fields present and correctly typed? If site_list is missing or mode isn't one of the allowed enum values, reject the call before it goes anywhere near the overlay.
  2. Authorization — Does this user or agent identity actually have permission to call this tool? An agent scoped to read-only monitoring shouldn't be able to invoke deploy_data_policy at all, regardless of what it "decides" to do.
  3. Safety checks — Is this action allowed right now? A change-freeze window, an active P1 incident, or an in-progress controller upgrade are all reasons to block an otherwise-valid call.

Only after all three pass should the tool actually run against the controllers.


Quick Reference: Tool Category vs. Oversight Needed

CategorySD-WAN ExampleOversight Level
Retrievalget_tunnel_status, get_sla_performanceMinimal — safe to run freely
Executiondeploy_data_policy, update_device_templateHigh — preview mode, confirmation, rollback
Communicationsend_slack_alert, page_oncallLow — but should avoid alert fatigue
Perceptionsummarize_alarms, correlate_app_degradationLow — but accuracy matters, since downstream decisions rely on it

Final Thoughts

Tools are what make an AI agent useful on a real SD-WAN overlay instead of just a chatbot that can describe what an SLA class is. The four categories — retrieval, execution, communication, perception — map cleanly onto operations work you already do every day. The schema design determines whether tool selection is reliable or a guessing game. And the safety layer — preview-by-default execution, tight scoping, credential hygiene, and a validation pipeline before anything runs — is what determines whether you'd actually trust an agent near production vManage.

None of this replaces your judgment. It's what lets an agent earn a little bit of it, one well-scoped tool at a time.


FAQ

Q: Should an AI agent ever have direct, unsupervised write access to vManage? A: Generally no. Execution tools should default to preview mode and require explicit confirmation before applying, with authorization and safety checks run before every call — the same discipline you'd want from any junior engineer making overlay-wide changes.

Q: What's the biggest mistake in designing SD-WAN tool schemas for an agent? A: Vague tool descriptions. If two tools' descriptions sound interchangeable, the agent will eventually pick the wrong one — and on SD-WAN, the wrong tool call can mean a fleet-wide policy change instead of a simple status check.

Q: Are perception tools (like alarm summarization) risky the same way execution tools are? A: Not in the same way — they don't change the overlay — but their accuracy still matters, because a bad summary can lead a human or a downstream tool call to the wrong conclusion.


Related Reading on Networklearner:


Need help with SD-WAN, Cisco ACI, Nexus, data center networking, or network automation?

I am a CCIE Data Center engineer with 18+ years of enterprise networking experience, working hands-on with production SD-WAN and ACI environments.

Contact me for consulting, troubleshooting, design reviews, and project support: rockingoa@gmail.com

How AI Agents Actually "Touch" Your Cisco ACI Fabric: Tools, Schemas, and Guardrails Explained

 An AI agent that can only talk isn't much use to an ACI operations team. It can summarize a fault log you paste in, sure — but it can't tell you Leaf-104's health score right now, and it definitely shouldn't be pushing a contract change just because it "reasoned" its way there. What separates a chatbot from something you'd actually let near APIC is tools — the functions an agent is allowed to call to read from, act on, or talk about your fabric.

This post breaks down how tools work in an AI agent, using Cisco ACI as the running example throughout, so the concepts map directly onto things you already manage — health scores, contracts, EPGs, and change windows.

Why "Just Talking" Isn't Enough on ACI

Picture an agent with no tools at all:

Agent: "I think Leaf-104 might be unhealthy, but I have no way to check."

Not useful. Now give it a single tool that can query APIC's health score API:

Agent: [calls get_node_health(node="Leaf-104")] → "Leaf-104's health score is 62. Two contract-related faults are currently active on this leaf."

Same question, completely different value. That's the entire point of tools — they turn an agent from something that speculates into something that can actually verify against the fabric's real state.

The Four Tool Categories, Mapped to ACI

Every tool an ACI-aware agent might use falls into one of four buckets. Knowing which bucket a task belongs to tells you immediately how much oversight it needs.

1. Retrieval Tools — Read-Only Lookups

These pull information without changing anything: health scores, fault counts, contract relationships, endpoint locations.

ACI examples:

  • get_node_health(node) — pull a leaf or spine's current health score
  • get_epg_faults(epg) — list active faults on a given EPG
  • get_endpoint_location(ip_or_mac) — find which leaf/port an endpoint is learned on
  • get_contract_relationships(tenant, epg) — list which contracts an EPG consumes/provides

Retrieval tools are the safest category — an agent can call these freely without much risk, which is exactly why they're the easiest place to start trusting AI in ACI operations.

2. Execution Tools — Tools That Change the Fabric

These make actual changes: pushing a contract, modifying a Bridge Domain setting, restarting a service, triggering a firmware upgrade. Because ACI's policy model propagates changes fabric-wide, execution tools deserve the most design care of any category.

ACI examples:

  • deploy_contract(tenant, epg_consumer, epg_provider, filter, mode)
  • update_bridge_domain(tenant, bd, setting, value, mode)
  • trigger_firmware_upgrade(node, target_version, mode)

Notice the repeated mode parameter — more on that below. It's the single most important detail in an ACI execution tool's design.

3. Communication Tools — Looping in Humans

These don't touch the fabric at all — they notify people. Sending a Slack alert about a degraded leaf, opening a ServiceNow ticket for a recurring fault, paging the on-call engineer when a Multi-Site link drops.

ACI examples:

  • send_slack_alert(channel, message) — e.g., posting when a leaf's health score drops below threshold
  • create_servicenow_ticket(summary, severity, affected_tenant)
  • page_oncall(team, reason) — for something like an APIC cluster losing quorum

These tools are how an agent stays useful even when it shouldn't act on its own — escalating to a human is often the correct behavior, not a fallback.

4. Perception Tools — Making Sense of Raw Data

These interpret information rather than fetch or change it: parsing a wall of fault codes into a plain-English summary, correlating a traffic spike with a recent contract change, summarizing a week of APIC audit logs.

ACI examples:

  • summarize_faults(tenant, time_range) — turn 200 raw fault codes into three actionable findings
  • correlate_traffic_anomaly(epg, time_range) — check if a spike lines up with a recent policy push
  • parse_audit_log(node, time_range) — extract what actually changed and who changed it

Perception tools are what let an agent reason well before it decides whether a retrieval or execution tool is even needed.

How the Agent Actually Picks a Tool

The agent doesn't understand ACI the way you do — it reads tool descriptions and matches them against the question. This is why the wording of a tool's description matters as much as the code behind it.

Say someone asks: "Is EPG WEB-EPG healthy right now?"

The agent scans its available tools and finds get_epg_faults described as "Retrieve current fault count and severity for a given EPG." That's a strong match — it extracts WEB-EPG as the input and calls it.

A vague description like "Gets ACI stuff" would leave the agent guessing between three different tools that all sound plausible. A precise description — naming exactly what the tool returns and for what object type — is what makes tool selection reliable instead of a coin flip.

Anatomy of an ACI Tool Schema

Take deploy_contract as a worked example of what a well-designed execution tool schema looks like:

  • Name: deploy_contract — the unique identifier the agent calls
  • Description: "Push a contract between two EPGs in a tenant. Warning: makes real fabric changes." — the explicit warning matters; it tells the agent (and anyone reviewing its plan) that this isn't a harmless lookup
  • Input schema:
    • tenant — which tenant this applies to
    • epg_consumer / epg_provider — the two EPGs the contract connects
    • filter — the port/protocol filter being applied
    • mode — constrained to an enum of "preview" or "apply", defaulting to "preview"

That last field is the load-bearing detail. A default of preview means the agent's first call only shows what would happen — the actual zoning-rule programming on the leaves doesn't happen until a human (or a separate, explicit step) chooses apply.

Schema design principles worth carrying into any ACI tool you build:

  • Write descriptions specific enough that two tools never sound interchangeable
  • Type every input (don't let a device ID be passed as a free-text string with no validation)
  • Use enums to constrain choices like mode, severity, or scope
  • Default to the safe option, never the destructive one
  • Mark required fields so the agent can't fire a call with the tenant or EPG left blank

Safety Considerations for ACI-Facing Tools

A tool that can act on a shared fabric needs guardrails baked in from the start — not bolted on after the first incident.

Destructive actions. A tool like update_bridge_domain can affect flooding or ARP behavior fabric-wide. Mitigation: default to preview mode, require explicit confirmation before applying, and never let an agent's very first call against a tool be a live push.

Credential exposure. If a tool logs its inputs and one of those inputs happens to include an APIC admin token or a TACACS credential passed through for auth, that's a real exposure. Mitigation: never log sensitive fields — scrub credentials before anything gets written to a log or a transcript.

Ineffective guardrails. A tool scoped to "manage the entire fabric" is too broad — it hands an agent far more blast radius than any single task requires. Mitigation: scope each tool tightly — a contract-deployment tool shouldn't also be able to touch fabric access policies.

Cascading failures. ACI workflows chain tool calls — the output of a health check might feed into whether a firmware upgrade proceeds. If one tool returns bad or stale data, a downstream tool can act on it. Mitigation: build rollback into any tool that changes state, and don't let a single failed check silently get treated as a pass.

The Validation Pipeline Before Anything Executes

Before an execution tool is allowed to actually touch APIC, three checks should pass, in order:

  1. Schema compliance — Are all required fields present and correctly typed? If tenant is missing or mode isn't one of the allowed enum values, reject the call before it goes anywhere near the fabric.
  2. Authorization — Does this user or agent identity actually have permission to call this tool? An agent scoped to read-only monitoring shouldn't be able to invoke deploy_contract at all, regardless of what it "decides" to do.
  3. Safety checks — Is this action allowed right now? A change-freeze window, an active P1 incident, or an in-progress firmware upgrade are all reasons to block an otherwise-valid call.

Only after all three pass should the tool actually run against APIC.

Quick Reference: Tool Category vs. Oversight Needed

CategoryACI ExampleOversight Level
Retrievalget_node_health, get_epg_faultsMinimal — safe to run freely
Executiondeploy_contract, update_bridge_domainHigh — preview mode, confirmation, rollback
Communicationsend_slack_alert, page_oncallLow — but should avoid alert fatigue
Perceptionsummarize_faults, correlate_traffic_anomalyLow — but accuracy matters, since downstream decisions rely on it

Final Thoughts

Tools are what make an AI agent useful on a real ACI fabric instead of just a chatbot that can describe what a health score is. The four categories — retrieval, execution, communication, perception — map cleanly onto operations work you already do every day. The schema design determines whether tool selection is reliable or a guessing game. And the safety layer — preview-by-default execution, tight scoping, credential hygiene, and a validation pipeline before anything runs — is what determines whether you'd actually trust an agent near production APIC.

None of this replaces your judgment. It's what lets an agent earn a little bit of it, one well-scoped tool at a time.

FAQ

Q: Should an AI agent ever have direct, unsupervised write access to APIC? A: Generally no. Execution tools should default to preview mode and require explicit confirmation before applying, with authorization and safety checks run before every call — the same discipline you'd want from any junior engineer making fabric changes.

Q: What's the biggest mistake in designing ACI tool schemas for an agent? A: Vague tool descriptions. If two tools' descriptions sound interchangeable, the agent will eventually pick the wrong one — and on ACI, the wrong tool call can mean a fabric-wide policy change instead of a simple status check.

Q: Are perception tools (like fault summarization) risky the same way execution tools are? A: Not in the same way — they don't change the fabric — but their accuracy still matters, because a bad summary can lead a human or a downstream tool call to the wrong conclusion.

Related Reading on Networklearner:

Need help with Cisco ACI, Nexus, data center networking, or network automation?

I am a CCIE Data Center engineer with 18+ years of enterprise networking experience, working hands-on with production ACI fabrics.

Contact me for consulting, troubleshooting, design reviews, and project support: rockingoa@gmail.com