AI Agents for Business: Opportunities, Limits, and Use Cases
An AI agent is not a smarter chatbot: it decides steps and acts on real systems. Definitions, use cases by function, five concrete risks, a documented court case, and a governance model built on levels of autonomy.
Table of contents
- Chatbot, workflow, agent: three different things sold under one label
- Four questions to decide whether you need an agent
- Use cases by business function
- Five concrete risks, not generic ones
- Who is liable when the agent gets it wrong: the Air Canada case
- Minimum viable governance: levels of autonomy and controls
- From pilot to production: four gating criteria
- Common mistakes
- Next step
An AI agent is not a smarter chatbot. It is a system that, given a goal, decides what steps to take and uses tools (query the CRM, draft and send a message, update a record, issue a refund) to get there. That ability to act is what makes agents valuable and what makes them risky. A chatbot that gets it wrong gives a wrong answer. An agent that gets it wrong can take a wrong action.
Our position: most needs in a mid-sized company are better served by rules and fixed workflows, some of which include AI steps. An agent earns its place when the path cannot be set in advance, the outcome can be verified, and the cost of an error is bounded. And whenever one is deployed, accountability stays with the company. This article covers the whole business: support, sales, operations, finance, and internal services. For the marketing-specific view, see AI Agents for Marketing: When to Implement and When to Hold Off.
Chatbot, workflow, agent: three different things sold under one label
The market calls almost any AI product an “agent.” Before evaluating a vendor, separate three concepts.
| Concept | What it does | Who decides the steps | What can go wrong |
|---|---|---|---|
| Chatbot / conversational assistant | Answers questions in a conversation, often from a knowledge base. | The person chatting; the model only responds. | A wrong or invented answer that the customer takes as true. |
| Automated workflow (with or without AI) | Runs a fixed sequence of steps; at certain steps it may use a model to classify, extract, or summarize. | The person who designed the workflow. | A model step fails, but the path is predictable and auditable. |
| AI agent | Pursues a goal, chooses which tools to use and in what order, and acts on external systems. | The model, within the limits it has been given. | A wrong action on a real system, compounding errors, misuse of permissions. |
Definitions based on OpenAI, A practical guide to building agents (2025), and Anthropic, Building effective agents (2024). The risk column is Maccam Network's editorial standard.
The two leading labs agree on the essentials. OpenAI describes agents as systems that independently accomplish tasks on your behalf, where an LLM manages workflow execution and makes decisions, and which have access to tools to interact with external systems, always operating within clearly defined guardrails. It adds a useful precision: applications that integrate LLMs but don’t use them to control workflow execution, such as simple chatbots, single-turn LLMs, or sentiment classifiers, are not agents. Anthropic separates workflows, orchestrated through predefined code paths, from agents, where the LLM dynamically directs its own process and tool usage.
The practical consequence is that “we have an agent” says nothing about the risk level. The list of actions it can take, and the permissions it holds, does. To see how this applies to one concrete sales flow, read AI marketing automation: from inquiries to opportunities.
Four questions to decide whether you need an agent
OpenAI suggests prioritizing workflows that have resisted traditional automation, with three signals: complex decision-making involving judgment and exceptions (its example: refund approval in customer service), rules that have become unwieldy and costly to maintain (vendor security reviews), and heavy reliance on unstructured data (processing a home insurance claim). Its explicit advice is to validate that your use case meets these criteria clearly, and otherwise a deterministic solution may suffice.
We add a management condition of our own. An agent is justified if you can answer yes to all four questions:
- Can the path not be fixed in advance? If it can, a workflow is cheaper, faster, and easier to audit.
- Can you verify whether the result is correct? Anthropic notes that coding agents work well in part because solutions can be verified with automated tests. Without a way to verify, you will not know when the agent degrades.
- Is the cost of an error bounded and, ideally, reversible? If not, the action must go through human approval.
- Do volume and value justify configuring, evaluating, and supervising it? An agent is not a product you install. It is a process you operate.
Use cases by business function
The table lists plausible examples and, for each, the autonomy level we recommend at the start. It is a decision aid, not a promise of results: value depends on data quality, process design, and supervision.
| Function | Use case | Why it fits | Recommended starting autonomy |
|---|---|---|---|
| Customer support | Resolve inquiries with access to the knowledge base and the customer's history; handle simple changes. | Conversation plus tool access. Anthropic cites support as a natural case, with success measured by resolutions. | Answers from verified information. Actions with financial impact need approval, and escalation to a person is always available. |
| Sales and lead handling | Qualify inquiries, prepare account briefs, propose meetings. | Free text plus fit criteria, with a measurable outcome in progression to opportunity. | Proposes and prepares; the seller decides. More in the inquiry-to-opportunity flow. |
| Marketing | Monitoring mentions and competitors, campaign analysis, report preparation. | Research tasks with reviewable output. | Read-only plus recommendations. See AI agents for marketing. |
| Operations and back office | Document review, data reconciliation, processing requests that involve rules and exceptions. | Unstructured data and complex rules; OpenAI's example is processing an insurance claim. | Prepares the file; executes only reversible actions; everything else needs approval. |
| Finance and controls | Anomaly detection, review of suspicious transactions, close support. | OpenAI illustrates with fraud analysis: a rules engine works like a checklist, while an agent evaluates context. | Flags and explains; does not block or pay without approval. |
| Technology and engineering | Code assistance backed by automated tests, incident triage. | Results are verifiable with tests, which lets the agent iterate. | Works in sandboxed environments; changes go through review before production. |
| People and hiring | Scheduling, answering internal questions from current policy. | Low-risk, high-volume tasks. | Be cautious with any decision about people: the European Commission lists employment among the sensitive areas under the AI Act. Requires specific legal review. |
Sources for the examples: OpenAI (2025), Anthropic (2024), European Commission (2026). Autonomy levels are a Maccam Network recommendation.
Look at what the working cases share: unstructured information, decisions with exceptions, and an output you can verify. And what the failing cases share: a promise of “full autonomy” with no clear way of knowing whether the agent is right.
Five concrete risks, not generic ones
1. Compounding errors. An agent that makes a small mistake at an intermediate step can build the following steps on top of it. Anthropic warns of higher costs and the potential for compounding errors in agents, and recommends extensive testing in sandboxed environments with appropriate guardrails. Countermeasures: step limits, intermediate validation, and checkpoints.
2. Excessive permissions and manipulation. The OWASP project lists “Excessive Agency” (LLM06:2025) among the risks of LLM applications. It arises when a model can call functions or interface with other systems and is granted too much functionality, too many permissions, or too much autonomy, and it can lead to damaging actions from unexpected, ambiguous, or manipulated outputs, whether the cause is a model error or a prompt injection attack. Its recommended mitigations include minimizing the tools available, avoiding open-ended tools such as shell commands, limiting permissions to what is essential, and requiring human approval for high-impact actions. For risks specific to agentic systems, the OWASP GenAI Security Project published the Top 10 for Agentic Applications 2026 on December 9, 2025, built with input from more than a hundred specialists; it is a technical reference for your security team or vendor.
3. Liability. The company answers for what its agent says and does. See the next section.
4. Transparency toward the customer. If the agent interacts with people, there are transparency obligations and expectations. In the European Union, Article 50(1) of the AI Act requires people to be informed that they are interacting with an AI unless that is obvious from the context. Even where no such rule applies, hiding it erodes trust; we cover this in how to use AI to win customers without losing authenticity.
5. Drift and cost. Agents depend on models, tools, and data that change. Without a set of test cases and periodic review, quality can degrade unnoticed and resource consumption can creep up. It is not a dramatic risk, but it is the most common one.
Who is liable when the agent gets it wrong: the Air Canada case
Among documented decisions on liability for a conversational AI, one of the most discussed is Moffatt v. Air Canada, 2024 BCCRT 149, decided on February 14, 2024 by British Columbia’s Civil Resolution Tribunal in Canada.
The facts, as stated in the decision itself: a customer asked the airline’s chatbot about fares after a family member died, and the chatbot told him he could apply retroactively for a bereavement fare. The airline’s actual policy did not allow it. He bought the ticket relying on that information and the airline later refused a refund. Air Canada argued, among other things, that the chatbot was a separate legal entity. The tribunal rejected that, reasoning that a chatbot is part of the company’s website and that Air Canada is responsible for all the information on it, whether it comes from a static page or a chatbot. It found negligent misrepresentation and ordered Air Canada to pay C$812.02 in total: C$650.88 in damages, C$36.14 in pre-judgment interest and C$125 in tribunal fees (paragraph 44 of the decision).
Three cautions in reading it. It is a small-claims tribunal decision, not binding precedent, and its reach depends on jurisdiction. The chatbot in that case only answered questions; it was not an agent taking actions. And the amount is modest. What matters is the reasoning: neither technical complexity nor a system’s autonomy shifts responsibility away from the company that puts it in front of the customer. If that holds for a chatbot that only talks, the reasonable expectation for an agent that acts is higher, not lower. None of this is legal advice; consult an attorney.
Minimum viable governance: levels of autonomy and controls
You don’t need a department to govern an agent. You need an explicit decision about how much autonomy it has and which controls come with it. We propose four levels.
| Level | What the agent does | Minimum controls |
|---|---|---|
| 0. Assistant | Suggests; a person executes. | Sample-based review of suggestion quality. |
| 1. Prepare and wait | Prepares the action (draft, case file, order) and waits for approval. | Prior human approval; a log of what was approved and rejected. |
| 2. Execute the reversible | Performs low-risk, reversible actions within clear limits. | Minimal permissions, full action logging, periodic sampling, rollback and stop mechanisms. |
| 3. Execute the critical | Sensitive, irreversible, or high-impact actions. | Not recommended in a first phase. If ever considered: human approval, quantitative limits, and legal review. |
Maccam Network working framework. OpenAI's guide recommends human oversight for high-risk actions (for example, canceling orders, authorizing large refunds, or making payments) and escalating to a person when failure thresholds are exceeded.
To sequence these decisions, a widely known reference is NIST’s AI Risk Management Framework (AI RMF 1.0), released January 26, 2023, voluntary, and organized around four functions: Govern (who decides and who answers), Map (what it is used for and in what context), Measure (how performance and risks are assessed), and Manage (what is done about identified risks). NIST released a Generative AI Profile (NIST-AI-600-1) on July 26, 2024, and its page notes that the AI RMF is being revised as part of the White House AI Action Plan, so check for the current version (page checked October 9, 2026).
For a mid-sized company’s agent, that framework translates into five simple decisions:
- Govern: a named owner for each agent and a list of what it can and cannot do.
- Map: the process, the data it touches, the people affected, and what happens if it fails.
- Measure: a set of test cases with known answers and a periodic review of real samples.
- Manage: failure thresholds that escalate to a person, a kill switch, and an incident log.
- Document: which model, which version, which tools, and which permissions it has at any given time.
From pilot to production: four gating criteria
A well-designed pilot limits scope (one process, one channel), defines in advance what success means, and stays at level 0 or 1 while you measure. Before granting more autonomy, we recommend requiring:
- Results measured against a set of test cases and against human decisions, with each error analyzed.
- Actions and permissions reviewed: nothing the agent doesn’t need.
- An escalation and stop procedure that has been tested, not just written.
- An owner accountable for outcomes, and legal review wherever customers, employees, or personal data are involved.
Common mistakes
- Buying “an agent” before defining the process. Without a documented process, the agent automates ambiguity. It is the same lesson as the three preconditions before automating.
- Starting with the highest-impact case. Start where errors are cheap and reversible.
- Granting broad permissions “to make it work.” It is the classic origin of excessive agency.
- Measuring only time saved. Also measure errors, escalations, and the satisfaction of the people affected.
- Forgetting the customer. If the customer doesn’t know they are talking to an AI, or can’t reach a person, the savings are paid for in trust.
- Treating adoption as a technology project. It is a change in process and responsibilities; we cover it in how to implement AI in a mid-sized company without burning your budget.
Next step
If you are evaluating AI agents for your company, we start with what matters before the technology: which processes have the right characteristics, what level of autonomy is reasonable, and what governance they need. Learn about our AI Agents service or get in touch.
Sources
- OpenAI (2025). A practical guide to building agents. cdn.openai.com
- Anthropic (2024). Building effective agents (December 19, 2024). anthropic.com
- OWASP GenAI Security Project (2025). LLM06:2025 Excessive Agency, in OWASP Top 10 for LLM Applications 2025. genai.owasp.org
- NIST (2023). AI Risk Management Framework (AI RMF 1.0) and NIST (2024) Generative AI Profile (NIST-AI-600-1). nist.gov
- British Columbia Civil Resolution Tribunal (2024). Moffatt v. Air Canada, 2024 BCCRT 149 (February 14, 2024). decisions.civilresolutionbc.ca (the decision’s original text; the tribunal’s server may block automated downloads).
- OWASP GenAI Security Project (2025). OWASP Top 10 for Agentic Applications for 2026 (December 9, 2025). genai.owasp.org
- European Union. AI Act, Article 50: Transparency obligations (consolidated text, AI Act Service Desk). ai-act-service-desk.ec.europa.eu
- European Commission (2026). AI Act: regulatory framework for AI (updated August 3, 2026). digital-strategy.ec.europa.eu
Editorial note: this article offers general management criteria and is not legal advice. Sources verified as of October 9, 2026.
Preguntas frecuentes
A chatbot answers within a conversation. An agent also decides which steps to take and uses tools to act on external systems until a goal is met. OpenAI defines agents as systems that independently accomplish tasks on your behalf, and says applications that integrate LLMs without using them to control workflow execution (simple chatbots, single-turn LLMs, sentiment classifiers) are not agents. Anthropic describes agents as systems where the LLM dynamically directs its own processes and tool usage, in contrast to workflows, which follow predefined code paths.
Tasks that involve unstructured information, decisions with exceptions, and steps that cannot be fixed in advance, as long as the result can be verified and the cost of an error is bounded: customer support with tool access, document review, reconciliations, report preparation, and software development with automated tests. Sensitive, irreversible, or high-stakes actions (payments, large refunds, cancellations) should require human approval, as OpenAI's guide recommends.
In practice, the company that deploys it. In Moffatt v. Air Canada (2024 BCCRT 149), British Columbia's Civil Resolution Tribunal rejected the idea that a chatbot could be treated as a separate entity and held that Air Canada was responsible for the information on its website, including the chatbot's. It is a small-claims tribunal decision, not binding precedent, but it illustrates a likely line of reasoning. This is not legal advice; consult an attorney in your jurisdiction.
The best known is NIST's AI Risk Management Framework (AI RMF 1.0), released January 26, 2023, voluntary, and organized around four functions: Govern, Map, Measure, and Manage. NIST also published a Generative AI Profile (NIST-AI-600-1) on July 26, 2024. In the European Union, the AI Act sets transparency obligations for systems that interact with people and stricter rules for certain high-risk uses. None of these replaces business judgment or legal advice.
It depends on the impact of its actions. As a working rule, an agent that only suggests needs after-the-fact review; one that prepares actions needs approval before execution; one that executes low-risk, reversible actions needs logging, limits, and sampling; and irreversible or high-impact actions should keep human approval. It is also wise to set failure thresholds after which the agent hands the case to a person.
When a simple rule or fixed workflow solves the problem, when you cannot verify whether the result is right, when an error is costly or irreversible and cannot be supervised, or when volume does not justify the effort of configuring, evaluating, and monitoring it. Anthropic recommends starting with the simplest solution and adding complexity only when it demonstrably improves results.
New ideas, analysis and research — directly to your inbox.
Subscribe to receive new Insights publications and other selected content from Maccam Network. No spam. Unsubscribe at any time.
Shall we talk about your business?
Let's talk about what your business needs.
A 30-minute conversation is enough to understand the context, identify the problem and see if we are the right team to help you.