AI agents: what they are, how they work, and when to use one

Idea Labz · Oct 5, 2026 · 18 min read

  • AI STRATEGY & EXECUTION
  • AI AUTOMATION

The assistants you use now come with agent modes, and software vendors are selling agents that promise to take work off your hands. As of October 2026, ChatGPT has Work, which OpenAI describes as an agent for longer, multi-step work. Anthropic offers Claude Cowork, Google offers Gemini Spark, and Microsoft’s Copilot can carry out multi-step tasks inside Word, Excel and PowerPoint. Most explanations stop at “autonomous software that uses tools”, which doesn’t tell you what an agent would actually do in your business.

This guide explains how an AI agent works, follows one agent through a real task step by step, and covers when a workflow is the better choice, how much to let an agent do, and how to build your first. Agents are one part of AI automation, and we’re a UAE-based team that builds them into the tools companies already use. The main example below is one of ours.

  • What is an AI agent?
  • How does an AI agent work?
  • One agent run, step by step
  • AI agent or workflow? When you don’t need an agent
  • How much should an AI agent be allowed to do?
  • Types of AI agents
  • AI agents you can use today
  • How to build your first AI agent
  • AI agents: common questions
  • What to do next

What is an AI agent?

An AI agent is software that’s given a goal, uses a language model to decide its own next step, acts through tools such as a database or an email system, checks what came back, and repeats until the goal is met or it needs a person. Choosing the next step is what makes it an agent. A script or a workflow follows steps someone wrote in advance, while an agent picks its steps as it goes.

For example, if you ask a chatbot why a customer stopped ordering, it answers from what it already knows. Give the same question to an agent with access to your order system, and it looks up the customer’s orders, notices which product they stopped buying, reads the support tickets about that product, and then answers.

Four kinds of AI tool are easy to mix up, and the clearest difference between them is who decides the next step:

Who decides the next stepCan it act on other systems?Example
ChatbotYou, one message at a timeNoAnswers a question about your returns policy
AI assistantYou, with the assistant’s suggestionsSometimes, when you ask it toDrafts a reply in your inbox when you ask
Workflow with an AI stepWhoever designed the workflow, in advanceYes, along a fixed pathReads every expense receipt and files it the same way
AI agentThe model, within limits you setYes, through the tools it’s givenFinds out why a customer’s orders dropped and drafts a follow-up

You’ll also see the term agentic AI. It usually describes the wider approach of building systems from agents, plus the tools, memory and checks around them.

How does an AI agent work?

An AI agent works in a loop. It takes a goal, makes a plan, takes one action, looks at the result, and decides what to do next, until it reaches a stop condition or hands the task to a person:

  1. Goal
  2. Plan
  3. Act
  4. Observe
  5. Repeat, stop, or hand over

IBM’s guide describes a common pattern for this loop, ReAct, in which the model alternates between reasoning about the next step and acting on it. Every agent built this way has five parts.

The model

The model is a large language model (LLM) that reads the goal, the instructions and everything gathered so far, and decides the next step. A stronger model plans better and costs more for every step, so it’s worth matching the model to the task.

The instructions

The instructions, often called the system prompt, tell the agent who it’s working for, what it’s trying to achieve, which tools it may use, and when to stop or ask. For example, an instruction might say that the agent may draft replies but never send them, and that it must say so when it can’t find a record.

The tools

Tools are how an agent acts. Each tool is a function the agent can call, such as searching the order database, reading a support ticket, checking a delivery, or saving a draft. The agent can only use the tools it’s given, which makes the tool list its most important safety setting.

The Model Context Protocol (MCP) is the common standard for connecting tools to agents. Anthropic open-sourced it on 25 November 2024, and a tool offered over MCP works with any agent or assistant that supports the protocol. If your data lives in a backend you own, an MCP server is how an agent reaches it, as we show in our guide to connecting AI assistants to Payload CMS over MCP. A second open standard, Agent2Agent (A2A), covers agents communicating with other agents. Google developed it and donated it to the Linux Foundation.

Memory

Memory is what the agent knows while it works. Within one task, that’s its context window, which holds the goal, the instructions and every result so far. For documents too large to hold at once, retrieval (often called retrieval-augmented generation, or RAG) fetches the relevant passages when the agent needs them. Some agents also keep notes between tasks, such as a customer’s preferences, so the next run starts with them.

Guardrails

Guardrails are the limits around the loop. They include a cap on the number of steps and on spend, points where a person must approve, and a log of every tool call and result. Without them, an agent that misreads its goal can keep going, and keep costing money, long after a person would have stopped.

Now that you know the parts, the clearest way to see them working together is one real run.

One agent run, step by step

We built an agent for an online store that we call the post-order agent pipeline. It works out what each customer’s next engagement email should be, writes it for that customer, and puts it in front of the team for approval. The store and its figures stay private, so the stage names, checks and timings below show how a run like this works. They aren’t the store’s own settings.

A run has eight steps:

  1. An order completes
  2. The agent reads the customer
  3. It follows the customer’s stage
  4. It qualifies the next step
  5. It reads the customer’s recent orders
  6. It writes a draft
  7. The team approves it
  8. A rejection comes back with a comment

1. An order completes

The run starts when a customer’s order completes. In a build like this, the trigger might be the order being delivered, or a set number of days after delivery, depending on when the store wants to follow up.

2. The agent reads the customer

The agent pulls the customer’s record first. The store defines its customer stages by order value, lifetime value, previous orders and similar factors, so those are what the agent reads. For example, it might find that this is the customer’s third order in four months, and that their lifetime value has just passed the store’s threshold for its top stage.

3. It follows the customer’s stage

The store has defined its stages, and the agent picks each next step according to the stage the customer is at. A stage might be a first-time buyer, a repeat buyer, or a high-value customer, and each one opens a different set of possible next steps.

4. It qualifies the next step

Before it writes anything, the agent runs several checks to decide what this customer should be sent next. For example, it might check whether the customer has had an email this week, whether their order included something that usually needs a refill, and whether they’ve left a review. Two customers who placed the same order can get different next steps, because the agent chooses its path from what it finds for each one.

5. It reads the customer’s recent orders

The agent then looks at what the customer ordered in the last few days and writes the email around it, so it reads as written for that customer.

6. It writes a draft

Every email the agent produces is a draft, and sending stays with the team. The simplest way to enforce a rule like this is to leave sending out of the agent’s tool list, so the limit doesn’t depend on the agent following its instructions.

7. The team approves it

The drafts sit in an approval queue built into the store’s backend dashboard. The team works through them quickly, one after the next, and a draft they approve goes out.

8. A rejection comes back with a comment

When a reviewer rejects a draft, they add a comment, the agent redoes that step, and the draft comes back for approval. Kept together, those comments become test cases, because each one records a real mistake to check the next version of the agent against.

StepWhat happensWhat it readsWho acts
1. Order completesA run starts for this customerThe orderThe system
2. Read the customerGathers the stage factorsThe customer’s order history and valueThe agent
3. Follow the stageNarrows the possible next stepsThe store’s stage definitionsThe agent
4. Qualify the next stepDecides what the next email should beRecent emails, order contents, reviewsThe agent
5. Read recent ordersShapes the email around themOrders from the last few daysThe agent
6. DraftWrites the email and stopsEverything gathered so farThe agent
7. ApproveSends an approved draft onThe draftThe team
8. Reject and redoRewrites the step from the commentThe reviewer’s commentThe agent, then the team

When the review step can switch off

The manual review belongs to the stage when the store is still building confidence in the agent’s output. A confidence score checker and a set of rules run alongside it, and a draft that clears a high score and passes the rules is treated as good. Once the business is confident in the results, it can turn manual review off and let the emails go out automatically, as long as the rules and the evaluation keep running often enough to catch mistakes. At hundreds of orders a day, approving every email by hand stops being practical. The agent moves towards almost fully automatic sending as it learns, and once it gets there, the team’s job shifts from approving drafts to evaluating and observing the agent.

Where a run like this can go wrong

  • The wrong stage. If order data syncs late, the agent can place a customer in last month’s stage. Reading live records at the start of every run prevents it.
  • A repeat email. Without a record of what was already sent, two runs can produce near-identical emails a week apart. Read access to sent messages fixes it.
  • An invented detail. A model can write a confident sentence about a product the customer never bought. Limiting the email to facts from the order records, and asking the reviewer to check them, catches it.
  • Too many emails. Every run can decide an email is due. A frequency cap per customer, applied outside the model, keeps the total sensible.

AI agent or workflow? When you don’t need an agent

You don’t need an agent when the steps are the same every time. Anthropic, the company behind Claude, advises finding the simplest solution possible and adding complexity only when it’s needed, which can mean not building an agent at all. In its terms, workflows are systems where models and tools follow predefined code paths, while agents direct their own process and tool use. Agentic systems often trade speed and cost for better results, so the trade has to make sense for the task.

Four questions settle it for a given task:

  1. Is the path the same every time? If it is, a workflow is cheaper, faster, and easier to test.
  2. Does the next step depend on what the system finds? That’s where an agent helps.
  3. What does a wrong step cost, and can it be undone? The higher the cost, the narrower the agent’s permissions should be.
  4. Can you test it on real past cases? If you can’t say what a good result looks like, the agent can’t either.
TaskBest fitWhy
Reading expense receipts and posting themWorkflow with AI stepsThe steps never change, and the model only reads
Pulling delivery dates from supplier emails into the stock systemWorkflow with one AI stepThe path is fixed, and only the reading needs a model
The post-order agent aboveAn agent inside a defined workflowThe business fixes the stages, and the next step varies by customer
Planning a client visit across calendars, flights and a budgetAgentEach answer decides the next thing to look up

Reading expense receipts is the worked example in our step-by-step guide to how AI automation works, where a model reads each receipt and set rules decide where it goes.

How much should an AI agent be allowed to do?

An AI agent should be allowed to do what its results on real cases have shown it does safely, with anything that can’t be undone kept behind a person until its record and regular evaluation show it can act alone. In practice, that comes down to its level of autonomy and the permissions on each tool.

Four levels of autonomy

  1. Suggest. The agent proposes, and a person does the work.
  2. Act after approval. The agent prepares the action, a person approves it, and then it runs. The post-order agent starts at this level for sending.
  3. Act and report. The agent acts on its own where the action can be undone, and a person reviews a summary afterwards.
  4. Act alone within limits. The agent acts without review, inside tight limits on what it can do, how much, and how often, while the team evaluates and observes its results. The post-order agent can move to this level once its scores and rules show the store it’s ready.

An agent moves up a level when its record on real cases supports it, and moves back down after a serious mistake. The checks a person makes at each level, before the action, on exceptions, or after the fact, are set out in our guide to where people stay in the loop.

Permissions on every tool

Give each tool the narrowest access the task needs. Read access comes first, write access only where the task requires it, and delete access almost never. Irreversible actions, such as sending an email, issuing a refund, or changing a price, start behind a person’s approval and come out from behind it only on evidence.

For example, the post-order agent reads orders and customer records and writes drafts, while sending stays with the team.

Prompt injection

Prompt injection is text inside the content an agent reads that tries to give it new instructions. Any agent that reads emails, support tickets or web pages can meet it. For example, OpenAI’s help page on its agent describes an agent planning a group dinner from your calendar and emails, which meets a comment on a web page telling it to fetch a password reset code from Gmail and send it to a malicious site. Prompt injection is first on the OWASP Top 10 for LLM applications for 2025.

The defence is to treat everything the agent reads as data, never as instructions, and to keep risky tools behind approval, so a successful injection has nothing dangerous to call.

Logs and limits

Log every tool call and every result, so anyone can replay why the agent did what it did. Set a maximum number of steps and a spending limit for each run, and keep a way to switch the agent off.

Judgment about what matters stays with the person using the agent, as our founder argues in AI agents can’t do 100% of your work.

Types of AI agents

The textbook types describe how an agent decides. IBM lists five, from simple reflex agents to learning agents, and AWS adds hierarchical agents and multi-agent systems to make seven. Each one has an everyday business version:

TypeHow it decidesA business example
Simple reflexResponds to the current input with a fixed ruleAn auto-reply that sends opening hours when someone asks for them
Model-based reflexKeeps track of things it can’t see right nowA stock alert that knows what’s on order as well as what’s on the shelf
Goal-basedChooses actions that move it towards a goalAn agent choosing each customer’s next step towards a repeat order
Utility-basedWeighs options and picks the best trade-offA scheduler choosing the delivery slot that balances cost and speed
LearningImproves from feedbackA support agent whose corrected replies are added to its examples
HierarchicalA lead agent splits work for othersA manager agent that hands research and writing to two sub-agents
Multi-agentSeveral agents work towards one resultA research agent, a writing agent and a checking agent on one report

At work, you’re more likely to hear agents described by the job they do.

Conversational agents

A conversational agent talks to customers or staff, looks things up, and acts within limits. For example, an HR agent answers a staff member’s question about annual leave, checks their remaining balance, and passes the leave request to their manager to approve.

Research agents

A research agent gathers information from several sources and writes it up. For example, before a supplier negotiation, it gathers a year of orders, late deliveries and price changes for that supplier, and writes a one-page brief with each figure linked to the record it came from. A person reads the brief before acting on it.

Coding agents

A coding agent writes, tests and fixes code in a repository. A developer reviews every change before it’s merged, because the agent can produce code that passes its own tests and still does the wrong thing.

Browser and computer-use agents

A browser agent operates websites the way a person does, by clicking, typing and filling in forms. It suits sites with no API, and it needs the closest watching, because the pages it reads can carry prompt injection.

Multi-agent teams

A multi-agent team splits one job across agents with separate roles, such as one that researches, one that drafts, and one that checks. Every hand-off adds cost and another place for an error, so a single agent is the better start.

AI agents you can use today

Agents come in four main forms, from features inside the assistant you already use to toolkits developers build with. As of October 2026, these are the main places to find one:

KindExamplesWho it suitsWhat you control
Agent features in assistantsChatGPT Work, Claude Cowork, Gemini Spark, Microsoft 365 CopilotPeople and teams doing research and one-off tasksIts instructions and which apps it may connect to
Agents inside business softwareSalesforce Agentforce, Microsoft Copilot StudioTeams whose process lives in that vendor’s softwareIts set-up, within the vendor’s data and actions
Workflow platforms with agentsZapier Agents, Make AI Agents, n8nTeams connecting many apps with little codeThe steps around the agent, and which apps it uses
Developer toolkitsOpenAI Agents SDK, Claude Agent SDK, Google Agent Development Kit (ADK), LangGraphDevelopers building into their own systemsEverything, from the model and tools to permissions, review screens and hosting

Each row trades speed of setting up against control. An assistant’s agent mode works the same day, while a developer toolkit takes a build but lets you decide every permission and where the data goes.

How to build your first AI agent

Build your first AI agent on a task where the path varies and mistakes are cheap, then widen what it may do as it proves itself. Seven steps take it from an idea to a running agent:

  1. Pick a task where the path varies and mistakes are cheap
  2. Write the goal and the stop condition
  3. List the tools, read-only first
  4. Write the instructions
  5. Collect real past cases as a test set
  6. Run it with a person approving every action
  7. Widen its permissions as the results earn it

1. Pick a task where the path varies and mistakes are cheap

Look for work that reads a lot and drafts something, such as preparing a brief before a client call, drafting replies to common requests, or sorting incoming requests by what they need. If the steps never change, build a workflow instead.

2. Write the goal and the stop condition

Write down what a finished result looks like and when the agent should stop or ask. For example, a briefing agent might stop when every claim in the brief links to a record, or say which record it couldn’t find.

3. List the tools, read-only first

List every system the agent needs and give it read access only. Add write access one tool at a time, when a step can’t be done without it.

4. Write the instructions

Write the agent’s role, its rules, what it must never do, and how it should say it doesn’t know. Short, specific instructions are easier for the model to follow and easier for you to fix.

5. Collect real past cases as a test set

Gather real examples of the task, the awkward ones included, and write down what a good result would have been for each. Run every change to the instructions or the model against them before it goes live.

6. Run it with a person approving every action

Start with every action behind approval, as the post-order agent does for sending. Each rejection, with its reason, becomes another test case.

7. Widen its permissions as the results earn it

Move a single action to the next level of autonomy only when the agent’s record on that action supports it. Irreversible actions move last, and only with regular evaluation running behind them.

Where to build it

There are three places to build an agent:

  • Inside an assistant you already pay for, by giving it instructions and connecting your apps. It’s the fastest start, and it’s limited to what the assistant can reach.
  • On a workflow platform’s agent builder, when the agent needs to act across several apps you already use.
  • As a custom build, when the agent needs your own systems, its own review screen, or data that can’t leave your control. The post-order agent is this kind of build, with its approval queue inside the store’s own dashboard. It’s the kind of work our AI strategy and implementation service covers. If you’re weighing outside help, our guide to what to ask an agency before you hire one sets out the questions and what you should own when the work ends.

AI agents: common questions

Is ChatGPT an AI agent?

ChatGPT is a chatbot in its Chat mode and an agent in its Work mode. OpenAI describes Work as an agent designed for longer, multi-step work and finished deliverables, while Chat answers and waits for your next message.

Who are the big 4 AI agents?

There’s no official list. The phrase usually means the agents from OpenAI (ChatGPT), Google (Gemini), Microsoft (Copilot) and Anthropic (Claude), the four companies most often named in agent comparisons. It’s also used for the agent platforms of the four big accounting firms.

What are the types of AI agents?

The five textbook types are simple reflex, model-based reflex, goal-based, utility-based and learning agents. Lists of seven add hierarchical agents and multi-agent systems.

What is the difference between AI agents and agentic AI?

An AI agent is one system that pursues a goal by choosing its own steps. Agentic AI is the broader term for building with agents, including several agents working together and the tools, memory and checks around them.

How much does an AI agent cost to run?

An agent’s running cost is its model usage, its tools, and the time spent maintaining it. Each step in the loop is at least one model call, so an agent that takes many steps per task costs more than a workflow that does the same task in fewer. Step and spend limits on each run keep the cost predictable.

Can I get an AI agent for free?

You can build one without a licence fee. n8n’s self-hosted Community Edition is free, and toolkits such as LangGraph and the OpenAI Agents SDK are open source under the MIT licence. In assistants, agent features such as Claude Cowork (from Claude’s Pro plan) and Gemini Spark (on Google AI Pro and Ultra) come with paid plans, and any agent you run yourself still pays for its model usage.

Are AI agents safe to use with business data?

They can be, when each tool has the narrowest access the task needs, anything irreversible is approved by a person or covered by regular evaluation, and every action is logged. Check whether your model provider uses your data for training, and treat what the agent reads as data, so a prompt injection can’t steer it.

What to do next

Start by deciding whether your task needs an agent at all. If the path is the same every time, build a workflow. If the next step depends on what the system finds, write the goal and the stop condition, give the agent read-only tools, and keep a person approving every action. Widen its permissions only as its results on real cases earn it.

Use the seven steps above as a checklist for your first agent, starting with the task you’d most like off your desk this month.

References

All links checked and active as of 5 October 2026.

READY TO MAKE A REAL CHANGE?

Let's build it together