Beginner level7 min readTo learn at another level, choose it before you start the course.
After this lesson you canDecide whether a task is safe to hand to an agent, and set the access, approvals, limits and logs it needs.
The short answerIt can be, if the agent can only do what the task needs, asks you before important steps and keeps a record you can check. Start with small tasks that are easy to undo.
In simple words
- Start with tasks that are easy to undo.
- Give it only the access the task needs.
- Make it ask before it pays, sends or deletes.
- Check what it did, at least at the start.
The short answerLetting an agent act for you is a risk decision: weigh what it can reach, how easily its actions can be undone and how you will notice mistakes. Least access, approval for high-impact actions, spending limits and logs make most everyday uses reasonable.
In simple words
- Rate the access: which data and accounts can the agent reach?
- Rate the actions: can they be undone, and what does a mistake cost?
- Control: least access, approval for high-impact steps, spending limits.
- Notice: logs, alerts and regular checks of what it did.
The short answerTreat an agent deployment like any privileged automation: threat-model it, minimise its blast radius, gate irreversible actions and monitor it. Model quality lowers the error rate; architecture decides how much an error or attack can cost.
Key points
- Threat model: assets, untrusted inputs, outbound channels, attacker goals.
- Blast radius: scoped credentials, allowlists, separate accounts, budget caps.
- Gates: tiered approval by action class, with tamper-evident logs.
- Rules: EU AI literacy and AI disclosure, US voice-call consent; frontier incident duties sit with developers.
Risk: High
It can do things that are hard to undo without asking you first.
- Make it ask you before it pays, deletes or sends in your name.
- Set a spending limit on the card or wallet it uses.
- Choose an agent that shows a log of every step, and check the first runs.
Words to know
- Irreversible action
- A step that cannot be undone, such as paying or sending a message.
- Read-only access
- Permission to look at data, but not to change, delete or send it.
- Log
- A record of every step an agent took, so you can check its work.
Five questions before you say yes
What can it reach: your email, your files, your bank account? Can its actions be undone, or are they irreversible? Will it ask before it pays, sends or deletes something?
Can you see a log, a list of every step it took? And who pays if it makes a mistake? If you cannot answer these questions, start with a smaller task.
Start small
Good first tasks are easy to check and easy to undo: comparing prices, drafting an email that you send yourself, or sorting files into a new folder.
Wait with tasks that move money, send messages in your name or delete things. Give the agent more freedom only after it has done well many times.
Keep the keys
Use read-only access where you can, so the agent can look but not change anything. Set a spending limit on any card it uses, and turn off access you no longer need.
Remember that the agent acts in your name, so its mistakes can quickly become your problem. In California, a company cannot escape blame in court by saying its AI acted on its own.
Try it yourself
Pick a task you might give an agent, and tick what it could do in the checklist. Then tick what protects you, and read how the risk and the advice change.
Check yourself
You want to try an agent for the first time. Which task is the best start?
Your agent is planning a birthday dinner for you. Before which step should it ask you?
You let an agent book train tickets with your card. Why set a spending limit on that card?
Words to know
- Irreversible action
- A step that cannot be undone, such as paying or sending a message.
- Read-only access
- Permission to look at data, but not to change, delete or send it.
- Log
- A record of every step an agent took, so you can check its work.
Map what it can reach
List every account, file and tool the agent will use. Private data plus text from strangers plus a way to send data out is the mix that makes prompt injection dangerous.
Prefer narrow, read-only access, such as one folder instead of your whole drive. Remove access when the task is done.
Sort actions by how easily they are undone
Reading and drafting are easy to undo. Sending, paying, deleting and posting are not. OpenAI’s guide says sensitive, irreversible or high-stakes actions should involve a person until the agent proves reliable.
A practical rule: let the agent prepare everything, and keep the final click for yourself when money, messages or deletions are involved.
Make mistakes visible
Choose agents that show a log of every step and tool call. Check the first runs closely, then sample later ones, and set alerts for payments or new contacts.
Set hard limits: a spending cap, a maximum number of steps and a time limit. They stop a confused agent before it does much damage.
At work, rules apply
In the EU, businesses must help staff who use AI for them understand it well enough to use it safely; this duty is called AI literacy. Since August 2, 2026, AI systems that talk with people must tell them they are talking to an AI, unless that is obvious.
In the US, calls with AI-generated voices need the called person’s prior consent under federal robocall rules. California also bans bots that pretend to be human to sell things or influence votes, and a company there cannot defend itself in court by saying its AI acted on its own.
Try it yourself
Pick a task you might give an agent, and tick what it could do in the checklist. Then tick what protects you, and read how the risk and the advice change.
Check yourself
Your agent handles supplier invoices. Following OpenAI’s guide, which step should wait for a person?
An agent needs to sort the files in one project folder. What access should it get?
Overnight, your agent kept retrying a failing booking site hundreds of times. Which setting would have stopped it?
Threat-model the deployment
Start with assets, untrusted inputs, outbound channels and attacker goals. Any path from untrusted input to an outbound channel that carries private data is a candidate exfiltration route.
Write down what the agent must never do, then enforce it outside the model. Instructions in the prompt are guidance, not controls.
Minimise the blast radius
Use separate service accounts with narrow OAuth scopes, domain and tool allowlists, and spending caps. Run browsing and code in disposable sandboxes without standing access to secrets.
Time-box credentials and revoke them after the task. A compromised agent should find little to steal and few places to send it.
Gate actions by class
Classify actions: read, draft, reversible write, irreversible write and external communication. Auto-approve reads and drafts, log reversible writes so they can be undone, require confirmation for the last two classes, and alert on anything new.
Log each tool call with its inputs, outputs and the approval decision, ideally append-only. Without logs, a misbehaving run cannot be reconstructed or explained.
The legal frame
In the EU, businesses must support the AI literacy of staff using AI. Since August 2, 2026, providers must make AI that talks with people say so, unless that is obvious. High-risk uses, such as hiring, bring more duties from December 2027.
In the US, AI-voice calls need prior consent under the FCC’s 2024 ruling. Incident-reporting laws in California, New York and Illinois bind frontier developers, not the businesses using their agents. California also bars defendants from arguing that the AI acted on its own and so caused the harm (AB 316, in force since January 2026).
Try it yourself
Pick a task you might give an agent, and tick what it could do in the checklist. Then tick what protects you, and read how the risk and the advice change.
Check yourself
An attacker fully controls your agent through an injection. Which design limits what they can do?
Your support agent reads tickets, drafts replies and emails customers. If you gate actions by class, which step needs confirmation?
A shop in the EU builds a support agent that chats with its customers. Which of these duties applies?
Sources
- A practical guide to building agents (OpenAI)
- LLM06:2025 Excessive Agency (OWASP Top 10 for LLM Applications)
- Building effective AI agents (Anthropic, December 2024)
- FCC Declaratory Ruling 24-17 on AI voices in calls (FCC)
- AB 316: no defence that the AI acted on its own (California Legislature)
- Nonprofit sues OpenAI over its agents’ Hugging Face hack (Silicon AI News)
This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.