Learn · Second course
AI agents and safety
What agents are, how they fail in real use, and how to decide whether to let one act for you.
2 Start the course
The lessons
What is an AI agent?
An AI agent is an AI that does things for you, not just talks. You give it a goal, and it takes the steps on its own, such as searching, clicking or filling in forms.
How do AI agents fail? Real incidents
Real AI agents have gone further than asked, leaked private data and slipped out of the test areas they were meant to stay in. Most cases caused no known harm, but they show why agents need limits.
What is prompt injection, and how do permissions help?
Prompt injection is a trick: someone hides orders in a web page, email or file, and an agent follows them as if you had asked. The best protection today is limiting what the agent can do.
Is it safe to let an agent act for me? A checklist
It can be, if the agent can only do what the task needs, asks you before important steps and keeps a record you can check. Start with small tasks that are easy to undo.
After this course you can
- Describe how an AI agent uses tools in rounds to reach a goal, and say which steps should wait for your approval.
- Name the main ways real AI agents have failed, and explain why each case shows that agents need limits.
- Explain how hidden text can make an agent obey a stranger, and choose permissions that limit the harm.
- Decide whether a task is safe to hand to an agent, and set the access, approvals, limits and logs it needs.