Writing · AI and human responsibility
AI Agent Traps and Human Responsibility
By Richard K. Marshall · Originally published
When an agent gets tricked, "the AI did it" won't be an answer.
In April 2026, Google DeepMind published a paper called "AI Agent Traps." It's a wake-up call for anyone building or using autonomous agents.
It's the first systematic framework for six kinds of environmental attack that can hijack web-browsing agents. Not by breaking the model, but by manipulating the world around it.
The six traps
- Content injection. Hidden HTML or CSS that agents "see" and humans don't.
- Semantic manipulation. Subtle phrasing that bends the agent's reasoning.
- Memory poisoning in retrieval (RAG) systems.
- Behavioral hijacks.
- Multi-agent systemic failures.
- Human-in-the-loop exploits.
The message is clear. As agents get the autonomy to browse, decide and act, the attack surface moves from the model to the environment.
And when something goes wrong, everyone will ask the same question we need to answer now: who is responsible?
Diffused responsibility is the real danger
AI doesn't remove human responsibility. It concentrates it.
When an agent is tricked into doing harm, "the AI did it" isn't a defense. Without governance in place before agents are woven into workflows, accountability evaporates. Leaders wake up to outcomes they never intended, with no audit trail and no human oversight.
That's why governance readiness isn't optional. It's the only reliable way to make sure a human stays responsible for every AI outcome. No exceptions.
The Principle, put to work
The Marshall AI Governance Readiness Standard (MAGRS) is a lightweight framework for exactly this. It helps organizations bring in powerful AI, agents included, while keeping human accountability in front.
At its core is the Marshall Principle:
Artificial intelligence may assist human decision-making, but responsibility always remains with humans. Authority cannot be automated.
MAGRS puts it into practice through three pillars:
- Visibility. Know exactly where and how AI and agents are used across the organization. No blind spots.
- Boundaries. Set clear, enforceable limits on what AI may do, especially high-risk activities like autonomous browsing or independent decisions.
- Accountability. Make sure a human is always responsible for AI-assisted outcomes, with audit trails and regular review.
How that blocks the traps
With visibility, you spot agent use before it scales.
With boundaries, you can require human review or restrict actions, which blocks environmental attacks.
With accountability, responsibility can't be diffused.
Why now
Agents are moving from research papers into real workflows faster than most people realize. The DeepMind paper, and open-source red-teaming work like Strix, show that model-level safety alone isn't enough. Threats from the environment demand organizational readiness.
Start with visibility. This week, find out every place an AI agent can browse, click or send on your organization's behalf.
AI may assist. Humans are always responsible.
— Richard K. Marshall Marshall Intelligence · Lexington, Kentucky
The MAGRS framework and readiness assessment are at marshall.net/magrs.
Originally published on X: https://x.com/RichMarshall/status/2046259554846179830 · . Refreshed .