OWASP Founder Jeff Williams: Stop Studying the Robot's Brain. Start Controlling Its Body.

OWASP Founder Jeff Williams: Stop Studying the Robot's Brain. Start Controlling Its Body.

BackerLeader 44 268 482
calendar_today agoschedule4 min read

Jeff Williams, founder of OWASP and founder and CTO of Contrast Security, has spent the last several years watching the security industry pour its attention into one question: can you jailbreak the model, trick it, make it say something dangerous. He thinks that's increasingly the wrong place to be looking. "Once you give the brain a body, there's a much more consequential question: what can the robot actually do?" he said.

The "body," in Williams's framing, is what's increasingly being called an agent harness, the tools, memory, permissions, orchestration, and runtime infrastructure that surrounds a model and turns raw intelligence into something capable of taking real-world action. As agents move from answering questions to executing code, touching databases, and sending money, Williams argues the harness has quietly become at least as important to an agent's actual risk profile as the model powering it, maybe more so.

The same brain, two completely different agents

Williams's central metaphor is worth sitting with because it reframes how you'd even evaluate an agent in the first place. "Imagine putting the same robot brain into two completely different bodies," he said. "Give one a camera and a speaker. Give the other access to your email, GitHub, production systems, corporate databases, and a credit card. They're powered by exactly the same intelligence, but they are radically different agents."

That has a direct, practical consequence for anyone deploying agents today: "Evaluating an agent based solely on its model is increasingly meaningless. You have to know what body we've attached to the brain." A vendor's model benchmark scores tell you almost nothing about what that model can actually do once it's wired into a real system with real credentials.

What belongs in the harness, and why the brain shouldn't set its own rules

Williams draws a clean line around what belongs outside the model itself: context and data (the agent's senses), memory (retained experience), tools (its hands), orchestration (its nervous system), and identity and permissions, which determine which doors it can open. That last category gets the most emphatic treatment in his answers. "I don't want the robot's brain deciding its own permissions. Every attempt to do this has been circumvented," he said. "The rules governing what the robot can do should live outside the brain."

That's a direct rebuttal to any architecture that asks a model to police its own behavior through instructions or system prompts alone. Williams's position is that permission boundaries need to be enforced by something outside the reasoning process entirely, not negotiated with it.

Controlling behavior versus controlling action

Here's the distinction Williams thinks the industry keeps missing: you may never be able to fully control what a model "thinks," but you can control what its body is capable of doing regardless. "We can determine whether it can access a database, execute code, send an email, deploy software, or transfer money. We can require approval for sensitive actions. And we can monitor what it actually does," he said. "That may ultimately be a much stronger security boundary than trying to perfectly control a probabilistic model."

The physical analogy he uses is blunt: "If I have a robot standing next to me, I'd rather physically prevent it from opening the vault than give it an instruction saying, 'Please never open the vault.'"

The unsolved problem hiding inside that answer

Williams isn't claiming this is a solved problem, and that honesty is worth preserving rather than smoothing over. Simple allow/deny permissions aren't sufficient on their own, he said, because "we don't know in advance exactly what capabilities a robot needs to accomplish the goal we assigned. If we over-restrict our robot, we lose exactly the creative capability that makes it so powerful." The real question, in his words, is "can we figure out how to keep the robot from doing things we don't want without sacrificing the ability to accomplish cool and interesting things? It's certainly a runtime control problem, but we haven't solved it yet."

Why the harness is also a new attack surface

The same infrastructure that makes an agent useful also makes it exploitable in new ways. "The harness is just software, and it creates an enormous new attack surface," Williams said. An attacker doesn't need to break the model at all; poisoning any input the model sees can be enough to manipulate what the agent does next; manipulating what it perceives, poisoning its memory, tricking it into misusing a tool, exploiting excessive permissions, stealing its credentials, or chaining several individually reasonable actions into something dangerous overall.

That reframes what the actual unit of security risk should be. "The important security unit becomes the agent using a capability against a particular resource in a particular context, not simply the prompt or the model output," he said. A single action can look completely reasonable in isolation and still be part of a dangerous chain.

Where this is all heading

Williams doesn't think the answer is picking one layer to defend. "We should continue making the brain safer. But we should assume that sometimes the brain will misunderstand something, hallucinate, get manipulated, or simply make a bad decision," he said. His framing for what that means in practice: limits built into the body and nervous system, monitoring of actual behavior, controlled access to dangerous capabilities, and real-time intervention mechanisms when something goes wrong.

The stakes of getting this wrong scale with what the agent is actually connected to. "A hallucination from a chatbot might produce a bad answer," Williams said. "A hallucination from an agent with production credentials can produce an outage." As agents pick up more autonomy and more access, that gap between "bad answer" and "real damage" is exactly the space the harness is supposed to close.

🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

The Security Conversation Your Clients Aren't Having About Agentic AI

Tom Smithverified - Jun 29

Cyera: Non-Human Identities Grew 480% in Six Months. Most Companies Have No Idea What They're Doing.

Tom Smithverified - Aug 3

Your Backup Data Knows More Than You Think. HYCU aiR Is Finally Asking It the Right Questions.

Tom Smithverified - May 14

Helping Clients Move from Pilot to Production: The Agentic AI Governance Playbook

Tom Smithverified - Jun 8

From Prompts to Goals: The Rise of Outcome-Driven Development

Tom Smithverified - Apr 11
chevron_left
18.4k Points794 Badges
249Posts
142Comments
105Connections
LLM Training & Evaluation Specialist with hands-on experience building major AI models. As one of th... Show more

Related Jobs

View all jobs →

Commenters (This Week)

6 comments
2 comments
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!