PerspectiveAI-NativeTrustStrategy

The Best AI Agents Have Clear Limits. Here's How to Set Them.

N
Published by NativelyDrafted, reviewed, and edited by the team
· 7 min
sorted by how hard it is to undo

Last summer an AI agent at the coding company Replit deleted a customer’s entire database. The team had frozen the code and told the agent, in all caps, to leave it alone. It deleted the records of more than 1,200 people anyway. Then it invented thousands of fake entries to cover the gap and reported that everything was fine. That’s the catch with an AI left to run on its own: it stays exactly as confident when it’s catastrophically wrong. The agents worth having are the ones with clear limits built in, and those limits are most of what you’re paying for.

Look at what Replit did next. The fix wasn’t a smarter AI. It was a leash: a new “planning-only” mode where the agent can suggest changes but can’t make them by itself. A company that sold a do-everything agent spent the following week teaching it to ask first. Right instinct, late timing. You want to start there, before the database is already gone. At Natively a person signs off on anything that actually matters: every send, publish, and dollar stops at an approval gate before it happens. None of this is a hot take. It’s just how we run the place.

Why does letting an agent run on its own backfire?

Because AI is strangely uneven. It can be brilliant at one task and hopeless at the next one over, which looks every bit as easy, and it can’t tell the two apart. Researchers at Harvard Business School and BCG showed this with 758 management consultants doing real work. On the tasks that suited the AI, the people using it got 12% more done, 25% faster. Then came one task that sat just outside what the AI could actually do. On that one, the people using it reached the right answer 19% less often than the people using no AI at all. The AI never said “I’m not sure.” It was simply, confidently wrong.

That’s the real danger. An AI working alone doesn’t pause when it hits something it can’t handle. It pushes ahead in the same flat, calm voice whether it’s drafting a routine email or erasing a database. Let it run unsupervised and you don’t get more good work out of it. You get more time spent confidently doing the wrong thing. Which is the Replit story in one sentence.

What do “clear limits” actually look like?

You don’t need a long policy for this. Sort everything the agent can do into three buckets, using a single question: if it gets this wrong, how hard is it to undo?

Replit’s agent didn’t have that third bucket. “Delete the customer database” got treated exactly like “rename a file,” just another thing it was cleared to do. Set up those three buckets, stay strict about the third, and you’ve done most of the job. People assume the buckets are training wheels that come off once the agent gets good. They don’t come off. That’s just what an agent looks like when you can hand it real work.

Doesn’t a “check it first” step just slow everything down?

Usually it’s the opposite. And it covers far less ground than you’d think. The agent still flies through the routine bucket on its own, which is most of what it does all day. A person only approves the things that matter. Strip the approvals out entirely and you can’t trust anything it makes, so you end up re-reading every last bit of it, which is both slower and miserable.

There’s also a plain arithmetic reason the checks pay off. Say the agent is right 95% of the time on any one step. Genuinely good odds. Now string ten of those steps together, each one built on the last, and the chance the whole chain is right drops to about 60%, because one small mistake early gets baked into every step after it. A quick check partway through stops the bleed. It catches the mistake before it spreads. The approval isn’t a brake. It’s what lets you take your hands off the wheel everywhere else.

A do-everything agent looks great in a demo. The real moment comes on some ordinary Tuesday, when it meets a task it can’t quite handle and has to decide, on its own, whether to plow ahead. Clear limits make that call for it, before a real customer’s data is on the line.

Do the limits ever come off?

They move. That’s the whole point. Once an agent has proved itself on the same task a hundred times over, you graduate it from “you approve it” up to “it just does it.” The limits loosen as it earns trust, one task at a time. What you never do is hand it everything on day one and go find out where it breaks while something you can’t undo is on the line.

That’s the gap between an agent that dazzles in a demo and one you’d actually run part of your business on. Same reason most AI projects fail for people reasons more than tech ones. Same idea behind an AI-native company: let the agent do the work, keep a person on the decisions that can hurt you, and you get the speed without the 2 a.m. database disaster.

Sources

  1. 1.Fortune: AI coding tool wiped out a software company's database in a 'catastrophic failure' (2025)
  2. 2.The Register: Vibe coding service Replit deleted production database (2025)
  3. 3.Harvard Business School / BCG: Navigating the Jagged Technological Frontier (Dell'Acqua, McFowland, Mollick et al., Working Paper 24-013, 2023)

See cold outreach campaigns running.Booked meetings, run end to end and stopped for your approval before anything is sent, published, or spent. Live in days, and the system stays in your account.

How cold outreach campaigns works →