On the morning of June 10th, our AI SDR use case worked through a list of 47 accounts our team had been sitting on for three weeks. By noon it had researched every company, scored each one against our ICP, written and queued a first touch for 34 of them, and flagged 13 as not ready. Everything in draft, nothing sent. A human reviewed the queue, moved three to the back, adjusted one subject line, and approved the rest. By end of day, 34 first touches were out the door.
That was day one. Here’s the full 30 days.
9
qualified meetings booked, 30 days on Natively's own outbound pipeline
3,487
touchpoints across 312 sequenced accounts (420 ICP accounts pulled)
10.9%
account-level reply rate (34 replies from 312 accounts in sequence)
$174
all-in monthly cost: inference + orchestration + data APIs combined
Natively’s own production data, running our AI SDR use case on our outbound pipeline, June 10 – July 10, 2026. Not a study. Inference: Claude Sonnet 4.6 standard rate ($3.00/1M input, $15.00/1M output). Cost per meeting: $174 ÷ 9 = $19.33 all-in.
The hardest thing to fake in an AI pitch is a real number. Anyone can say their AI “dramatically improves” reply rates or cuts cost per meeting by a double-digit multiple. Nobody publishes the actual month-one data from running it on their own pipeline, against their own ICP, with their own results.
We do. We ran our own outbound sales use case on our own outbound motion for 30 days, start to finish: signal monitoring, account research, sequence writing, qualification. A human approved every message before it sent. This is what we got.
We started with 420 ICP accounts: funded B2B SaaS companies, 10–150 employees, actively hiring in sales or ops. Of those, 312 entered sequences in week one after initial research triage. The other 108 went in week two once the first batch was reviewed and the sending infrastructure had a few days of warm-up behind it.
A few things worth naming before reading these as benchmarks:
Email warm-up constrained week-1 volume. New sending infrastructure takes two to three weeks to build domain reputation, so we limited daily sends for the first two weeks. A company already operating on warmed sending infrastructure would have had higher week-1 touchpoints and probably more meetings in the first two weeks.
We also had three accounts where the buying signal that triggered the reach-out was misclassified: events that looked like growth signals but turned out to be internal reorgs. Those accounts got inappropriate first touches. All three were caught by the approval reviewer before they sent. That’s worth noting: the pre-send review step is not just overhead. It’s where misreads get caught before they reach an inbox.
Three of our 34 replies were automated responses from other companies’ AI outbound agents. We’re not sure if that’s a sign of where B2B outreach is headed, or just Tuesday.
Inference for the month, every research pull, every sequence draft, every qualification exchange, came to $76. That’s consistent with what we published in the per-utility cost post, which also reported approximately 3,500 touchpoints for $76 at Claude Sonnet 4.6 standard pricing.
Inference: Claude Sonnet 4.6 standard rate, $3.00/1M input, $15.00/1M output (Anthropic, July 2026). Data API costs are actuals from our Apollo account and email tooling for June 2026. SDR headcount: U.S. salary medians from BLS occupational employment statistics; 1.4× loaded is an industry rule of thumb for benefits + payroll taxes + tooling, our own estimate.
The obvious number, 9 meetings in month one vs. a ramped human SDR, isn’t the comparison that moves. Any new outbound motion, human or AI, runs below full rate in the first month: the sending domain is warming, the ICP definition is being refined, the signal taxonomy is being tuned. Month one is where you find out if the system works at all, not where you read its ceiling.
The comparison that matters is cost per meeting booked. Bridge Group’s SDR Metrics & Compensation Report (2023, 365 B2B companies) puts the average SDR quota at 19 meetings set per month, with 53% attainment, meaning roughly 10 actual qualified meetings per month from an average performer, at a fully loaded cost of $100K–$160K/year (RepVue US median: $90K OTE; Bridge Group median: $76K OTE; 1.4× loaded for benefits, payroll taxes, and tooling). Derived cost per meeting: $750–$1,400. Our first month: $19.33. Even at four times our current meeting volume, the gap doesn’t close. And the human SDR cost doesn’t include the data subscriptions, email tooling, and management overhead, all of which are already inside our $174.
Gartner’s 2024 research found that sellers who partner with AI are 3.7× more likely to meet quota, that’s about AI-augmented human reps, not AI running the motion end to end. Our model is structurally different: AI handles the outbound execution, a human handles the review and judgment calls, and the combination produces qualified pipeline at a cost floor no augmented rep team reaches.
The sending infrastructure is warm. The 14 “not now” replies re-enter the flow when their next signal fires. The 312 accounts that never replied get a final-touch message in week one before cycling out. We’re also adding a second ICP segment, SaaS-adjacent professional services firms, that we excluded from month one to keep the test clean.
If the meeting count doesn’t increase meaningfully in month two, the issue isn’t the AI. It’s the ICP definition or the signal taxonomy, and we’ll know which one because we can isolate it. The month-two numbers will be here in August.
Month one is never the ceiling. It’s the floor: the answer to whether the system works at all, whether the signals are real, whether the personalization lands, whether the approval step is fast enough to matter. This one did. Month two is where we find out what the ceiling is.
This is Natively’s outbound sales use case, and these are its numbers. The August post will have six more weeks of production data behind it. If you want to see how these numbers translate to your pipeline before then, that’s what the demo is for.
See signal-based outbound running.Replies from in-market buyers, run end to end and stopped for your approval before anything is sent, published, or spent. Live in days, and the system stays in your account.