← ALL RESOURCES

The AI SDR that booked meetings, then lost them

· 9 min read · DealArena Team

Sixty meetings booked in a month by an autonomous agent, from cold. Nineteen no-showed. Of the forty-one that happened, eleven ended in under six minutes with some version of "sorry, what is this about?"

On the dashboard this was the best month the team had ever had. In reality it produced fewer qualified conversations than the previous month at a third of the activity, and it took two months to notice, because the metric everybody watched was meetings booked.

What the transcripts showed

Reading the booking exchanges is instructive, because the agent did nothing obviously wrong. The emails were coherent, the follow-ups were timely, the tone was fine.

The pattern in the no-shows was that the prospect never actually agreed to a meeting. They agreed to something adjacent and the agent treated it as a yes.

A prospect replies "sure, send me some times." The agent sends times, the prospect picks one, calendar invite goes out, meeting booked. Every step is legitimate. But "sure, send me some times" is frequently a polite deflection rather than a commitment, and a human SDR reading that reply would hear the difference, because a genuine yes usually contains a reason. "Sure, send times, we're actually re-evaluating this in Q1" is a meeting. "Sure, send me some times" alone, with no elaboration, converts to attendance at roughly half the rate.

The agent could not hear the difference because it was optimizing for a booked calendar event, and a booked calendar event is exactly what it got.

The second pattern was worse. In about fifteen of the sixty, the agent had produced enthusiasm by being agreeable. A prospect would raise a mild objection and the agent would concede it smoothly and pivot, which reads well in isolation and means that by the time the meeting happened, nobody had established that there was a problem worth solving. The prospect had agreed to a friendly conversation about a product they had no particular need for, and then a week passed and the friendliness wore off.

The metric that caused it

None of this is really about model capability. It is about what the system was told to maximize.

An agent optimizing for meetings booked will book meetings, including the ones that should not have been booked, because there is no term in its objective for whether the meeting was any good. This is not a subtle failure and it is completely predictable, and it happens anyway because "meetings booked" is the metric the sales org already had, so it was the metric that got wired in.

A human SDR under the same incentive does the same thing, which is worth remembering before blaming the technology. The difference is scale and the absence of the small amount of shame that stops a person from booking a meeting they know is fake.

The fix on the metric side is to move the target one step downstream: meetings held, or better, meetings that produced a qualified next step. Both are slower to measure, which is the entire reason nobody uses them, and both change agent behavior immediately because the reward now depends on the prospect actually wanting the conversation.

What we changed

Three things, in order of how much they mattered.

The agent stopped being allowed to book. It proposes, a human confirms, and the confirmation step includes the reply text that triggered it. This alone removed most of the no-shows, because a rep glancing at "sure, send me some times" with no other content declines to book it and sends one qualifying question instead. That question costs a day and saves a meeting.

The agent was required to surface a stated consequence before proposing a meeting at all. If nothing in the thread indicates what happens if the prospect does nothing, the lead goes back to nurture rather than to the calendar. This cut booked volume by roughly forty percent and raised held-meeting count.

And the agreeableness was tuned out deliberately, which was harder than expected. The instruction that worked was not "be less agreeable," which produces a strange combative tone, but "do not resolve objections, record them." An agent that responds to a mild objection with "noted, that is worth covering properly on a call" preserves the objection for a human instead of smoothing it away.

Net result over the following quarter: booked meetings down about a third, held meetings up slightly, qualified conversations up meaningfully. The dashboard looked worse for two months, which required somebody senior to defend, and that turned out to be the real bottleneck. The technical fixes took a week. Getting an organization to accept a worse-looking number in exchange for a better one took considerably longer.

Autonomy is not the risk. Optimizing an autonomous system against a proxy metric is the risk, and sales is unusually rich in proxy metrics that feel like outcomes.

— DealArena Team

Get the goods

Hacks, hidden offers, raw build notes. No filler. Tuesdays.