---
title: "Do AI Setters Actually Work? What the Evidence Says"
slug: do-ai-setters-work
author: "Leonardo Maldonado"
category: "AI Setter Operations"
articleType: thought-leadership
tags: ["do AI appointment setters actually work","are AI setters worth it","AI setter results","AI setter evidence"]
publishedAt: 2026-08-20T09:00:00.000Z
updatedAt: 2026-08-20T09:00:00.000Z
canonical: https://setluca.com/blog/do-ai-setters-work
---# Do AI Setters Actually Work? What the Evidence Says

> Do AI setters work? Yes, for speed, follow-up, and qualifying at volume. The best independent evidence is a 5,179-agent trial published by the NBER in 2023. It measured a 14% lift in AI-assisted customer conversations, with almost all of it going to weaker performers. Nuanced objections and closing still need you.

## Key takeaways

- The strongest independent evidence is narrow but real. A 2023 NBER study of 5,179 support agents measured a 14% lift in issues resolved per hour, rising to 34% for the least experienced staff.
- Most published AI setter numbers are vendor-reported opinion, not measured outcome. Learn to tell the two apart before you buy.
- AI setters work well for speed, follow-up, and qualifying at volume. They measurably don't work for price negotiation, emotional objections, or thin inboxes.
- Four conditions predict success: inbound volume, offer clarity, your current response gap, and how much diagnosis your sale needs.
- The safest default is human review. With Luca, every AI reply waits for your approval unless you turn auto-send on.

---

Do AI setters work? For most coaches, yes, but not for everything, and the honest answer needs evidence rather than a vendor promise. An AI setter is good at the mechanical parts of DM sales: replying fast, asking the same qualifying questions, chasing leads who went quiet. Those are the parts humans skip when they get busy. Where they get oversold is the messy human stuff, like the hesitation behind a "let me think about it."

## Do AI appointment setters actually work?

Yes, at one specific job: getting more qualified leads onto your calendar, faster. Whether sales teams use AI at all stopped being an interesting question a while ago, since most of them do. Whether it pays for a coach selling in the DMs is the argument worth having, and the rest of this post is that argument. For the mechanics rather than the evidence, our [what is an AI setter](/blog/what-is-an-ai-dm-setter) primer covers what one does.

## What's actually been measured, and what's just vendor claim?

The strongest evidence for AI in sales conversations comes from outside the industry that sells it. In 2023, economists Erik Brynjolfsson, Danielle Li and Lindsey Raymond published *Generative AI at Work* through the National Bureau of Economic Research. It tracked 5,179 customer-support agents as an AI conversation assistant was rolled out to them in waves. Issues resolved per hour rose 14%.

Where that gain landed matters more than the headline. New agents improved 34%. The most experienced barely moved. The tool worked by spreading the habits of the best performers to everyone else, which means an AI setter lifts a weak DM operation toward competent. It won't push a strong closer past their ceiling.

Nobody has run a controlled trial of AI setters in coaching DMs. What we have instead is survey data and vendors' own usage numbers. Salesforce's 2026 report says 89% of sellers feel AI deepens customer understanding, which measures a feeling. The same report discloses its own agents contacting 130,000 leads to create 3,200 opportunities in four months, roughly 2.5%.

| What you're reading | What it is | Weight |
| --- | --- | --- |
| "5,179 agents, 14% more issues/hour" (NBER, 2023) | Controlled study, independent | High |
| "Purchases fell 23.7% to 4.8%" (Marketing Science, 2019) | Peer-reviewed field experiment | High, but calls, not DMs |
| "89% of sellers say AI deepens understanding" (Salesforce, 2026) | Opinion survey, 4,050 people | Low. Sentiment, not results |
| "130,000 leads, 3,200 opportunities" (Salesforce, 2026) | Vendor's own operating data | Medium. No control group |
| "Our clients book 3x more calls" (any vendor) | Selected customers, no baseline | None, until the denominator |

Ask any vendor, us included, one question: compared to what? A number without a baseline is decoration.

## Where do AI setters work well?

AI setters work best where the task repeats and a slow reply costs you money. Buyers care about speed far more than the businesses serving them think they do: 54% of consumers call a fast response critical, against 29% of businesses (Forrester for Google, 2020). That gap is your opening.

**Speed.** A lead who DMs you at 11pm doesn't wait politely. See our [DM response time and speed to lead](/blog/dm-response-time-speed-to-lead) breakdown.

**Follow-up.** This is where the folklore is worst. You'll see "80% of sales need five follow-ups" repeated everywhere with no traceable study behind it. The measured version comes from RAIN Group's Top Performance in Sales Prospecting research, a survey of 489 sellers. It takes an average of 8 touchpoints to land a first meeting, about 5 for top performers. Those top performers converted 52 of every 100 contacts against 19 for everyone else.

**Chart:** 44% give up after one attempt, 22% after two, 14% after three, 12% after four. Most quit before the fifth touch that most sales require. (Drop-off in follow-up effort by attempt number. Source: Cirrus Insight, B2B Sales Follow-Up Statistics (2025).)

**Qualifying at volume.** When dozens of DMs land in a day, asking the same three or four questions gets boring, and bored means sloppy. By Thursday you're skipping the budget question. An AI setter asks your [lead qualification questions](/blog/lead-qualification-questions-for-coaches) the same way every time, in the same order, at 2am.

## Where do AI setters measurably fail?

The honest limit: AI setters are weak exactly where sales gets human, and there's evidence for it rather than just a vibe. The clearest result is a 2019 *Marketing Science* field experiment by Xueming Luo and colleagues, covering more than 6,200 customers of a financial services firm.

Chatbots nobody had flagged as chatbots sold about as well as skilled human agents, and four times better than inexperienced ones. Then the researchers started telling customers they were talking to a bot before the conversation began. Purchase rates fell from 23.7% to 4.8%. Analysis of the calls showed the bots were doing the same job as well as before. Customers rated them as knowing less and caring less anyway.

So the skill is real and so is the cost in trust, and you can't dodge the second one by staying quiet. Meta's messaging rules, California's SB 1001 and Article 50 of the EU AI Act all say you have to tell people when a machine is answering. The drop was far smaller when that came later in the thread, which is the case for letting AI carry the early factual part and a human take the decision.

People simply give a human more benefit of the doubt. Asked who they'd trust for a recommendation, 54% picked a human agent and 32% picked AI (Gartner, 2026). You don't argue with that. You design around it.

A Canadian tribunal settled the liability question in 2024. In *Moffatt v. Air Canada*, the British Columbia Civil Resolution Tribunal made the airline pay damages over a refund policy its chatbot had invented, and threw out the argument that the bot was some separate legal entity. If your setter promises a payment plan you don't offer, that promise is yours. See our [AI setter vs human setter](/blog/ai-setter-vs-human-setter) breakdown and [objection handling in DMs](/blog/objection-handling-in-dms).

## Which four conditions decide if AI setters work for your business?

Do AI setters work for you specifically? It depends on four things you can check in ten minutes, and none of them is the tool. The NBER lift came from raising the floor, so your gain depends on how far your floor sits below your ceiling.

**Inbound volume.** Below roughly 5 conversations a day, the gain is convenience. Between 10 and 40, consistency breaks down and an AI setter pays. Above 40, you're already dropping leads.

**Offer clarity.** If you can say your price range, how you deliver, and who it's for in two sentences, an AI setter can qualify people against it. If your offer changes with every client, there's nothing steady to qualify against.

**Your response gap.** Measure your median time to first reply for a week. Under ten minutes, the speed benefit is small and you're buying follow-up consistency. Over two hours, that's your lift.

**How much diagnosis the sale needs.** A $200 program sells on fit and logistics. A $15k container often sells on a conversation nobody can script. More diagnosis, earlier handoff.

| Your situation | What to do | Why |
| --- | --- | --- |
| 10-40 DMs/day, clear offer, replies over 2 hours | Run an AI setter, review on | Your floor is the delay |
| Under 5 DMs/day, high-ticket, heavy diagnosis | Fix lead volume first | Nothing to scale yet |
| 40+ DMs/day, calendar already full | AI for triage, a human for calls | Past what review alone absorbs |
| Offer changes per client, no price band | Fix the offer first | Nothing fixed to qualify against |
| Replies under 10 minutes, no follow-up | Run it for cadence, not speed | Your gap is the 8 touches, not the first |

If you're weighing this against a person, our [hire a setter vs AI](/blog/hire-a-setter-vs-ai) comparison lays out where each earns its keep.

## What separates a good setup from a bad one?

Past those four conditions, results come down to setup rather than the logo on the tool. In the NBER study the gain came from spreading what the good agents already did, and that only helps if the system knows what good looks like in your voice. Which is why the safest default is simple: no AI reply goes out without you seeing it first. With Luca, auto-send is off by default.

| Bad setup | Good setup |
| --- | --- |
| Generic, off-the-shelf voice | Trained on how you actually write |
| Auto-sends everything | Review queue, you approve or edit |
| Tries to close and handle objections | Qualifies, warms up, hands off to you |
| No visibility into what it sent | You see every draft and reply |

## A worked example: Priya, mindset coach

**Illustrative composite, not a customer result. Swap the numbers for your own.**

Priya runs a mindset practice with a $3,200 twelve-week container. She gets around 55 inbound DMs a day from Reels, and spent a week measuring instead of guessing. Median time to first reply: 6 hours. DMs she never answered: about 18 a day. Calls booked that week: 4.

She turned on an AI setter for speed and follow-up, review queue on. Her qualifying questions were the three she already asked: what they're working on, what they've tried, whether they're ready to start within a month. Her handoff rule was blunt. Anything touching price, anything emotional.

A typical thread ran like this. Lead: *"hey saw your reel about morning anxiety, do you work with people 1:1?"* Draft reply, approved in one tap: *"I do. That reel came out of something I worked through myself. What does the morning usually look like for you right now?"* Two exchanges later the lead asked the price, and the thread was flagged for Priya instead of answered.

After 30 days: median first reply under two minutes, zero unanswered DMs, 9 calls booked in the final week against her baseline of 4. She edited roughly 90% of drafts in the first fortnight and about half by the end, and still closed every call herself.

The honest caveat: her gain came from a 6-hour reply gap and 18 dropped conversations a day. A coach replying in ten minutes with a working follow-up sequence would see a fraction of it. Our [AI appointment setter cost](/blog/how-much-does-an-ai-appointment-setter-cost) breakdown covers the money.

## Edge cases most coaches don't plan for

**The 2am reply from another timezone.** Sending at 3am their time reads as a bot. Reply inside their waking window, not yours, and cap send windows per timezone.

**The person who already bought.** An existing client asking a logistics question should never enter a qualification flow. Tag customers before you launch, or your setter will pitch a program to someone already paying.

**The lead in genuine distress.** Coaching inboxes get messages about grief, burnout, and worse. No qualification, no booking link, no cadence. Route it to you and answer as a person.

**The WhatsApp 24-hour window closing.** Once a lead's last message is more than 24 hours old, WhatsApp will only let you send pre-approved template messages. A follow-up sequence built for Instagram just quietly stops working there. Check the rules on each channel before you run one sequence across both.

**The lead who reschedules twice.** That's a qualification signal, not an admin task. Have the setter surface it instead of quietly re-booking.

## Troubleshooting: when the results don't show up

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| Replies fast, calls flat | It's answering questions instead of moving to a booking step | Add a booking ask after the second qualifying answer; measure calls, not replies |
| Leads go quiet after the first AI reply | Voice mismatch; it reads nothing like your Reels | Retrain on 30 of your own sent DMs, not website copy. See [make AI DMs sound like you](/blog/make-ai-dms-sound-like-you) |
| Plenty of calls, terrible show-up rate | Qualifying is too loose, so unready people book | Add a timing question and a budget-range question before the link |
| Still editing every draft after a month | Rules are fine, examples are thin | Feed your edits back as training examples |
| Follow-ups feel pushy, you get blocked | Cadence too tight or too uniform | Spread touches across days, stop at 5. See [DM follow-up sequence](/blog/dm-follow-up-sequence) |

## Which mistakes make AI setters look like they don't work?

**Turning auto-send on in week one.** You haven't learned what it gets wrong yet, and week one is when it gets the most wrong. Read a fortnight of drafts before you hand over the send button.

**Measuring replies instead of calls.** Reply volume always goes up. That's the easy metric and it proves nothing. Booked, held calls is the number.

**Training it on your polished writing.** Your sales page voice is not your DM voice. Feed it your sent messages, typos and shorthand included.

**Letting it handle price.** The *Marketing Science* disclosure result is a warning about this exact moment. Price is where perceived competence gets tested hardest.

**Testing during a launch.** Launch traffic is warmer, so any tool looks brilliant. Test in an ordinary month against a baseline you measured, because most coaches guess their reply time and guess low.

## How do you run a 30-day test with defined success metrics?

Run the test as an experiment with a baseline, not a trial you eyeball. Most people buy a tool because someone told them they were behind, then judge it four weeks later on a feeling. Write the numbers down first and the answer stops being a matter of opinion.

**Week 0, baseline.** Do nothing differently. Record four numbers daily: inbound conversations, median time to first reply, conversations you never answered, calls booked.

**Weeks 1-4, run it.** Review queue on. Same offer, same content schedule, no launch. Set the thresholds below before you have feelings about the tool.

| Metric | Pass threshold at day 30 | Benchmark |
| --- | --- | --- |
| Median time to first reply | Under 5 minutes | 54% call fast response critical (Forrester for Google, 2020) |
| Unanswered conversations per day | Zero | -- |
| Touches per lead before booking | 5-8 sustained | 8 average, 5 for top performers (RAIN Group) |
| Calls booked per week | Above week 0, holding at week 4 | -- |
| Show-up rate | No worse than week 0 | A fall means loose qualifying |
| Edit rate on drafts | Lower in week 4 than week 1 | The 34% novice gain (NBER, 2023) came from this curve |

If calls booked rise and show-up rate holds, it worked. If replies rose and calls didn't, your qualifying or your booking ask is the problem, not the AI. If nothing moved, check whether your baseline was already good, because that's the most common answer and it isn't a failure.

So, are AI setters worth it? If you're losing leads to slow replies and dead follow-up, and you'll keep a human on the close, yes. If you expect it to replace your sales skill, no. See our [AI setter for coaches](/ai-setter) guide.


## FAQ

### Do AI setters actually book calls?

Yes. A well-run AI setter qualifies the lead, handles the back-and-forth of finding a time, and drops the booked call on your calendar. RAIN Group's research found it takes an average of 8 touchpoints to land a first meeting, which is exactly the persistence most coaches skip. The close still belongs to you.

### Is there real evidence that AI setters work, or is it all vendor marketing?

Both exist. The strongest independent evidence is a 2023 NBER study of 5,179 customer-support agents, which measured a 14% rise in issues resolved per hour. No controlled trial has been run on AI setters in coaching DMs specifically, so treat any vendor multiple without a stated baseline as marketing.

### Do AI setters sound like a bot?

They can, if you don't train them. A setter running a generic script reads like one. A 2019 Marketing Science field experiment found undisclosed bots sold about as well as proficient human agents, so the capability is there. Voice training and a review queue are what keep it from sounding robotic.

### Will leads buy less if they know an AI replied?

Some will, and you still have to disclose. That same 2019 study found purchase rates fell from 23.7% to 4.8% when bot identity was revealed before the conversation. The penalty shrank sharply when disclosure came later, which is the argument for AI handling the early thread and you handling the decision.

### Are AI setters worth it for a small coaching business?

If you're losing leads to slow replies or nonexistent follow-up, usually yes, even at low volume. Below about five inbound conversations a day the gain is mostly convenience. The NBER data suggests the lift goes to whoever has the biggest gap between their worst day and their best. Check Luca's pricing.

### Can an AI setter handle objections?

Simple, factual ones, yes, like "how long are your calls" or "what's included." Emotional or nuanced objections, no. Gartner's 2026 survey found 54% of customers trust a human more than AI for recommendations, against 32% who trust AI more. A good setup flags those moments and hands them over.

### Will an AI setter replace a human closer?

No. Gartner predicts that by 2027, half the companies that attributed headcount cuts to AI will rehire for similar work under different titles. Only 20% had actually cut service headcount when surveyed in 2026. An AI setter replaces the delay and the dropped follow-up, not the conversation that sells.

### Who's responsible if my AI setter promises something I don't offer?

You are. In February 2024, the British Columbia Civil Resolution Tribunal held Air Canada liable for a refund policy its chatbot invented, rejecting the argument that the bot was a separate entity. Keep pricing, guarantees, and refund terms out of automated replies, or approve them yourself.


## Sources

1. [Erik Brynjolfsson, Danielle Li & Lindsey R. Raymond -- "Generative AI at Work," NBER Working Paper 31161 (2023)](https://www.nber.org/papers/w31161)
2. [Xueming Luo, Siliang Tong, Zheng Fang & Zhe Qu -- "Frontiers: Machines vs. Humans: The Impact of Artificial Intelligence Chatbot Disclosure on Customer Purchases," Marketing Science 38(6) (2019)](https://pubsonline.informs.org/doi/10.1287/mksc.2019.1192)
3. [Gartner -- "Gartner Predicts Half of Companies That Cut Customer Service Staff Due to AI Will Rehire by 2027" (February 2026)](https://www.gartner.com/en/newsroom/press-releases/2026-02-03-gartner-predicts-half-of-companies-that-cut-customer-service-staff-due-to-ai-will-rehire-by-2027)
4. [Gartner -- "85% of Service and Support Leaders are Expanding Human Agent Responsibilities Despite Expectations of Mass AI Layoffs" (April 2026)](https://www.gartner.com/en/newsroom/press-releases/2026-04-28-gartner-survey-finds-eighty-five-percent-of-service-and-support-leaders-are-expanding-human-agent-responsibilities-despite-expectations-of-mass-ai-layoffs)
5. [Salesforce -- "State of Sales" (2026)](https://www.salesforce.com/news/stories/state-of-sales-report-announcement-2026/)
6. [Forrester Consulting for Google -- "What Businesses Need To Know About Communicating With Consumers" (2020)](https://developers.google.com/business-communications/business-messages/files/google-what-businesses-need-to-know-about-communicating-with-consumers.pdf)
7. [RAIN Group Center for Sales Research -- "How Many Touchpoints Does It Take to Make a Sale?"](https://www.rainsalestraining.com/blog/how-many-touchpoints-does-it-take-to-make-a-sale)
8. [American Bar Association -- "BC Tribunal Confirms Companies Remain Liable for Information Provided by AI Chatbot" (February 2024)](https://www.americanbar.org/groups/business_law/resources/business-law-today/2024-february/bc-tribunal-confirms-companies-remain-liable-information-provided-ai-chatbot/)
9. [Cirrus Insight -- "B2B Sales Follow-Up Statistics: Touches, Timing & Reply Data"](https://www.cirrusinsight.com/blog/sales-follow-up-statistics)

---

Published by SetLuca, the company behind Luca.