BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Comparisons/The Things an AI Sales Agent Will Not Do for You
ComparisonsAugust 12, 2026
Read · 5 min
ai sales agent · ai sdr

The Things an AI Sales Agent Will Not Do for You

Worked funnel arithmetic on automated outreach, the Gmail spam threshold that ends a domain, and which sales tasks automation genuinely improves.

Key takeaways
  • Generated outreach at volume is a commons problem. Every additional sender lowers the response rate for everyone, including the sender.
  • Work the funnel arithmetic before buying anything. Halving the reply rate halves the meetings, and reply rates are moving in that direction.
  • Gmail requires bulk senders to keep spam complaints below 0.30 percent. On a list of 2,000 that is about six people pressing report, which is not many.
  • The durable value sits in research, enrichment and keeping records straight. First touch generation at volume is where the damage happens.
  • Domain reputation outlasts the campaign. A bad send costs you the channel for months, long after the list is forgotten.

The pitch is meetings on autopilot. It is also, if you check your own inbox, the thing you delete without reading. Both of those are true at once, and the gap between them is the whole subject.

An AI sales agent is sold as a headcount replacement and priced as one. Whether it is worth anything depends on which part of selling you point it at, and the honest answer splits cleanly rather than sitting on a fence.

What does the funnel actually produce?

Build the model before you build the list. Take a plausible set of numbers for cold outreach and follow them through, and the result is smaller than the pitch implies even when every assumption is generous.

Start with 2,000 contacts. Assume 85 percent of messages reach an inbox rather than being blocked or filtered, which gives 1,700 delivered. Assume 3 percent of those delivered produce a reply of any kind, which is 51 replies. Assume a quarter of replies are positive rather than a polite no or an unsubscribe, which is about 13. Assume 60 percent of positive replies become a booked meeting, which is 8.

Eight meetings from two thousand contacts. That is the good case, and it is a real result if your product is worth a meeting.

Now halve the reply rate to 1.5 percent, which is the direction cold reply rates have been moving as volume rises. Everything else held constant, you get roughly 4 meetings. The relationship is linear and unforgiving: reply rate is the only lever in that chain that has collapsed by half in living memory, and it is the one that scales directly into the output.

Note

Those percentages are a model, not a measurement. Substitute your own if you have them. The point is the structure, which is that the number at the end is a product of five fractions, and generation only helps with one of them.

Why does sending more make it worse?

Because attention is fixed and the cost of producing a message fell to almost nothing. When drafting a personalised email took ten minutes, volume was self limiting. It no longer is, so the constraint moved from the sender's time to the recipient's patience, and the recipient's patience is shared across every sender.

This is the shape of a commons problem. Each additional sender makes a rational individual decision and the aggregate effect is that the channel degrades for everybody. You cannot opt out of the degradation by being the good sender, and you also cannot fix it by sending more.

What you can do is notice which side of the arithmetic you are on. If your outreach is indistinguishable from the volume, you are paying the cost and receiving the average. If it is distinguishable, the collapse in everyone else's reply rate is not entirely bad news for you.

Where does automation actually belong?

Sort the job into tasks, and the answer stops being controversial.

Set diagram of six sales tasks: list building, research, first touch, follow up, scheduling and CRM hygiene
TaskDoes automation help?WhyWhat it costs when it goes wrong
List building and enrichmentYes, stronglyStructured lookup against public records, easy to verifyWasted sends on stale contacts
Research before contactYes, stronglyReading and condensing is the task models are best atA wrong fact in the first line, which reads worse than no fact
First touch at volumeNoThe output is what recipients have learned to deleteSpam complaints, domain reputation, the channel itself
Follow up sequencingPartlyTiming and reminders are mechanicalChasing someone who already said no, which ends the relationship
Meeting schedulingYesA closed problem with a definite answerDouble bookings, which are embarrassing and recoverable
Record keeping and CRM hygieneYes, stronglyTedious, high volume, checkableAlmost nothing, which is why nobody sells it

Notice what the two strongest rows have in common. Research and record keeping are both about handling information you already have a right to, and both produce output a person reviews before it reaches anybody. They are also the two nobody markets, because answers your questions about your own pipeline sells worse than books meetings while you sleep.

The weakest row is the one the category is named after. That is not an accident of current quality. Generated first touch fails because its output is competing against every other generated first touch for the same three seconds, and the model has no information the other models lack.

What does a bad send actually cost?

More than the campaign, and the mechanism is worth understanding because it is measured against a published threshold rather than left to judgement.

Google's email sender guidelines set the requirements for reaching personal Gmail accounts. All senders need SPF or DKIM authentication, valid forward and reverse DNS records, TLS in transit, and spam rates reported in Postmaster Tools kept below 0.3 percent. Senders above 5,000 messages a day to Gmail need SPF and DKIM and DMARC, alignment between the From domain and the authenticated domain, and one click unsubscribe on marketing messages.

Statistic card showing six spam reports out of two thousand recipients reaches the published Gmail spam rate threshold

Do the arithmetic on that threshold, because it is smaller than it sounds. Out of 2,000 recipients, 0.30 percent is six people. Six recipients pressing the report button, out of a list where you expected fifty replies, puts you at the line. This is the actual constraint on volume outreach, and it is not a matter of taste.

The consequence is also not confined to the campaign. Reputation attaches to the sending domain, which is usually the domain your invoices and your customer support replies come from, and inbound support is a queue with its own rules about which tickets can safely be answered without a person. A campaign that goes badly in March degrades your ability to reach existing customers in May, and there is no undo. Sending cold outreach from a separate domain is the standard mitigation and it is worth doing on the first campaign rather than the third.

Where the reply rate goes when volume rises

It is worth being precise about why the reply rate is the fragile term, because it is the one the vendor's own model treats as fixed.

A recipient's decision is comparative. They are not judging your message against a blank inbox, they are judging it against the fourteen others that arrived the same morning, and the threshold they apply moves with the volume. When the volume doubles, the bar for reading past the first line rises, because the cost of being wrong about which messages to open has gone up.

That is why the usual advice to write better copy has limited reach. Better copy competes on the margin within a mechanism whose baseline is falling. It helps, and it does not restore the number, and any projection that assumes today's reply rate holds for a twelve month campaign is projecting the wrong variable.

The counterargument is worth stating fairly. If reply rates fall for everyone and you are one of the few sending something recognisably human, your relative position improves even as the average declines. That is real, and it is an argument for fewer, better messages rather than for more of them. It is not an argument that volume tools stop working, since they still produce a number, it is an argument that the number they produce is falling while their price is not.

What the tool cannot know

One limit sits underneath all the others. An outreach agent has access to what is public about a company and nothing about whether that company is in a position to buy.

Buying happens when something changed: a person moved into a role, a contract came up for renewal, a system broke, a rule changed. Almost none of that is visible from a website, and the signals that are visible, a hiring post or a funding announcement, are visible to everyone at the same moment, which is precisely why a founder who has just raised receives two hundred messages that week.

So a generated message arrives with no information about timing, which is the variable that decides the outcome. That is not a quality problem to be fixed by a better model. It is missing data, and the only reliable sources of it are a conversation, a referral, or somebody arriving on your site with a problem. Which is the argument for spending the same budget on being findable rather than on being loud.

The compliance floor

Nobody selling you an outreach tool will lead with this, and it is short enough to read.

In the United States, the FTC's CAN-SPAM compliance guide is direct about scope: the law covers all commercial messages, not just bulk, and makes no exception for business to business email. The requirements are accurate header information, a subject line that reflects the content, clear identification that the message is an advertisement, a valid physical postal address, a clear opt out mechanism, and honouring opt outs within ten business days. The guide states penalties of up to $53,088 per email in violation.

Europe works from the other direction. Consent is one of the lawful bases, and Article 7 of the GDPR sets conditions on it: where processing is based on consent, the controller must be able to demonstrate that the person consented, the request must be clearly distinguishable and in plain language, withdrawal must be as easy as giving consent, and the person must be told they can withdraw before they give it. A purchased list does not satisfy any of that, and being able to demonstrate consent is a records requirement rather than an assertion.

The practical reading for a small seller is that the American regime regulates the message and the European regime regulates the list. A tool that generates compliant looking messages solves half of one of them.

What good AI assisted outreach looks like

The difference is where the model sits in the process. Used to write the message, it produces the average. Used to prepare the person writing it, it produces something the recipient has not seen forty times this week.

Before. Hi Sam, I noticed you are the operations lead at Fairhill Supplies. Many businesses in your industry struggle with inventory visibility. Would you be open to a quick chat this week about how we help companies like yours streamline operations?

After. Sam, your site lists next day delivery on the ceramics range but the checkout quotes three to five days on anything over 4kg. If that gap is deliberate, ignore me. If it is a carrier rule you inherited, we fixed the same thing for a pottery supplier last year and I can send you what they did in two paragraphs. No meeting needed.

The second one took four minutes, three of which were the model reading the site and reporting what it found. Nothing in it was generated as prose. That is the split worth internalising: the model does the reading, you do the writing, and the reason is that the reading scales and the writing is the part that has to be yours.

It also does not scale to two thousand people, which is the honest cost. Four minutes each across two hundred contacts is thirteen hours. Against eight meetings from two thousand generated sends, that is a trade many small businesses should take, and it is one nobody models because the tool is priced per send.

Questions people ask

Does personalisation still work if everyone does it?

Inserted variables stopped working some time ago, and that is what most tools mean by personalisation. A first name and a company name signal a merge field, not attention. What still works is a specific observation the sender could only have made by looking, because that is expensive to fake at volume and recipients can tell.

Should I warm up a new domain?

Yes, and treat it as a requirement rather than a growth hack. A new domain sending two thousand messages on day one behaves exactly like a spam operation from a filter's point of view, because that is what a spam operation looks like. Build volume gradually and keep the authentication records right from the start.

What about AI agents that answer inbound instead?

Different economics entirely, and usually better. Someone who contacted you has already spent attention, so a fast accurate answer is worth something to them rather than being an imposition. The failure modes are also cheaper: a wrong answer to an inbound question gets corrected, while a wrong claim in cold outreach gets you reported.

Can I use it to write follow ups only?

That is the better use of the two, with one condition. Follow ups have context the first message lacked, since the recipient has now done something or nothing, and drafting against a real signal is a task the model is good at. The condition is that the sequence must stop when somebody says no, and that has to be a rule in the system rather than a habit in a person.

What to do with the budget instead

If the money was going to buy sending volume, the alternatives are unglamorous and better documented.

Your existing customer list outperforms any cold list, because the people on it have already bought. The stage where models genuinely move the number there is segmentation rather than copy, which we worked through in which stage of an email programme AI actually improves. If the constraint is the writing rather than the list, the honest question is how much editing a generated draft needs before it ships, and our measurement of that editing distance is the number to plan against.

The wider version of this decision, which task to hand over and which to keep, is set out in where a founder's hour really goes. And if you are weighing the cost of assisted work against a per seat sales tool, our plans and credit pricing makes the comparison concrete rather than notional.

The part that will not change

Nothing in the arithmetic above depends on the current generation of models. Reply rate is a function of how many messages a person receives, message quality only helps relative to the alternatives in the same inbox, and a domain reputation takes months to build and one campaign to lose.

So the durable position is the one that gets less attention: use the model where its output is checkable and its input is information you already have a right to. Research, enrichment, follow up drafting, records. Keep the first sentence to a customer, because that sentence is the only part of the whole chain where being different from everyone else is worth anything.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building