
If your AI voice agent's outbound calls are showing up as "Spam Likely," the carriers are not detecting the AI. They can't — the label is applied at ring time, before a single word is spoken. What they're scoring is the number: how fast it dials, how long its calls last, how many people hang up, whether it has a registered caller name, and how strongly it's authenticated. An AI agent burning through 400 leads from one unregistered number is producing the exact fingerprint those models were built to catch. Fix the number's registration and dialing pattern, and the label clears without touching the agent.
We run outbound voice agents for client acquisition and rehash campaigns, and this is the single most common reason a pilot "fails." The agent handles objections fine. The transcripts are good. The answer rate just quietly collapses from 22% to 6% over two weeks, and everyone blames the script.
What actually decides whether a call gets labeled
Three separate systems are involved, and agencies routinely confuse them.
STIR/SHAKEN attestation is cryptographic signing. Your originating carrier attaches a token saying how confident it is about the call's origin. A-level (full) attestation means the carrier knows you and confirms you have the right to use that specific number. B-level means it knows you but can't verify the number. C-level means it's just passing the call along. Attestation answers "is this caller who they claim to be" — it does not answer "does the recipient want this call."
CNAM is the caller name database. If you never set one, your number arrives as "Unknown" or bare digits. An unregistered number calling a mobile is a free spam signal you're handing over for nothing.
Analytics engines — Hiya, First Orion, and TNS — are what actually paint the label on the screen. They sit behind Verizon Call Filter, T-Mobile Scam Shield, and AT&T Call Protect, and they score every number continuously on behavioral data: dials per hour, call duration distribution, answer rate, block rate, user complaint reports, and whether the number appears in their caller registry.
The label comes from the third system. Most agencies only invest in the first one. That's why "we have A attestation, why are we still flagged" is such a common support ticket. A-attestation is an input to the score, not a veto over it. We've watched fully A-attested numbers get labeled inside 48 hours purely on volume.
Why AI voice agents trip these models harder than humans do
Nothing about a synthetic voice is visible to the scoring system pre-answer. What is visible is everything that makes an AI agent attractive in the first place.
Velocity. A human SDR makes 60 to 80 dials a day and takes breaks. An agent can do 400 before lunch, evenly spaced, with machine-perfect timing. Even pacing across hours is itself an anomaly — human dialing is bursty and irregular.
Short calls. Agents get hung up on more often in the first ten seconds, especially early in a deployment while the opener is still being tuned. A cluster of sub-15-second calls is one of the strongest dialer signals there is.
Connect delay. If there's a pause between answer and the agent's first word — a common symptom of a cold TTS stream or a webhook round-trip — recipients hang up, and the pattern looks like a predictive dialer. Cutting that delay is the highest-leverage fix in the whole list, because it improves both the label risk and the conversation.
Number reuse. Teams buy one number, point the agent at it, and scale the lead list instead of the pool. Reputation is per number. Scaling volume on a fixed pool concentrates all the risk in one place.
No local presence. A 212 number calling Dallas homeowners at 2pm reads as a call center. Area-code matching meaningfully improves answer rate, which then improves the reputation score — the two reinforce each other.
The setup that keeps numbers clean
This is the checklist we run before an outbound agent makes its first production dial. It takes about a day and saves weeks of remediation.
1. Confirm A-level attestation in writing. Ask your voice provider — Twilio, Telnyx, Vonage, or whoever sits under your Retell or Vapi account — what attestation your numbers sign with and whether the business entity is verified. If you're on a reseller layer, this is worth being blunt about: it's the difference between a fixable reputation problem and a structural one.
2. Set CNAM to your real public business name. It has to match the name on your website and your Google Business Profile. Mismatched or generic caller names ("SALES CALL," an LLC nobody recognizes) don't help and can hurt.
3. Register every number at FreeCallerRegistry.com. Free Caller Registry is the joint portal Hiya, First Orion, and TNS launched specifically so call originators could register once instead of three times. It's free, it takes about fifteen minutes, and skipping it is the most common omission we find in a broken outbound stack.
4. Size the number pool to the dial volume. Cap each number at 25 to 50 dials per day. For 500 dials a day, that's 10 to 20 registered numbers in rotation. Rotate on a schedule, not randomly, so each number has a stable, plausible pattern rather than sporadic spikes.
5. Match area codes to the lead's region. Provision local numbers for the markets you actually call into. This is a real cost — more numbers, more registration overhead — and it's the one that pays back fastest in answer rate.
6. Kill the connect delay. Measure the gap between answer and first audio. Under 500ms should be the target. Pre-warm the TTS stream, keep the first turn short, and don't put a webhook call in front of the greeting.
7. Suppress and honor opt-outs across every channel. A complaint is the heaviest single input to the score. If someone told your SMS agent to stop, the voice agent must never dial them. If you're running both, the suppression list has to be shared — this is the same discipline that keeps A2P 10DLC campaigns from getting rejected, and it belongs in one place, not two.
Remediating a number that's already flagged
Order matters here. Filing a remediation request before you've fixed the behavior gets you a temporary clear and a re-flag within days.
Start by confirming which carriers show the label — it's per-engine, so a number can be clean on AT&T and "Scam Likely" on T-Mobile. Keep two or three prepaid test handsets on different carriers and call them weekly. It's a $30 diagnostic that tells you exactly what your prospect sees, which no dashboard will.
Then rest the number. Stop dialing from it entirely for one to two weeks. Reputation models weight recent behavior heavily, so rest is what makes the remediation request credible. During that window, register it in the Free Caller Registry if you haven't, fix CNAM, and file remediation through each analytics engine's business portal. Expect days to a couple of weeks per engine.
Plan around the rest period rather than fighting it. If a client needs calls going out next week, rotate in fresh, properly registered numbers and treat the flagged ones as a slow parallel repair. Trying to rush a flagged number back into production is how agencies end up with a whole pool burned instead of one number.
What changes in 2026: branded calling
The longer-term answer is branded calling, where your business name, logo, and a call reason are delivered to the handset at ring time instead of a bare number. It's already available as a paid service through the analytics engines and several carriers, and it lifts answer rates substantially when the brand is one the recipient recognizes.
The regulatory side is moving too. In October 2025 the FCC issued a Further Notice of Proposed Rulemaking on call branding and caller ID authentication, looking at requirements for providers to transmit and display verified caller identity via Rich Call Data — name, logo, and call purpose carried inside the STIR/SHAKEN framework (FCC 25-76). Deployment through 2026 has been pilots and early production among Tier 1 carriers and enterprise platforms rather than anything universal.
Our read: branded calling is worth budgeting for if outbound volume is core to the business, and it is not a substitute for the fundamentals. Branding a number with a bad behavioral score doesn't clear the label — you're paying to put your logo on a call that still gets a spam warning next to it. Get registration and dialing right first, then brand.
The metric to actually watch
Answer rate per number, weekly. Not aggregate answer rate — per number.
Aggregate hides the failure, because a pool of twenty numbers with three burned ones still looks acceptable in a monthly rollup. Per-number answer rate shows you a specific number sliding from 20% to 8% two weeks before anyone would have noticed, which is enough time to rest and remediate it instead of losing it. Pair that with average call duration and you have an early-warning system that costs nothing to build.
Most teams find this out the hard way, after a campaign underperforms and someone finally borrows a colleague's phone and sees the label. Wire it into the call tracking you already run for attribution and it's one more column, not a new system.
None of this is exotic. It's registration hygiene and dialing discipline — the same category of unglamorous groundwork as getting voice agent compliance right before launch or sizing cost per minute honestly. It just tends to get skipped because the agent demo works perfectly on the founder's own phone, which is, of course, the one number every carrier already trusts.
Running an outbound AI voice agent and watching answer rates fall? We audit the whole stack — attestation, CNAM, registration, number pool sizing, connect latency, and the suppression logic underneath it. Get a free automation audit and we'll tell you which numbers are salvageable and which ones to retire.
