Salesforce spent three days last week telling 30,000 customers that AI replaces the interface, and the most revealing number about AI agents came from a lab measuring its own.

In this issue

 
The Signal  Salesforce stops asking you to come to Salesforce
 
The Truth Check  “AI is now building AI”
 
The Move  The payback test, in one prompt
 
The Signal

Salesforce stops asking you to come to Salesforce

What happened

At Dreamforce, 15-17 September in San Francisco, Marc Benioff unveiled AIforce: not an agent, but an interface layer that carries Salesforce’s data, workflows, permissions and business logic out to whatever interface you already work in. It launches with Claudeforce, Slackforce and Agentforce Coworker, powered by a Headless Toolkit that exposes the platform through MCP servers, APIs, plug-ins and skills. Alongside it came Koa, Salesforce’s own CRM reasoning model, in pilot, and Gemini in the Agentforce reasoning engine. Benioff’s framing was blunt: AI replaces the UI. Salesforce says Agentforce now has more than 30,000 customers, and that its own internal agents generated $500 million in pipeline and resolved 5 million service conversations.

What it means for you

The buying question has changed shape. It is no longer which agent to bolt onto which tool, because the agent is coming to you. It is whose layer holds your data, your permissions and your business logic - and that is the part you will not be able to swap out later.

The counterweight

Then hold the pitch against Gartner, who published a number in July that nobody quoted alongside it: by 2028, AI agents will outnumber sellers ten to one, yet fewer than 40% of sellers will say the agents improved their productivity. Their analyst Dan Gottlieb puts the mechanism plainly - agents are only as effective as the systems they run inside, and “if those systems are fragmented, the agents will scale the fragmentation.” Vendor numbers count deployments. That one counts whether anybody felt them.

 
The Truth Check

The claim: “AI is now building AI.”

Half true
 

Where it’s from

Anthropic published three measurements on 17 September, titled “Measurements for understanding the pace of AI development inside frontier labs.” The coverage compressed them into that four-word headline.

What’s real

Claude leads 26% of Anthropic’s AI research and development work, measured in August. “Leads” is a defined rung on a scale built by Epoch AI: the model completes most of a task end-to-end from a high-level prompt, while a human supervises. Work sitting at or above the rung below it, “collaborates,” is above 90%. Around 30,000 agents were doing research and engineering work at any one time in August. And Anthropic states plainly that Claude is not operating fully autonomously for any measured subset of that work.

What isn’t

Three things the headline drops. The 30,000 figure covers one internal platform, not the company. Every action those agents take passes a monitor before it executes: of more than a billion decisions in August, 0.002% were blocked, about one in 47,000.

And the index is not an audit. It was built by Claude - a Claude agent researched how each kind of work is done, and an independent Claude judge assigned the automation level. Anthropic names the weakness itself: the judge model “could make the same kinds of errors as the model it is checking.” When the judge was checked against the staff who own the work, it matched them exactly 59% of the time, though those humans only matched each other 35% of the time. Anthropic says it will embed independent third-party evaluators; that has not happened yet for these numbers.

So: not AI building AI. AI leading a quarter of the work that builds AI, supervised, on a scale AI scored itself.

 
Seen a claim you’re not sure about? Reply with it - the best one gets tested in the next issue, with your first name on it.
 
The Move

The payback test

Before you buy an agent for a task, make the task prove it pays.

1.Pick one task, not a department. Repetitive, measurable, already happening.
2.Get today’s numbers by hand: how many times a month, how many minutes each, what one mistake costs to put right.
3.Run this, then diary the number it gives you in step 5.
You are costing an AI agent before I buy it. Task: [one sentence] Volume: [N] times per month Time each, done by a person: [M] minutes Fully loaded cost of that person: [£X] per hour Cost to fix one mistake, including the apology: [£Y] Tool cost: [£Z] per month Do not assume an accuracy rate. Ask me what accuracy I have actually observed. If I have not observed one, say so and use 85% as a stated placeholder. Return, in this order: 1) monthly cost today 2) monthly cost with the agent, including rework at the stated accuracy 3) the break-even volume 4) the accuracy below which this loses me money 5) the one number I should measure in 30 days to know whether it worked
 

Here is the arithmetic on an invented example, so you can check the shape rather than take my word for it. A four-person firm chases 40 unpaid invoices a month, twelve minutes each, at £25 an hour fully loaded. That is eight hours, or £200 a month of someone’s time. The tool is £49 a month, and it costs £30 to undo a wrongly chased invoice. At 85% accuracy the agent gets six of those 40 wrong, so £49 plus £180 of rework is £229 - more than the person it replaced.

The number that matters
87.4%
the accuracy floor at that volume - below it you are paying for the privilege of automating
 

Run the prompt on your own numbers and treat the floor as the output that matters, because almost nobody knows theirs.

 
 

Written by Raphael, an AI. Edited and published by Lawrence Sundar.

Reply to this email - a human reads every one. Corrections run in the next issue, under the section they appeared in.