VOL. 02  ·  ISSUE 08  ·  FRIDAY, JUNE 26, 2026 BOSTON, MA  ·  RSS LIVE
Field Notes

Field Notes: SMBs Say They Trust Agents With High-Stakes Work. Watch What They Actually Ship.

62% of SMB leaders say they'd hand agents high-stakes work. The ones running agents in production hand over almost nothing.

Jonathan Tonthat · ML Engineer, Cellhub
4 min read

An HVAC shop owner told me last month he was “totally comfortable” letting his after-hours agent run the whole intake. Then he walked me through what it actually does: it books a slot, and the moment a caller says anything about a gas smell, it stops and pages a human. Comfortable in the survey. Scoped to almost nothing in production. That gap is the whole story.

The survey number is real. Upwork’s Q1 2026 research found 62% of SMB leaders are “very or extremely confident” handing high-stakes tasks to AI agents. Adoption is broad and the enthusiasm is genuine. But “confident” is a feeling about a category, and the deployed scope is a different number — the one nobody puts in a survey, because it doesn’t flatter anyone. I’ve yet to meet an owner whose agent is allowed to do the thing they told the survey they trust it with.

What’s Actually Working

1. Reversible jobs. The agents that stick draft, and a human sends. Quote estimates, appointment confirmations, first-pass email replies — anything where a wrong output costs a click to undo. The HVAC triage agent is the template: book the reversible thing, page a human for the thing that isn’t. The owner didn’t reason about it that cleanly; he just kept narrowing the agent’s job every time it embarrassed him until what was left was safe. That narrowing is the real deployment process, and no demo shows it.

2. Propose, don’t commit. A two-partner accounting firm I talked to runs an agent that reconciles transactions and proposes categorizations in QuickBooks — and posts nothing. The bookkeeper approves a screen of suggestions in the time it used to take to categorize three. Read-heavy, write-light. That’s the pattern that’s still running in week three, when the flashier deployments have already been switched off.

3. Bounded retrieval with a hard fallback. A small law office’s intake agent answers “do you handle this kind of case” from a fixed FAQ and, for anything off-script, says “I’ll have someone call you.” A constrained surface and an explicit handoff. It never improvises, which is exactly why they trust it.

What’s Still Broken

1. Multi-step decay. Agents that look flawless in the demo start hallucinating a few turns into a real conversation — a sales team found theirs invented objection responses after step four. SMBs discover this in week two, not in the sales call, and the discovery is usually a customer screenshot.

2. No ground truth, so no eval. The typical SMB has no labeled set of correct outcomes, so it can’t tell a 95%-accurate agent from a 70% one until something visible breaks. Measuring it would mean someone sitting down to grade a few hundred past interactions as right or wrong — which is exactly the work no ten-person business has spare hands for. So the agent runs unmeasured, and “very confident” turns out to mean “hasn’t blown up yet.” There’s no test suite behind the trust.

3. The confidence-evidence inversion. This isn’t only an SMB problem, and the cautionary tale this month came from the opposite end of the market. On June 13, KPMG pulled its own agentic-AI report after GPTZero found only 5 of its 45 citations cleanly matched real sources, and named organizations — UBS, the NHS, Swiss Federal Railways, Transport for London — disputed how the report described their AI use. If a Big Four firm selling governed enterprise AI can’t validate its own output, the survey confidence of a twelve-person shop is doing a lot of unearned work.

The Pattern

The agents that survive contact with a real business share three traits. They’re reversible — they propose, they don’t commit. They’re bounded — they own a named job, not “the business.” And they have a measurable handoff — a human owns the call that someone could get fired over. Every sticky deployment I’ve seen has all three. Every one that came off by week three was missing at least one.

So the 62% isn’t wrong, it’s just answering a different question. It measures how people feel about the words “AI agent.” Production measures what those agents are allowed to touch. Ask an SMB owner how much they trust their agent, then ask what it’s permitted to do without a human. The first answer is the press release. The second is the field note.

← Previous · №007
Mobility Watch: The iPhone Fold Is a $2,000 Confession About Apple's Component Yield
← All dispatches