All articles
AI

Agentic AI for the Rest of Us: What's Actually Worth Doing (and What Will Quietly Wreck Your Budget)

Don't skip agentic AI. But, it means be specific about where you use it — and this week I wrote about the 3 places it actually pays off for a small business (not an enterprise), plus the 5 things that keep the bill predictable.

John SteinmetzJuly 31, 20265 min read
Agentic AI for the Rest of Us: What's Actually Worth Doing (and What Will Quietly Wreck Your Budget)

Agentic AI for the Rest of Us: What's Actually Worth Doing (and What Will Quietly Wreck Your Budget)

In May, OpenClaw creator Peter Steinberger posted his API dashboard: $1,305,088.81 spent in 30 days, 603 billion tokens, 7.6 million requests, across roughly 100 Codex coding agents running continuously. It wasn't a mistake — OpenAI, his employer, covers the bill, and he was deliberately running Codex's "Fast Mode" to see what software development looks like when token cost isn't a constraint. But he later noted that turning Fast Mode off would have cut the bill to roughly $300K. Same workload, same agents — a single configuration setting was worth 4x the cost.

That's the number worth sitting with, even though the story itself is an outlier. Nobody reading this is running 100 agents on someone else's budget. But the mechanism — one setting quietly multiplying your bill — is exactly what small businesses are about to run into with far less room for error. Uber is a more grounded version of the same lesson: the company blew through its entire 2026 AI budget in about four months and has since capped individual employees at $1,500 a month, per tool. If a company with a dedicated engineering org and a CTO watching the dashboard got surprised, an owner-operator running an agent off a personal API key has even less margin.

None of that means skip agentic AI. It means be specific about where you use it, and treat cost control as part of the setup, not an afterthought.

Where it actually pays off for a small business

Support ticket triage, in draft-and-approve mode. Full autonomy isn't the move yet for most SMBs — a human approving every response before it goes out gets you most of the value at a fraction of the risk. Vendors in this space report AI-assisted resolution costing roughly a tenth of a fully human-handled ticket, with well-tuned agents closing out 60–70% of first-contact queries end to end. Take those numbers as directional, not gospel — get your own baseline before you believe anyone's benchmark, including mine. Start with your three most repetitive ticket types, not your whole inbox.

Inbox and CRM hygiene. New lead comes in, their company size and role auto-populate, a follow-up sequence fires, and the second touch adapts based on what they actually clicked. This isn't a growth hack — it's a consistency fix. Most small businesses aren't losing deals to bad marketing; they're losing them to follow-up that a busy owner meant to send and didn't. This is the single cheapest, lowest-risk place to start.

Back-office document and inventory monitoring. Reconciling invoices against POs, flagging stock that's about to run out, drafting (not sending) reorder requests. Consultancies working with SMBs on this report double-digit hours back per week when the integration is done well — take the exact figure with a grain of salt, but the direction is right. The integration depth is the whole game here: an agent wired into your actual CRM, Stripe, and inventory system beats a standalone tool every time.

Notice what's missing: multi-agent orchestration, custom LLM fine-tuning, anything that needs a platform team. That's enterprise spend chasing enterprise problems. Skip it until you've outgrown the above.

How to not get a surprise bill

This is the part most guides skip, so here's what actually works:

  • Start in draft-and-approve mode, not full autonomy. Steinberger's 4x swing came from one setting on agents that were already running unsupervised at scale. One narrow workflow with a human checking output for the first few weeks catches both cost problems and quality problems before either compounds.
  • Set a hard daily spend cap before you turn anything on. Most platforms and API providers support this natively. While you're testing, cap it low — $20–50 a day is plenty to prove out a workflow. You can raise it once you trust the thing.
  • Watch for retry loops, not token prices. Runaway bills almost never come from the sticker price per token. They come from an agent stuck re-trying a failed step, or a workflow that quietly fans out into far more sub-tasks than you intended. That's a config and monitoring problem, not a model problem.
  • Route the boring steps to a cheaper model. You don't need your most expensive model reading and classifying every incoming email. Save the frontier-tier reasoning for the step that actually needs judgment.
  • Read the pricing model before you flip the switch. Several major platforms are moving from flat seat pricing to credit- or usage-based billing for agent features this year. That's a good moment to actually read what a "run" costs before your team starts running dozens of them a day out of habit.

Where to start this week

Pick one workflow — ticket triage or lead follow-up are the easiest wins — put a human in the approval loop, set a $20/day cap, and run it for two weeks before you touch the autonomy dial. You'll learn more from that than from any framework, including this one.

John Steinmetz logo
JOHN STEINMETZ
Fractional CDO

Fractional CDO services delivering data architecture, governance, and practical AI implementation that drives measurable enterprise value.

Navigate
Schedule a MeetingAbout
Connect
LinkedIn
Clients
Portal Login
© 2026 John Steinmetz / Fractional CDO
Message ConsentAdminArchitecting Intelligence