Skip to main content

The Great AI Arbitrage: How the Internet Got It Wrong, and How the Economy Is Actually Changing

Cheap intelligence is not a business model. Turning it into a result someone can trust might be.

Conceptual illustration of scattered online chatter and documents becoming an orderly workflow, with a human controlling the final approval gate.

From cheap generation to trusted delivery: the work between the two is where the business takes shape. AI-generated editorial illustration.

On the internet, AI arbitrage can sound like three businesses wearing the same trench coat.

One trades cryptocurrency. Another sells AI-assisted services. A third sells the dream of running the second without doing much work at all. In the sales-pitch version, you buy access to a model, find a client who has not discovered it yet, and keep the difference.

The missing slide is the one about what happens when the output is wrong.

In service businesses, AI arbitrage is a loose term for exploiting the gap between what AI-assisted work costs to produce and what the market still pays for it. The opportunity opens when production costs fall faster than prices and working practices adjust. Sustainable margins depend on quality, review, and the full cost of delivery.

It is usually not arbitrage in the strict financial sense: capturing a price discrepancy in the same or equivalent asset. Selling an AI-assisted service involves ordinary business risk. A low software bill does not make the profit automatic.

Still, beneath the dubious promise lies an important economic question. What happens when parts of knowledge work become much cheaper to produce, while dependable outcomes remain difficult to deliver?

The answer has less to do with finding a clever prompt than with what happens after the prompt returns an answer.

The gap between an answer and a service

The appeal is understandable. A tool can draft a report in minutes that would previously have taken someone hours. The apparent opportunity is to sell the report at the old price and produce it at the new cost.

That can be a real efficiency gain. It is not, by itself, a durable business.

A client may discover the tool. A competitor may lower the price. The draft may require extensive correction. Or the customer may never have wanted a report in the first place; they wanted a decision they could defend.

The pitch treats those as details. In a services business, they are the business.

Part of the confusion is that “AI arbitrage” gets applied to three different activities:

Three meanings of AI arbitrage
Use of the termWhat is being sold or pursued?What actually determines success?
AI-assisted tradingPrice differences between markets, sometimes including crypto exchangesExecution, fees, liquidity, and market or platform risk
An AI arbitrage agencyServices delivered partly with AI, such as research, document processing, or content productionFinding customers and delivering useful work reliably at a sustainable cost
AI-enabled workflow redesignA repeatable business outcome produced through a different mix of software and human workEnd-to-end quality, integration, economics, and accountability

The last two overlap. A good agency may build excellent workflows. An internal operations team may capture the same efficiency without selling anything externally. Neither needs to pretend it has discovered a financial loophole.

“Digital arbitrage” is another broad online-business label, used for opportunities involving digital platforms, distribution, and differences in cost or price. Adding AI describes a production method; it does not remove the need for a viable customer proposition.

Nor is every course or AI agency a scam. The appropriate target of skepticism is the claim, not the job title. In September 2024, the US Federal Trade Commission announced enforcement actions involving, among other businesses, online-storefront schemes it alleged had used deceptive AI-powered earnings claims. That is a concrete reason to question promises of effortless income, not evidence that all AI businesses are fraudulent. FTC announcement.

The useful question is not, “Can I charge more than the model costs?” Almost any plausible service can pass that test.

It is, “Can I repeatedly deliver something worth buying after paying for everything required to get it right?”

AI changes more than the cost of labor

The search for cheaper service work long predates the chatbot. For decades, one way to reduce its cost was to move it. A company could have similar tasks performed in a market with lower labor costs, while accounting for training, management, and coordination. This is labor arbitrage, also written as labour arbitrage.

In its simplest form, the work stayed much the same. Its location and wage bill changed.

AI introduces another possibility: change how much human execution the work requires.

Consider a hypothetical distributor processing supplier invoices. In the existing workflow, someone opens an attachment, copies the details into a system, finds the purchase order, checks the amounts, and routes discrepancies for investigation.

Offshoring that process changes who performs those steps and at what cost. Redesigning it with AI could automate the reading of varied documents, use ordinary software to match records and check totals, and send the exceptions to a person.

The contrast is not “expensive human versus cheap machine.” It is one arrangement of work versus another.

The familiar labor-arbitrage question is: Where can we buy these hours more cheaply?

The workflow question is: Which hours do we still need, and what must happen before we can remove the others safely?

Call the second approach capability arbitrage: using a newly affordable capability before an organization—or its competitors—has fully adapted its processes and prices. This is an explanatory lens, not a settled economic classification.

There are two possible gaps to exploit. One is temporary: knowing how to use a tool while a customer does not. The other is harder to close: being able to integrate that tool into a process that works under real operating conditions.

Neither advantage is guaranteed to last. But the second asks more of a competitor than buying the same subscription.

Software already gave businesses ways to grow without adding people in exact proportion. AI did not invent operating leverage. It can extend it to parts of language-heavy work that were harder to automate with rigid rules alone.

The uneven productivity gains

A study published in the Quarterly Journal of Economics examined AI assistance across 5,172 customer-support agents at one company. Workers with access to the assistant resolved 15% more issues per hour on average. Less-skilled and less-experienced workers saw gains of about 30%.

These were people working with assistance, not an autonomous system replacing the support department. Most agents in the sample worked outside the United States, mainly in the Philippines. The finding illustrates how AI and offshore service delivery can coexist rather than cancel each other out. The authors also caution that their study does not establish economy-wide employment or wage effects. Generative AI at Work.

The boundary matters just as much as the gain. In an experiment involving 758 consultants, participants using GPT-4 completed more tasks and took 25.1% less time on work within the model's capabilities. On a deliberately chosen task outside those capabilities, however, AI users were 19 percentage points less likely to reach the correct answer. The research appeared in its final published form in 2026; it was not a test of 2026 models. Navigating the Jagged Technological Frontier.

A tool can improve one part of a job while making another worse. Even the people using it may misjudge the effect. In an early-2025 randomized study of 16 experienced open-source developers doing 246 tasks in familiar repositories, METR found that AI access made completion take 19% longer, despite participants believing they were faster. That small, specialized study used tools available then—not today's models. METR's 2025 study.

METR's February 2026 follow-up suggested improvement, but selection effects and measurement difficulties prevented a reliable speedup estimate. These studies support a practical conclusion: measure the work you actually operate rather than borrowing someone else's productivity claim. METR's follow-up.

The larger economic change is therefore not a universal collapse in the need for people. It is a change in the amount and kind of human work needed for particular outcomes. A team might handle more demand without equivalent hiring, improve service with the same staff, or reduce staffing. Productivity evidence alone does not tell us which choice it will make.

Who keeps the savings?

A lower delivery cost can lead in three directions.

First, the provider may retain it as a higher margin, particularly while competitors struggle to match the service. This is the version promoted in the AI arbitrage pitch.

Second, competition may transfer some of the saving to customers through lower prices or better service for the same fee. A provider can become more productive without becoming more profitable.

Third, lower costs can expand what customers buy. A business that could afford one research report each quarter might commission a useful update every week. A service can gain a new market when its price falls within reach of smaller customers.

These possibilities can coexist. Which dominates depends on demand, competition, and how much customers value the difference between providers. They also explain why the employment result is uncertain: reducing the work required per order and increasing the number of orders pull in different directions.

The temporary arbitrage can disappear even while the technology succeeds. Once faster, cheaper delivery becomes the normal expectation, the business has to compete on the new terms. Yesterday's price is not an entitlement.

The workflow is the product

What, then, can an AI business offer once its competitors have access to the same tools? The simplest version puts a prompt behind a payment screen.

There is nothing inherently wrong with that. A well-designed interface can save people time. But if a customer's entire benefit is getting an answer they could obtain just as easily elsewhere, the business needs a reason they will stay.

The more interesting value sits in the orchestration layer: the software, rules, data connections, and human handoffs that turn a model's output into completed work.

For the distributor, the distinction is concrete. A chatbot can read an invoice. The business needs a correct record, matched to the right transaction, with discrepancies resolved and approvals preserved.

In our hypothetical workflow, that requires six connected stages:

  1. Receive. Accept documents through an approved channel, preserve originals, and enforce account access controls.
  2. Extract. Identify invoice fields and retain their source locations, using a document model where needed.
  3. Check. Use conventional software to verify arithmetic and compare purchase orders, receipt records, and previous invoices. The ledger is no place for improvisation.
  4. Route. Send missing records, conflicting amounts, possible duplicates, and changed payment details to a qualified reviewer.
  5. Approve. Present the evidence for a decision. Keep payment authorization separate from document preparation.
  6. Record and monitor. Preserve the result and its history. Track corrections and sample apparently clean cases as well as flagged exceptions.
Illustrative invoice workflow: receive documents, extract fields with AI, check records with fixed rules, route exceptions to human review, approve the record, and preserve its audit history. Payment authorization remains separate.

View this infographic at full size

The main processing path and the exception-review branch form one workflow. Preparing an accepted record does not authorize a payment.

The product is the completed, traceable record. The extracted text is an intermediate step.

The same distinction runs through the enterprise AI stack: a model supplies a capability; the surrounding operation must deliver the service.

None of this requires every step to be handled by an autonomous agent. An AI agent can choose tools and next steps dynamically; a predefined workflow follows a more constrained sequence. Anthropic's engineering guidance distinguishes the two and recommends starting with the simplest approach that meets the need, adding complexity when it demonstrably helps. Building effective agents.

For the invoice example, fixed rules should handle much of the routing and validation. An agent might investigate an unusual mismatch across approved records. Its findings still need to be checked against those records: another model agreeing with it is not independent verification.

“Workflow-first” is the useful starting point. Define the job, the evidence of success, and the limits of authority. Then choose the technology.

What a finished result actually costs

Token prices are easy to quote. Delivery costs are harder to calculate.

A useful operating metric is:

Cost per accepted outcome = total workflow cost, including failed attempts and rework ÷ outcomes that meet the agreed quality standard.

Define that standard before the pilot. Otherwise, it is too easy to claim success by counting generated outputs instead of completed work.

Here is a deliberately simplified planning example—not a measured case study or a market price estimate. Assume the distributor completes 1,000 invoice records per month to the same acceptance standard under either process. The manual baseline uses 100 hours at a loaded labor cost of $30 an hour, including routine checks and handling.

Illustrative monthly processing costs
Monthly delivery costManual processAI-assisted workflow
Human processing, review, and corrections$3,000$1,200
Model and document-processing usage, including retries$0$200
Additional workflow tooling and monitoring$0$300
Setup cost allocated across 12 months$0$500
Total within this comparison$3,000$2,200
Cost per accepted record$3.00$2.20
Hypothetical monthly costs for 1,000 accepted invoice records: manual processing costs $3,000; AI-assisted processing totals $2,200, made up of $1,200 human work, $200 model usage, $300 tooling and $500 setup allocation. The difference is $800, about 27 percent.

View this infographic at full size

The small model bill is only one part of delivery. Both bars assume 1,000 accepted records at the same quality standard; these are illustrative figures, not measured results.

The setup allocation is $6,000 spread over 12 months. Unchanged shared costs are excluded from both columns, so the zeroes do not imply a software-free manual operation. Consequential losses from undetected errors are not modeled; a real assessment must consider them separately.

Under these assumptions, the saving is $800 a month, or roughly 27%—not the 93% suggested by comparing $3,000 of labor with $200 of model usage alone.

If human review and correction rise from $1,200 to $2,000, the advantage disappears. If quality falls, the comparison is invalid even before the arithmetic changes.

Released capacity also differs from cash saved. With the same payroll, the immediate benefit may be time for other work. An agency would still need to cover sales, account management, and overhead before earning a profit.

Even when the numbers work, the design alone offers little protection from competition. Another team can copy an orchestration diagram. More defensible advantages can accumulate in permissioned customer context, reliable integrations, domain-specific tests, trusted relationships, and experience with exceptions.

Two invoice services might use the same model, yet only one fits a customer's approval procedures and handles its recurring edge cases without repeated intervention. That difference can be worth paying for.

Outcome-based pricing needs equal care. Charging for a “resolved ticket” creates a poor incentive if the system can close it while the customer still needs help. Buyers must be able to verify what they are buying.

A robust workflow is less like a fortress than a factory. It earns its value through acceptable work, repeatedly, including on the bad days.

Someone still has to take responsibility

Now suppose a familiar supplier submits an invoice with new bank details. The amounts match. The formatting looks ordinary. The extraction is flawless.

Should the business pay it?

That question exposes the limit of treating intelligence as execution speed. The system needs a policy for changed payment details, an independent verification route, and someone authorized to hold the transaction. Reading the invoice correctly does not establish that the request should be trusted.

The responsibility does not disappear when the reading and sorting are automated. It may sit with a domain specialist, an operations manager, a product team, or several people working together. What matters is that someone has the evidence and authority to make the decision.

In the distributor, a finance lead might own the acceptance rules and payment controls. An operations specialist could handle the unresolved queue, with authority to stop processing when something looks wrong. An engineering owner would investigate recurring failures and test changes before release.

The machine may remove the copying and sorting, while leaving those responsibilities intact. The organizational change is to make ownership explicit, rather than assuming the person who performed each manual step would notice the problem.

That also creates a new capacity question. If one expert must review every difficult case, the exception queue can become the next bottleneck. Staffing, training, and escalation routes have to grow with it. Giving a person responsibility for more automated work does not give them unlimited attention.

“A human is in the loop” can describe a meaningful safeguard or a person clicking through a queue they cannot realistically check.

For oversight to matter, reviewers need access to the underlying evidence, enough time to examine it, relevant expertise, and the power to reject or pause the system. Evaluation also has to include failures the automation did not flag.

NIST's generative-AI risk guidance treats governance, testing, monitoring, and incident response as ongoing responsibilities. They are not a one-time approval before launch. NIST Generative AI Profile.

For an operator, that means watching more than throughput. Track cost per accepted outcome alongside corrections, serious errors, unresolved exceptions, and what customers experience. More completed tasks are not an improvement if the system quietly shifts the cleanup to someone else.

There is a personal discipline here, too. As explored in how AI can shape the story we tell about ourselves, confident reassurance can be seductive. An operator needs evidence that the system works, not affirmation that they are a visionary for building it.

The work of learning to judge

It would be comforting to say machines will handle the boring work and everyone will graduate into judgment. There is no basis for promising that.

Some tasks may disappear. Some jobs may shrink or change. Access to training, managerial choices, bargaining power, and demand for services will affect who benefits.

The ILO–NASK's 2025 occupational-exposure research estimates that one in four workers worldwide is in an occupation with some exposure to generative AI. It sees transformation as more likely overall than complete replacement because many jobs still require human contributions. Exposure, however, is not a count of jobs already lost or a prediction that one-quarter will disappear. ILO–NASK research.

There is also an apprenticeship problem. If a company automates the routine tasks through which beginners learn, how will those beginners acquire the judgment needed to supervise the system later?

One practical response is to preserve deliberate learning: compare model work with source material, investigate exceptions, explain corrections, and practice without assistance. Senior judgment should not be treated as a resource that renews itself.

For someone deciding where to build skills, the useful starting point is a real domain and a real process. Learn what good work looks like, where it fails, and how to test it. Tool fluency helps, but it cannot substitute for knowing why an apparently convincing answer is wrong.

Seen from that perspective, the internet's promise of buying intelligence cheaply and selling it dearly misses much of what makes the work valuable.

The more consequential question is whether you can reorganize work around that cheaper capability without losing quality, trust, or control.

That is a harder question. It is also a better foundation for a business.

Start with one repeatable workflow. Establish the baseline. Specify what an acceptable result is. Test the smallest useful intervention, count the review and rework, and expand only when the evidence warrants it.

The emerging advantage is not simply working faster than another person. It is understanding a process well enough to decide what machines should do, what people must still do, and how the two become a service worth trusting.

Anyone can generate the invoice summary. Someone still has to decide whether to pay it.

Frequently asked questions

Is AI arbitrage a scam?

Not inherently. The term covers legitimate AI-assisted services as well as exaggerated business pitches. Evaluate the specific offer: who is the customer, what result is delivered, how is quality checked, and what costs are excluded from the earnings claim? Guaranteed-income promises warrant particular skepticism.

What is an AI arbitrage agency?

An AI arbitrage agency sells services it delivers partly through AI tools. Examples include document processing, research support, and content operations. Its commercial value comes from solving a customer problem reliably—not simply from having access to a model. The agency remains responsible for its agreed deliverables.

Can you make money with AI arbitrage?

Potentially, if customers value the service and revenue exceeds all delivery and business costs. AI can lower parts of the production cost, but it does not guarantee demand or remove sales, review, integration, and support expenses. Test a narrow service before assuming the economics will scale.

How is AI arbitrage different from labor arbitrage?

Labor arbitrage seeks cost advantages from differences in labor markets. AI-enabled workflow redesign changes the mix of human and machine work needed for an outcome. The two can coexist: an offshore team can use AI, and an AI-assisted process can still depend heavily on people.

Is AI arbitrage the same as crypto arbitrage?

No. Crypto arbitrage concerns price differences between markets. AI may be used as a tool in trading, but AI-assisted service delivery is a different business model. Neither the shared label nor the use of AI establishes that a particular offer is profitable or safe.

The weekly question

From obsession to clarity — one original question every week.

We answer one noisy topic at a time, in full. No daily roundup, no thread bait — just the question, the principles, and the system.

Continue reading

More in Evolving AI