Command Palette
Type to search. Press Enter for the top match.
runpoint.
Runpoint Research · July 2026

The evidence-based case for AI in the mid-market.

We reviewed the academic studies, government data, enterprise surveys, and implementation reports from 2017 through 2026. This is the short version: who is getting a return, which projects work, what changed this year, and what a mid-market CEO should do next.

UngatedEvidence gradedSources linkedBuilt for $20M to $200M companies
More work like this

Field Notes turns the week's AI noise into one useful call.

Every two weeks we send mid-market executives one idea, one number, and one thing worth doing on Monday. You can read this entire report without subscribing.

The Runpoint letter

Three useful minutes every two weeks. Written from active work with companies like yours. Read past issues →

Executive Summary

The headline: AI capability has advanced faster than organizational value capture. Controlled studies repeatedly find double-digit task gains, yet representative firm data still finds little realized productivity impact at most companies. The missing layer is not usually the model; it is workflow redesign, data and integration work, adoption, governance, and converting released time into cash or throughput.

Who gets the biggest ROI: not a revenue band by itself, but companies with high-volume digital work, expensive labor/SaaS/outsourcing baselines, verifiable outputs, usable data, a powerful process owner, and the ability to redesign roles. Information/software, professional services, and finance lead adoption; customer operations, IT/cyber, software engineering, and document-heavy back offices produce the most repeatable evidence.

The $20M–$200M thesis is plausible but not yet causally proven. Mid-market firms can combine enough process volume to matter with shorter decision chains and CEO-led change. But large firms still adopt and scale more often, and there is no rigorous study showing that this revenue band earns the highest ROI. Runpoint should position the segment as a mechanism-backed deployment sweet spot, not a settled empirical fact.

What changed in 2026: code generation crossed a meaningful threshold. Built-from-scratch internal tools, integrations, tests, CRUD applications, and bounded code changes are now technically feasible with agentic systems and expert oversight. But end-to-end SaaS replacement is only conditionally solved: identity, permissions, migration, auditability, uptime, maintenance, and organizational ownership remain the hard part. A precise formulation is: code production is increasingly solved; software stewardship is not.

Best strategic position: “rent the models, own the workflows” is well aligned with the evidence. Model prices, packaging, and leaders change quickly; owned data, process logic, evaluations, integrations, and source code preserve switching power and grow in value over time. Sovereignty is an economic architecture, not a requirement to train or host the frontier model.

AI use, % of orgs

88Source: McKinsey – The State of AI (late 2025)

Organizations reporting AI use in at least one function; global survey, late 2025.

2017 20Source: McKinsey – The State of AI (late 2025)
Organizations reporting AI use in at least one function; global survey, late 2025.Source: McKinsey – The State of AI (late 2025)

Global survey results on AI use, scaling, EBIT impact, company-size differences, workflow redesign, and high performers.

Any EBIT impact, %

39Source: McKinsey – The State of AI (late 2025)

Organizations reporting any enterprise-level EBIT impact attributable to AI.

Organizations reporting any enterprise-level EBIT impact attributable to AI.Source: McKinsey – The State of AI (late 2025)

Global survey results on AI use, scaling, EBIT impact, company-size differences, workflow redesign, and high performers.

AI high performers, %

6Source: McKinsey – The State of AI (late 2025)

Organizations meeting McKinsey's late-2025 high-performer definition.

Organizations meeting McKinsey's late-2025 high-performer definition.Source: McKinsey – The State of AI (late 2025)

Global survey results on AI use, scaling, EBIT impact, company-size differences, workflow redesign, and high performers.

No realized impact, %

80Source: NBER – Firm Data on AI (2026)

Representative international firms reporting no productivity or employment impact over the prior three years.

Representative international firms reporting no productivity or employment impact over the prior three years.Source: NBER – Firm Data on AI (2026)

Representative international survey of almost 6,000 CEOs, CFOs, and executives in the United States, United Kingdom, Germany, and Australia.

How AI ROI Progressed

2017–2021: narrow prediction and point solutions

AI programs were centralized, data-science-heavy, and aimed at forecasting, ranking, optimization, fraud, or recommendations. ROI was normally modeled as a case-level cost reduction, risk reduction, or incremental revenue lift. McKinsey adoption moved from 20% in 2017 to 56% in 2021, but the beneficiary was usually a workflow or decision–not the whole enterprise.

2022–2024: individual augmentation

Generative AI made language and coding broadly accessible. Laboratory and field studies showed meaningful gains in writing, customer support, consulting, and bounded coding. Enterprises often measured prompts, seats, hours saved, or pilot satisfaction. Those are leading indicators, not financial returns; unreallocated time and unaccepted output have a cash value of zero.

2024–2025: production integration and the pilot trap

Adoption surged, but scaling lagged. The most consistent lessons were to redesign workflows, focus on a small number of valuable use cases, build reusable data/governance foundations, and make senior leaders accountable. The common “95% of pilots fail” headline is directionally consistent with the value-capture gap but comes from a non-peer-reviewed MIT NANDA report with an opaque denominator; it should not anchor an investment case.

2026: agents, software production, and quality-adjusted economics

Frontier systems can now execute longer coding and computer-use tasks, while open-weight models drive inference prices down. The economic unit shifts from cost per token or seat to cost per accepted business outcome. The new constraints are orchestration, evaluation, review, system access, and who owns the evolving workflow.

Organizational AI adoption, 2017–2025
Organizational AI adoption, 2017–2025 data
YearOrganizations using AIDefinition noteContext
2.02K20%AI use in at least one functionMcKinsey global survey
2.02K47%Question/definition differs from later wavesMcKinsey global survey
2.02K58%Question/definition differs from later wavesMcKinsey global survey
2.02K50%Question/definition differs from later wavesMcKinsey global survey
2.02K56%AI use in at least one functionMcKinsey global survey
2.02K50%AI use in at least one functionMcKinsey global survey
2.02K55%AI use in at least one functionMcKinsey global survey
2.02K72%AI use in at least one functionMcKinsey global survey
2025a78%Early-2025 survey waveMcKinsey global survey
2025b88%Late-2025 survey waveMcKinsey global survey

Why Large Task Gains Have Not Become Large Company Gains

The studies are not contradictory; they measure different layers. A worker can complete a bounded task 25% faster while the company sees no P&L change because demand is fixed, review work grows, downstream bottlenecks remain, or management never removes cost or increases throughput.

The evidence supports a four-stage funnel: capability → accepted output → changed workflow → financial realization. Most public studies prove stage one or two. McKinsey, BCG, PwC, and NBER show that stages three and four remain scarce. In McKinsey's late-2025 survey, 88% reported AI use, roughly one-third had begun scaling, 39% reported any enterprise EBIT impact, and only about 6% met its high-performer definition. NBER's representative 2026 international survey found more than 80% of firms reported no productivity or employment impact over the prior three years, even as they expected future gains.

The practical implication is severe: time saved is inventory, not ROI. It becomes economic value only when it produces more accepted output, avoids hiring, reduces overtime or outsourcing, eliminates a license, shortens revenue cycle time, or enables role redesign.

Observed task-level AI effects in controlled and field studies
Observed task-level AI effects in controlled and field studies data
Study and taskReported task effectEndpointSampleInterpretation
GitHub Copilot – bounded coding55.8%Faster task completion95 professional developersNarrow JavaScript HTTP-server task; vendor-associated RCT
Noy & Zhang – professional writing40%Less time444 college-educated professionalsQuality also improved 18%; bounded writing tasks
BCG/HBS – consulting inside frontier25.1%Faster task completion758 BCG consultantsParticipants also completed 12.2% more tasks with >40% higher quality
NBER/QJE – customer support14%Issues resolved per hour5,172 agentsAverage gain; roughly 34% for novice/lower-skilled workers
METR – experienced OSS developers-19%Longer task time (shown as negative)16 developers, 246 issuesEarly-2025 tools, large familiar repositories; METR says estimate is now stale

The Company Archetype Most Likely to Win

The best predictor of ROI is a system of conditions, not company size. High-return candidates have: (1) frequent, repeatable digital transactions; (2) a costly current baseline; (3) observable quality and an accepted-output definition; (4) accessible systems and reasonably clean data; (5) a CEO or functional owner who can change the process; (6) a human checkpoint where error costs require it; and (7) enough demand to monetize released capacity.

Mid-market companies can be unusually attractive because they may have enterprise-like process volume without enterprise-length decision chains. Yet they also face thinner data/IT capacity. RSM's 2025 middle-market survey found 92% experienced implementation challenges, 62% said deployment was harder than expected, and data quality and expertise were leading problems. McKinsey found only 29% of companies below $100M had reached scaling, versus nearly half of companies above $5B. Those facts challenge any simple “smaller is better” claim.

Mid-market sweet-spot thesis: evidence and counterevidence

Source: Government and middle-market survey synthesis

Synthesis of representative firm-size evidence and self-reported middle-market surveys; associations are not treated as causal ROI.

Mid-market sweet-spot thesis: evidence and counterevidence
FactorImplication for $20M–$200MEvidenceConfidence
Decision speed and senior ownershipA CEO or functional owner can change roles, incentives, and process quickly.BCG 2026: 72% of surveyed CEOs said they were the main AI decision maker; McKinsey high performers show far stronger senior commitment.Medium – mechanism supported; samples skew large.
Enough volume for meaningful economics$20M–$200M firms often have material service, software, document, and back-office throughput.Census shows adoption rises with firm size; 32% of 100–249 employee firms used AI in 2026.Medium – volume/adoption, not ROI.
Large-firm scale advantageBigger firms can fund platforms, specialist teams, and change programs; mid-market nimbleness does not guarantee scale.McKinsey late 2025: 29% of sub-$100M companies had begun scaling versus nearly half of companies above $5B.High – direct survey counterevidence.
Lower organizational change distanceFewer layers can make workflow redesign and realization faster.Consistent with change-management theory and Runpoint experience, but no direct causal size-band comparison was found.Low – attractive hypothesis.
Observed performance correlationAI-adopting middle-market firms are growing faster, but readiness and selection likely explain part of the gap.NCMM 2026: 87% of adopters vs 66% of non-adopters grew revenue; average growth 12.9% vs 5.8%. Authors flag adopters as larger/faster-growing/readier.Low for causality; useful for segmentation.
Thin data and IT capacityIntegration, security, data quality, and maintenance can overwhelm an otherwise attractive use case.RSM 2025: 92% reported implementation challenges; 41% cited data quality and 39% lack of expertise.High – repeated self-report evidence.

Initiative Portfolio: Where ROI Is Most Repeatable

Prioritize workflows, not generic tools. The highest-confidence opportunities combine deterministic process control with AI for the parts that benefit from language, perception, classification, generation, or exception handling. Start with a baseline and an owner, then instrument acceptance, quality, cycle time, and downstream impact.

The ranking below favors repeatability and cashability for a $20M–$200M company, not novelty. Individual copilots are easy to deploy but often weak on cash realization. Owned workflow systems and selective SaaS replacement require more implementation work but can create durable, auditable economics when they remove a meaningful license, service, or labor baseline.

Initiatives ranked for repeatable mid-market ROI

Source: Cross-study initiative evidence synthesis

Evidence-graded ranking based on controlled task studies, enterprise operating-pattern surveys, and implementation feasibility for mid-market firms.

Initiatives ranked for repeatable mid-market ROI
RankInitiativeWhy ROI worksEssential control2026 status
1Source: Cross-study initiative evidence synthesisCustomer-service and employee assistHigh volume, clear handle-time/quality metrics, strong field evidence; biggest gains often accrue to newer workers.Knowledge grounding, acceptance/quality audit, escalation path.Solved / proven
2Source: Cross-study initiative evidence synthesisDeterministic workflow automation with AI stepsCombines reliable state and permissions with AI for extraction, classification, drafting, or exceptions; directly attacks cycle time and manual cost.Structured inputs/outputs, confidence thresholds, human approval for consequential actions.Solved / proven pattern
3Source: Cross-study initiative evidence synthesisSoftware engineering and internal toolsCode, tests, documentation, integrations, and built-from-scratch tools can remove queues and SaaS spend; capability improved materially in 2026.Repository harness, automated tests, review, observability, clear ownership.Solved for bounded work
4Source: Cross-study initiative evidence synthesisDocument and knowledge operationsContracts, claims, compliance packets, proposals, and research have observable inputs and outputs and often expensive review baselines.Rubrics, citations, exception routing, sampled human QA.Solved / conditional by risk
5Source: Cross-study initiative evidence synthesisOwned SaaS replacement, including core systemsCan eliminate recurring license and implementation cost while encoding company-specific workflow and preserving ownership. Core ERP or CRM replacement can work when the migration and three-year economics are explicit.Three-year TCO, migration plan, identity/permissions, support and maintenance owner.Conditional, increasingly attractive
6Source: Cross-study initiative evidence synthesisIndividual copilots for writing/search/analysisFast adoption and strong bounded-task evidence; good for scarce expert capacity.Usage telemetry and a realization plan; prevent workslop and review externalities.Technically solved; cash ROI conditional
7Source: Cross-study initiative evidence synthesisRevenue personalization and decision supportLarge upside where experiments and contribution margin can be measured.Holdouts, causal measurement, bias/privacy controls.Conditional / data dependent
8Source: Cross-study initiative evidence synthesisEnd-to-end autonomous multi-system agentsPotentially large labor and cycle-time leverage.Sandboxing, permissions, rollback, full audit, human checkpoints, reliability budget.Emerging / generally unsolved

Code Generation and SaaS Replacement: The 2026 Inflection

Runpoint's conviction is directionally right and should be sharpened. Frontier coding agents can now explore repositories, plan changes, write code and tests, execute tools, and iterate. METR's task-horizon work shows the duration of software tasks frontier agents can complete at a given reliability has been rising rapidly. OpenAI and Anthropic report internal agent-authored code at unprecedented scale.

But the neutral evidence is still mixed. GitHub's bounded Copilot experiment found a 55.8% speed gain; METR's early-2025 study of experienced open-source developers found AI made work 19% slower in large, familiar repositories. METR's 2026 update believes newer tools are faster but could not yet estimate the gain robustly because of selection effects. DORA's core finding is that AI amplifies the quality of the engineering system around it.

Solved enough now: built-from-scratch internal tools, CRUD and reporting apps, integrations, tests, documentation, contained refactors, and replacement of thin commodity SaaS where requirements are stable and the organization can own operations.

Not solved: deciding what to build, legacy migration, identity and permissions, security, compliance, observability, uptime, user adoption, long-term maintainability, and accountability when behavior changes. The relevant comparison is not “agent cost versus developer salary”; it is three-year quality-adjusted total cost of ownership versus the incumbent workflow or SaaS contract.

Solved, Conditional, and Unsolved

“Solved” here means a competent mid-market company can deploy the pattern repeatedly with established controls and a credible path to value. It does not mean zero implementation work or zero risk. “Unsolved” means the business and operating model is not yet standardized enough to promise repeatable results.

Solved versus unsolved AI implementation problems

Source: 2026 capability and implementation synthesis

Synthesis distinguishing technically feasible patterns from standardized organizational and operating-model solutions.

Solved versus unsolved AI implementation problems
DomainSolved / repeatableRemaining unsolved edgeVerdict
Business workflowsDeterministic orchestration with AI for fuzzy subtasks and explicit escalation.Unbounded probabilistic autonomy in high-stakes, irreversible operations.Hybrid pattern solved
CodingBuilt-from-scratch tools, tests, CRUD, docs, integrations, contained changes with strong harnesses.Large legacy changes, ambiguous requirements, stewardship, security, maintenance, accountability.Production crossed threshold; not autonomous software ownership
Enterprise ROI attributionUse-case baselines, experiments, accepted-output unit economics, phased rollouts.Causal enterprise-wide attribution across simultaneous technology and process changes.Project measurement solved; enterprise attribution hard
Individual productivityDrafting, summarization, search, first-pass analysis, translation, meeting/document assistance.Turning saved minutes into realized P&L; managing review burden and workslop.Capability solved; realization conditional
Model economicsRouting, caching, prompt/context optimization, open-weight deployment for stable tasks.Predicting quality-adjusted cost as models, prices, and workloads change.Engineering toolkit mature; forecast volatile
SaaS replacementThin, expensive, weakly differentiated apps and bespoke internal workflows.ERP/core CRM/regulatory systems, complex ecosystems, migration and 24×7 operations.Selectively solved, not universal
Workforce designHuman-in-the-loop roles, AI-assisted onboarding, expert review, center-led enablement.Career ladders, entry-level learning, spans/layers, incentives, accountability, and employment effects.Strategically unsolved

A Better Way to Project and Measure ROI

Use three ledgers and refuse to blend them prematurely.

1. Cashable value

Eliminated licenses, contractors, overtime, BPO spend, losses, or avoided hires. This is the strongest evidence.

2. Capacity value

Hours saved × fully loaded hourly cost × realization rate. Use a realization rate of zero until management identifies the added output, avoided work, or role redesign that captures the time.

3. Growth and strategic value

Incremental contribution margin, faster sales or product cycles, improved retention/quality, and option value. Establish a counterfactual with A/B tests, phased rollout, matched teams, or pre/post controls.

Gross annual benefit = cash savings + realized capacity + incremental contribution margin + risk-adjusted loss reduction.

Total annualized cost = build + integration/data cleanup + change/training + human review/rework + model/tool usage + security/evaluation/compliance + maintenance + migration/exit.

ROI = (benefit − cost) ÷ cost. Also report three-year NPV, payback months, and cost per accepted output.

A board-quality business case should show downside/base/upside scenarios for adoption, success rate, realization rate, model mix, review burden, and maintenance–not a single-point estimate.

ROI measurement stack

Source: Runpoint evidence-based ROI framework

Framework synthesized from process-redesign research, controlled studies, enterprise survey evidence, and standard investment analysis.

ROI measurement stack
Value ledgerExamplesCalculationEvidence required
Cashable savingsLicense removed; BPO/contractor reduction; overtime avoided; loss reduction; avoided hire.Verified baseline spend minus post-rollout spend, net of transition and residual costs.Invoice, payroll, budget, loss, or headcount reconciliation.
Fully loaded costBuild, integration, data cleanup, training, review, inference/tools, security, maintenance, migration/exit.All one-time and recurring cash costs plus internal labor opportunity cost.Named owner, budget, usage forecast, maintenance reserve, and downside scenario.
GrowthConversion, retention, price, faster launch, cross-sell, new product revenue.Incremental units × contribution margin, adjusted for cannibalization and confidence.A/B test, holdout, matched cohort, phased rollout, or credible counterfactual.
Realized capacityMore tickets, proposals, analyses, releases, or orders with the same team.Hours saved × loaded rate × realization rate, or accepted-output increase × contribution value.Demand and throughput evidence; specific avoided work or role redesign.
Risk and qualityFewer defects, fraud losses, compliance events, rework, or service failures.Change in event probability × loss severity, plus avoided rework cost.Stable denominator, severity history, audit sample, and confidence range.

Token Economics: When the Meter Matters

Token cost is usually secondary for light, high-value assistance and can dominate low-value, high-volume, long-context agents. Current GPT-5.6 API list prices illustrate the range: Sol is $5 per million input tokens and $30 per million output tokens; Luna is $1 and $6. At 100,000 runs per year, a light 5k-input/1k-output task costs about $5.5k on Sol; a long-horizon 5M-input/1M-output run profile costs about $5.5M. Model routing changes that extreme case by roughly $4.4M before caching, tool fees, or volume terms.

The “under 150 seats are subsidized, then per-token” idea should be presented as vendor packaging, not an industry rule. Claude Team has a 150-seat ceiling and current enterprise plans are usage based; OpenAI enterprise plans include baseline access and sell additional credits, while Codex moved to token-based rates in 2026. Subsidy and bundle structure can temporarily hide inference economics, but renewal and heavy usage expose them.

Open-weight models improve sovereignty and can reduce unit cost, especially for stable, high-volume, privacy-sensitive tasks. They are not free: hardware/cloud capacity, serving, monitoring, evaluation, upgrades, and quality gaps matter. The correct break-even is quality-adjusted accepted outcomes, not raw tokens.

Annual inference cost = runs × [(input tokens × input price + cached tokens × cache price + output tokens × output price) ÷ 1M + tool fees].

Local break-even volume = annual fixed local stack ÷ (API cost per accepted outcome − local variable cost per accepted outcome).

Illustrative annual model cost at 100k runs
Illustrative annual model cost at 100k runs data
Tokens per run (input/output)Annual model costModelInput tokens/runOutput tokens/runAnnual runsAssumption
Light – 5k / 1k$5.5KGPT-5.6 Sol5K1K100KList price; no caching, tools, retries, or discounts
Light – 5k / 1k$1.1KGPT-5.6 Luna5K1K100KList price; no caching, tools, retries, or discounts
Standard – 40k / 8k$44KGPT-5.6 Sol40K8K100KList price; no caching, tools, retries, or discounts
Standard – 40k / 8k$8.8KGPT-5.6 Luna40K8K100KList price; no caching, tools, retries, or discounts
Agentic – 500k / 100k$550KGPT-5.6 Sol500K100K100KList price; no caching, tools, retries, or discounts
Agentic – 500k / 100k$110KGPT-5.6 Luna500K100K100KList price; no caching, tools, retries, or discounts
Long-horizon – 5M / 1M$5.5MGPT-5.6 Sol5M1M100KList price; no caching, tools, retries, or discounts
Long-horizon – 5M / 1M$1.1MGPT-5.6 Luna5M1M100KList price; no caching, tools, retries, or discounts

AI Sovereignty: Rent the Models, Own the Workflows

The evidence favors architectural optionality. Model quality, price, context limits, hosting, and vendor packaging move too quickly to make one model the durable asset. The assets that grow in value are the workflow specification, source code, proprietary context, integrations, permissions, evaluation set, audit history, and user adoption.

For Runpoint, sovereignty should mean:

  1. A model abstraction layer with routing by task, quality, latency, privacy, and price.
  2. Owned workflow logic and source code in the client's environment.
  3. Portable data and context with explicit access controls.
  4. An evaluation harness that makes model switching measurable.
  5. A deterministic reliable business record around probabilistic steps.
  6. An exit path for SaaS, model, and implementation vendors.

This position is strongest when framed as financial risk management and intellectual property that grows in value–not as an ideological preference for local hosting. A frontier model may still be the best rented component.

Runpoint Thesis Scorecard

The table separates claims that can be stated directly from claims that need qualification. This is also a useful editorial guardrail for a video: be most forceful where independent evidence is strongest.

Runpoint claims: evidence-graded language

Source: Independent evidence compared with Runpoint thesis

Runpoint positions treated as hypotheses and compared with independent academic, government, and cross-company evidence.

Runpoint claims: evidence-graded language
ClaimVerdictWhat supports or challenges itRecommended language
CEO-driven companies will outperform.SupportedMcKinsey, BCG, and PwC consistently link senior commitment, workflow redesign, and focused investment with higher reported value.AI ROI is an operating-model decision led by the CEO and process owners, not an IT tool rollout.
Deterministic workflows with AI sprinkled in are the known pattern.Strongly supportedControlled systems, observable outputs, human escalation, and workflow redesign align with the strongest cross-study patterns.Keep state, permissions, and commitments deterministic; use AI where ambiguity creates value.
Enterprise accounts below 150 seats are heavily token-subsidized.Vendor-specific, not universalClaude Team's ceiling and usage-based Enterprise create a packaging cliff; OpenAI uses baseline access plus credits and separate token rates.Seat bundles can hide inference cost until limits or renewal; model unit economics must be measured independently of packaging.
Open/local models make AI very cheap.Conditionally supportedWeight cost can be zero and hardware requirements have fallen, but serving, utilization, operations, and acceptance quality determine economics.Open weights create a credible low-cost and sovereign option at sufficient stable volume; they do not create zero-cost intelligence.
Rent the models, own the workflows.Supported as architectureRapid model and price change plus open weights increase the value of portability; proprietary workflows, data, and evaluation history become more valuable.Sovereignty is owning the workflow, context, code, and evaluations, not necessarily hosting the frontier model.
SaaS replacement is solved in 2026.Partially supportedCoding capability crossed a threshold; neutral studies still show task and repository dependence, and lifecycle obligations remain.Core SaaS replacement is now a credible conditional bet when three-year economics, migration, controls, and a maintenance owner are explicit.
The $20M–$200M mid-market is the best place to play.Plausible; not provenMechanisms favor decisiveness and sufficient volume, but large firms scale more often and exact-band causal ROI evidence is absent.The mid-market is an underexploited deployment sweet spot when volume, ownership, and change capacity coexist.

Recommended Mid-Market Playbook

  1. Choose one economic starting point. Target a workflow with at least one cashable baseline: SaaS, outsourcing, overtime, avoided hiring, loss, or constrained throughput.
  2. Instrument before building. Record volume, labor minutes, acceptance/defect rate, cycle time, exception rate, cost, and downstream business outcome.
  3. Redesign the workflow. Remove steps and change decision rights; do not merely insert a chat window.
  4. Keep probabilistic work bounded. Use AI for generation, extraction, classification, search, or exception handling; keep permissions, state changes, approvals, and records deterministic.
  5. Build the eval set early. Test the real edge cases, including refusal, escalation, security, and regression.
  6. Create model optionality. Route light/high-volume work to cheaper or open-weight models and reserve frontier reasoning for cases where its acceptance lift exceeds its price.
  7. Finance the rollout in gates. Prototype → shadow mode → limited production → scale. Release the next tranche only when accepted-output economics and adoption thresholds are met.
  8. Measure realization quarterly. Reconcile hours saved to avoided cost, additional throughput, contribution margin, or a redesigned role.

The strongest initial candidates are likely software and information businesses, professional and business services, and document-heavy finance or insurance companies. Operational firms also fit when there is an obvious back-office or integration problem. Avoid leading with fully autonomous agents in regulated, irreversible, or weakly observable work.

Evidence Quality and Caveats

This review prioritizes representative government and academic evidence, controlled field experiments, then large cross-company surveys. Consultancy and vendor studies are useful for operating patterns but frequently survey large enterprises, rely on self-reported value, select advanced initiatives, or use changing definitions. Associations between AI adoption and revenue/productivity do not prove that AI caused the result; faster-growing, better-managed firms may adopt earlier.

The adoption series is directional because McKinsey changed questions and definitions over time. Task-study effect sizes are intentionally shown together to demonstrate the jagged frontier, but their endpoints, participants, models, and tasks are not directly comparable. Current model capabilities and prices are especially volatile; refresh the coding and token sections immediately before publication.

The middle-market evidence commonly uses $10M–$1B or employee-count bands rather than Runpoint's exact $20M–$200M segment. Conclusions about that band are therefore mechanism-based inferences, clearly labeled as such.

Selected Source Corpus

Representative and academic: U.S. Census Bureau, AI Use Among Businesses and Microstructure of AI Diffusion (2026); OECD, AI Adoption by SMEs; NBER, Firm Data on AI; Brynjolfsson, Li & Raymond, Generative AI at Work; Noy & Zhang, Experimental Evidence on the Productivity Effects of Generative AI; Dell'Acqua et al., Navigating the Jagged Technological Frontier; METR coding-productivity and task-horizon studies.

Enterprise and middle market: McKinsey State of AI (2017–2025 series); BCG AI Radar (2025, 2026); PwC AI Jobs Barometer and AI Performance Study; Deloitte State of Generative AI in the Enterprise; RSM Middle Market AI surveys (2024, 2025); National Center for the Middle Market (2026).

Engineering and economics: DORA research; OpenAI GPT-5.6 pricing and open-weight model documentation; Anthropic agent/coding research and enterprise consumption guidance; GitHub Copilot controlled experiment.

Runpoint point of view: Karp Is Right; Before the AI Project Starts; Runpoint's site and operating model. These are treated as the thesis being tested, not independent validation.

Sources

  1. Research scope and methodology2026-07-21

    Desk research completed July 21, 2026, prioritizing representative public data and primary academic evidence, then large cross-company surveys and vendor evidence. Runpoint sources are treated as hypotheses.

  2. McKinsey – The State of AI (late 2025)DuckDB · 2026-07-21

    Global survey results on AI use, scaling, EBIT impact, company-size differences, workflow redesign, and high performers.

  3. McKinsey – State of AI longitudinal seriesDuckDB · 2026-07-21

    Runpoint synthesis of McKinsey reported organizational AI adoption for 2017–2024, extended with 2025 survey results.

  4. NBER – Firm Data on AI (2026)DuckDB · 2026-07-21

    Representative international survey of almost 6,000 CEOs, CFOs, and executives in the United States, United Kingdom, Germany, and Australia.

  5. Academic and field-study synthesis – task productivityDuckDB · 2026-07-21

    Manually reviewed primary studies. Effect values preserve each study's reported task endpoint; they are plotted together only to show heterogeneity and must not be averaged.

  6. Government and middle-market survey synthesisDuckDB · 2026-07-21

    Synthesis of representative firm-size evidence and self-reported middle-market surveys; associations are not treated as causal ROI.

  7. Cross-study initiative evidence synthesisDuckDB · 2026-07-21

    Evidence-graded ranking based on controlled task studies, enterprise operating-pattern surveys, and implementation feasibility for mid-market firms.

  8. 2026 capability and implementation synthesisDuckDB · 2026-07-21

    Synthesis distinguishing technically feasible patterns from standardized organizational and operating-model solutions.

  9. Runpoint evidence-based ROI frameworkDuckDB · 2026-07-21

    Framework synthesized from process-redesign research, controlled studies, enterprise survey evidence, and standard investment analysis.

  10. OpenAI – GPT-5.6 API pricing (July 2026)DuckDB · 2026-07-21

    Illustrative annual list-price inference cost using GPT-5.6 Sol ($5 input/$30 output per million tokens) and Luna ($1/$6) at 100,000 runs; excludes caching, tools, discounts, retries, and human review.

  11. Independent evidence compared with Runpoint thesisDuckDB · 2026-07-21

    Runpoint positions treated as hypotheses and compared with independent academic, government, and cross-company evidence.

  12. U.S. Census Bureau – AI Use Among Businesses (2026)
  13. U.S. Census Bureau – Microstructure of AI Diffusion (2026)
  14. OECD – AI Adoption by Small and Medium-Sized Enterprises
  15. BCG – Closing the AI Impact Gap (2025)
  16. BCG – AI Radar 2026
  17. PwC – AI Performance Study (2026)
  18. PwC – AI Jobs Barometer (2025)
  19. Deloitte – State of Generative AI in the Enterprise
  20. RSM – Middle Market AI Survey 2025
  21. National Center for the Middle Market – AI Adoption and Performance (2026)
  22. METR – AI Coding Productivity Uplift Update (2026)
  23. DORA – State of AI-Assisted Software Development (2025)
  24. OpenAI – Introducing gpt-oss open-weight models
  25. Anthropic – Enterprise Consumption Guide
  26. OpenAI – Flexible enterprise pricing
  27. Runpoint – Karp Is Right
  28. Runpoint – Before the AI Project Starts
The practical next step

Find the first contract or workflow that can pay for the work.

Runpoint replaces expensive SaaS and manual workflows with software built around your business. You own the source code, your data stays under your control, and we measure the work against the contracts and labor it replaces.

Free Operating Stack ROI Audit

In 48 hours, we'll show you what may be worth replacing, what it costs to own, and where the savings are.

Request the free auditSee the work first →