AI workloads do not all belong in the same place: Vinod Bijlani’s guide for CIOs argues that cost, data boundaries, latency and operational capacity should decide what runs in the cloud, on specialist GPU services, on-premises or at the edge. Its most useful advice is practical: compare the cost of completing a business task, not just the price of an hour on a GPU.
Watch Desk analysis
What happened
Bijlani, an AI practice leader at Hewlett Packard Enterprise, sets out a framework for deciding where enterprise AI workloads should run. His guide covers managed cloud platforms, specialist GPU providers, self-hosted infrastructure and edge computing, while stressing that one workflow may span several of them. Read Bijlani’s guide.
The practical starting point is to identify constraints such as response deadlines, connectivity and permitted data flows. Then compare the remaining options against quality, expected demand, operational capability and total cost. Bijlani cites a Gartner forecast of $42 billion in AI-optimised infrastructure-as-a-service spending in 2026, with inference accounting for 55 per cent. Those are forecasts cited in the guide, not a guarantee about any particular organisation’s costs.
Our top picks
- Start with constraints
Response deadlines, connectivity and data rules can rule out an otherwise attractive deployment option. - Follow the data
Prompts, logs and outputs matter too; a private database does not keep retrieved passages private if they are sent to an external model. - Compare equivalent services
GPU memory, interconnects, storage, capacity guarantees and support all affect whether a cheaper hourly price is actually cheaper. - Measure the whole task
Include retrieval, data transfer, retries and human correction, not just compute, when comparing cost per successfully completed job. - Benchmark demand
Test realistic usage and peak loads at the required quality and response time before choosing where a workload should live. - Plan for change
Review placement as volumes, models, prices or requirements shift, and define approved fallbacks before an outage forces a decision.
Why it matters
A placement choice is also a decision about who operates the system, what data crosses a boundary and what happens when demand changes. A managed model endpoint leaves different responsibilities with a team than self-hosted infrastructure; neither ownership nor a compliance badge, by itself, settles whether an application is secure or compliant.
The guide’s strongest point is that “AI” is not a useful unit of infrastructure planning. A low-sensitivity assistant with irregular demand and a manufacturing inspection that must respond on time have different needs, even if both use models. One workflow might reasonably split between edge processing, private data retrieval and an approved cloud model, provided the extra boundaries earn their keep.
Our read
This is useful CIO advice because it turns “cloud or on-prem?” into a decision a team can test: specify the outcome, trace the data, model the demand and compare the full operating cost. Bijlani also proposes a shared layer for routing and governing access across providers and private infrastructure. That could make switching easier, but a tidy abstraction is no substitute for testing whether different models behave well in the actual application.
What to watch
- Whether teams benchmark cost per completed task, including retries and human correction.
- How often real workloads cross environments, and whether the added complexity brings a measurable benefit.
- Whether hybrid routing systems provide tested fallbacks, clear data controls and useful audit trails.
Discussion spark: When choosing where an AI workload runs, should data control and reliability outweigh the cheapest measured cost per completed task, or should teams let the workload decide case by case?
Sources and evidence
- Which AI Workload Goes Where? A CIO's Guide To Hybrid AI – Forbes (8 October 2026, 11:00 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.