Moving to the cloud doesn’t make software faster. Lift-and-shift migrations carry ossified legacy code onto servers that bill for every inefficiency. AI features, which touch data far more often than ordinary software, magnify the damage. Flexera’s 2026 State of the Cloud Report found that respondents estimated 29% of their IaaS and PaaS spend was wasted, the first increase in five years, as AI workloads made forecasting harder. Skills gaps ranked among the top challenges at 30%. For mid-sized firms without a cloud engineer, the first review often goes to an outside partner, such as the IT support team at Tech River or another managed provider. That review usually starts with three questions: what does this cost, where is it slow, and who can touch what?
Why AI makes cloud inefficiency worse
On-premises, a clumsy database query costs you performance. In the cloud, it costs money too, because the platform absorbs inefficiency by adding compute and then invoicing for it. AI compounds this in three ways:
- More round trips. A retrieval-based assistant may consult a search index, a database, and a model endpoint for one answer.
- Heavier payloads. Documents, embeddings, and images move between services constantly.
- Continuous operation. OpenAI’s recently launched always-on agents keep working after the user closes the chat, each on its own cloud computer.
That last point is the insidious one. Software that runs around the clock pays for a poor design around the clock.
What should you measure first?
Teams that skip the baseline fix whatever feels slow, which is rarely what costs most. NIST’s Cloud Computing Service Metrics Description (SP 500-307) doesn’t set targets. It gives a model for defining metrics precisely enough that provider and customer agree on what they’re measuring. For AI workloads, four numbers carry most of the signal:
| Metric | What it tells you | Warning sign |
| Cost per output | What one answer or processed document really costs | Nobody can state it |
| Tail latency (p95/p99) | How the slowest requests behave | Averages look fine while users complain |
| Accelerator utilization | Whether expensive hardware works or waits | Long low stretches outside known peaks |
| Queue time | How long requests wait before processing | Rising queue time with flat traffic |
Right-sizing compute for AI workloads
Elasticity only saves money when the resource fits the job. Light preprocessing doesn’t need accelerators, and a large model on general-purpose CPUs will lag regardless of tuning. AWS’s Well-Architected performance pillar encourages teams to experiment more often, comparing instance types, storage, and configurations. In practice, that means benchmarking your actual model on two or three instance families before committing.
Underused accelerators are usually the most profligate line on an AI bill. Look first at:
- development environments left running over weekends
- inference endpoints sized for traffic that never arrived
- finished experiments whose resources were never torn down
Low utilization isn’t automatically a reason to downsize, since a server at 20% may be holding headroom for a known spike. Sustained idleness still warrants a question.
Autoscaling needs the right trigger. A busy model can saturate its accelerator while CPU graphs look placid, so queue depth or request rate is usually the better signal. Scaling to zero saves money but forces a slow model reload. That’s fine for batch jobs and painful for a customer-facing assistant, where a small warm baseline often pays for itself.
How do you reduce inference latency?
An application in London calling a database in Tokyo pays a delay on every round trip. AI multiplies those trips: one retrieval-augmented answer may query a vector database, fetch documents, and then call the model. Co-locating the app, its data, and the model endpoint sounds obvious, yet teams often miss it when they adopt a hosted model in whatever region it launched. AWS also notes that multi-region deployment can cut latency for users at minimal cost.
| Technique | Best for | Watch out for |
| Response and embedding caching | Repeated questions, common lookups | Stale answers after data changes |
| Content delivery networks | Static and front-end assets | Little help for dynamic model output |
| Asynchronous queues | Long jobs like contract summaries | Users need status updates |
| Service isolation | Stopping one slow part from stalling the app | Harder failure tracing |
Storage and data pipelines decide AI speed
Put an active training set or retrieval index on cheap, slow storage and the compute starves. You pay full rate for accelerators that sit waiting. Match the tier to how often data is read:
- Hot (SSD or in-memory): active indexes, recent transactions, current training data
- Warm: reference data used occasionally
- Cold: old checkpoints, aged logs, archives
Watch egress fees too. A pipeline that copies one dataset to three regions pays three times, and these kludges pile up quietly until someone reads the invoice line by line.
Do AI applications need microservices?
Independent services let a team swap a model without redeploying the whole app. The costs get less airtime. In 2023, Amazon’s Prime Video team rebuilt its quality-monitoring service after the distributed, serverless design proved too expensive and scaled poorly, partly because it passed video frames through an S3 bucket. Merging the components into a single process cut infrastructure costs by over 90% and improved scaling. That case covered one service, not all of Prime Video, but the lesson fits AI pipelines that shuttle large payloads between stages.
| Keep it together when… | Split it apart when… |
| Stages constantly exchange large payloads | Components ship on different release cycles |
| Latency between steps dominates | Components scale on very different curves |
| One small team owns the pipeline | You need to swap models without touching the app |
Automated security speeds AI deployment
Security slows releases mainly when it arrives late. In Flexera’s survey, 53% of cloud leaders named security and compliance as their top challenge for AI initiatives. Controls built in from day one remove most of the friction:
- Least privilege by default for every endpoint, pipeline and agent, especially agents that can take actions.
- Hardened infrastructure-as-code templates with encryption, network rules and logging preconfigured.
- Policy checks in the pipeline that block misconfigurations before production.
This matters more now that AI coding agents can turn a plain description into a deployable app. Manual review can’t keep pace with that. Pipeline policy can.
Quick diagnostic
| Symptom | Likely cause | First fix |
| Bill rising faster than usage | Idle or oversized instances | Utilization audit, shutdown schedules |
| Slow only for some users | Cross-region round trips | Co-locate app, data and model |
| First request of the day crawls | Cold model loading | Small warm baseline |
| GPUs busy, output sluggish | Storage starving compute | Move hot data to a faster tier |
| Releases stall at review | Late security controls | Hardened templates, pipeline checks |
Where to begin
- Price one unit of AI output.
- Switch off idle accelerators and forgotten environments.
- Co-locate the model, data and application.
- Fix storage tiers for hot data.
- Only then revisit service boundaries.
Providers change instance types and pricing often enough that last year’s sensible setup can become this year’s waste, so schedule the review rather than treating it as a one-off.
FAQs
Q. Why is my AI app slower in the cloud than expected?
Usually distance and data movement, not raw compute. The model, data and app may sit in different regions, or hot data may live on slow storage.
Q. What wastes the most cloud spend on AI?
For most teams, it’s idle or oversized accelerators. Flexera’s 2026 survey put self-estimated waste across IaaS and PaaS spend at 29%.
Q. Microservices or monolith for AI?
Split components that change or scale independently. Keep stages that pass large payloads together.
Related: Forward-Deployed Cybersecurity: Closing the Execution Gap
