Cheaper AI models may not lower enterprise costs as agent usage grows: Report

Jul 15, 2026 - 07:09
Cheaper AI models may not lower enterprise costs as agent usage grows: Report

Enterprises are being advised to measure AI spending based on the cost of successfully completed workflows rather than advertised token prices.Falling prices for artificial intelligence (AI) models do not necessarily translate into lower costs for enterprises, as increasingly complex AI agents consume far more computing resources than traditional chatbots, Forbes reported.The trend comes as AI companies intensify price competition.

Meta launched its first paid Model API on July 9, pricing its Muse Spark 1.1 model at $1.25 per million input tokens and $4.25 per million output tokens, undercutting flagship offerings from OpenAI and Anthropic.

Earlier in the week, OpenAI rolled out its GPT-5.6 family of models, while SpaceXAI introduced Grok 4.5.

Chinese AI startup DeepSeek also made permanent a 75 per cent price cut for its V4-Pro model in May.Despite lower token prices, enterprises deploying AI agents may continue to see overall spending rise.

Unlike conventional chatbots, agentic AI systems execute multiple steps—including planning, retrieval, tool calls, validation and retries—to complete a single task, significantly increasing token consumption.Enterprises urged to measure workflow costsAs a result, enterprises are being advised to measure AI spending based on the cost of successfully completed workflows rather than advertised token prices.The report stated that organisations should distinguish among four metrics: the price per token, the number of tokens consumed per attempt, the cost of completing a successful task, and total enterprise AI spending.

These figures can move independently.

A higher-priced model that resolves requests accurately on the first attempt may ultimately prove cheaper than a lower-cost model that requires multiple retries or human intervention.The publication also highlighted the growing demand for AI inference, in which trained models generate responses rather than being retrained.

Goldman Sachs estimates token consumption will increase 24-fold between 2026 and 2030, reaching 120 quadrillion tokens per month, largely driven by always-on enterprise AI agents.Reasoning-intensive AI models are expected to further increase usage.

Meta’s Muse Spark 1.1 bills internal reasoning tokens at output-token rates, while OpenAI’s GPT-5.6 Ultra mode deploys multiple AI agents simultaneously to solve complex tasks, trading higher token usage for improved speed.AI deployment costs extend beyond model APIsAI model API charges represent only part of the overall cost of deploying enterprise AI.

Additional expenses include search and retrieval, vector databases, reranking, browser automation, code execution, monitoring systems and human review.The publication also said evaluating AI models has become more challenging.

OpenAI recently disclosed that it no longer recommends SWE-Bench Pro as a leading coding benchmark after identifying flaws in roughly 30 per cent of its tasks, underscoring the importance of enterprises conducting their own workflow-specific evaluations.Lower list prices may also offer limited negotiating leverage where enterprises rely on proprietary tools, have long-term volume commitments or operate under strict compliance requirements that limit provider choice.The publication concluded that enterprises should focus on measuring AI costs per successful outcome, monitor complete workflows before scaling deployments and regularly update financial models as AI pricing continues to evolve rapidly.

ସ୍ପଷ୍ଟୀକରଣ: ଏହି ବିଷୟବସ୍ତୁଟି ସୂଚନାମୂଳକ ଉଦ୍ଦେଶ୍ୟରେ Enterprise AI ରୁ ସ୍ୱୟଂଚାଳିତ ଭାବରେ ସଂଗ୍ରହ କରାଯାଇଛି। ମୂଳ ଲେଖାଟି ପଢ଼ିବା ପାଇଁ, ଦୟାକରି ଏଠାରେ ଦେଖନ୍ତୁ।

indianiaiac

IAIAC.IN is India's first Safe, Trusted & Reliable AI Applications Center, dedicated to championing Responsible AI practices throughout society.