Datadog’s Yadi Narayana on why GPU visibility gaps are emerging as the ‘next big AI opportunity’ for partners
Enterprises scaling AI are struggling with cost attribution, utilisation and governance as GPU inefficiencies begin reshaping infrastructure strategy.
As enterprises move AI workloads into production, GPU infrastructure is becoming difficult to manage efficiently. Organisations are struggling to track utilisation, attribute costs across teams and workloads, and understand whether rising AI infrastructure spend is delivering measurable business value.
The challenge becomes more pronounced as AI deployments expand across teams, models and environments, forcing organisations to rethink how GPU infrastructure is allocated, monitored and optimised at scale.
Speaking to CRN India, Datadog’s field CTO for Asia Pacific Japan (APJ), Yadi Narayana, said the inflection point does not come at the pilot stage, but when organisations begin scaling AI across multiple teams and environments.
“In pilot phases, GPU usage is deliberate and relatively contained. The inflection point comes as teams move into limited production, where multiple models, teams and environments begin to share infrastructure,” he said.
At that stage, limited visibility into workload levels leads to systematic inefficiencies.
Without clear insight, organisations are unable to understand which models are driving costs, how efficiently infrastructure is being used, or how performance links to business outcomes.
“By the time organisations reach scaled deployment, GPU costs don’t just grow, they become structurally inefficient,” said Narayana.
Without visibility into utilisation and ownership, capacity planning becomes reactive, and GPU spend starts to disproportionately influence overall compute strategy.
Where AI spend breaks first
Narayana said the first failure point in scaling AI infrastructure is not provisioning or tooling, but cost attribution.
“The first failure point is almost always cost attribution. Organisations lose the ability to map GPU consumption to specific teams, models, or business outcomes,” said Narayana.
“Once that breaks, accountability follows. Without clear ownership, optimisation becomes nobody’s priority. Scheduling inefficiencies and tooling gaps then surface, but they’re downstream effects,” he added.
The most significant impact is on delivery velocity.
Teams operate under uncertainty, unsure whether constraints are due to insufficient capacity or poor utilisation, so they default to requesting more GPUs.
That introduces procurement delays, fragments infrastructure planning, and ultimately slows the pace at which AI systems can be iterated and deployed.
AI observability changes this dynamic by reconnecting cost, performance, and ownership so teams can optimise intelligently rather than scaling blindly.
For partners managing AI environments, this shift is defining a new services layer.
Addressing GPU waste is the immediate and tangible opportunity
The demand is already moving beyond one-time optimisation exercises.
According to Narayana, customers are looking for “continuous visibility” into how AI systems behave over time, as models evolve, data changes and workloads fluctuate.
Customer demand is shifting from one-time optimisation projects to continuous AI observability and governance.
“From a monetisation standpoint, addressing GPU waste is the most immediate and tangible opportunity,” Narayana said.
He said this delivers clear near-term ROI by identifying idle or underutilised resources, making it an effective entry point for partners.
However, the longer-term opportunity lies in “managing AI systems across their lifecycle”.
This includes observability across workloads, ongoing performance optimisation, and governance models that link infrastructure consumption directly to business outcomes.
For partners, this represents a shift from managing infrastructure to managing AI performance and cost.
When visibility changes decisions
The impact of this shift is already visible in enterprise decision-making. Narayana said organisations initially “plan additional GPU purchases” to address what they perceive as capacity constraints. However, workload-level visibility often reveals a different problem.
“The constraint is not capacity, it’s utilisation. When AI observability is introduced, organisations often discover underutilised GPUs and inefficient workload distribution,” he said.
The visibility allows organisations to reallocate workloads, improve scheduling and optimise model execution, often deferring or completely avoiding additional hardware investments.
Troubleshooting moves from trial-and-error to precise diagnosis, allowing teams to identify whether issues originate in the model, data pipeline or infrastructure layer.
This directly improves iteration speed and shortens the time required to move AI systems into production.
India’s efficiency-first AI model
In India, the shift is playing out differently compared to more mature global markets.
Narayana said Indian enterprises are approaching AI with a stronger focus on efficiency from the outset, driven by tighter budgets and a higher reliance on partners.
“In India, we’re seeing a more efficiency-first approach to AI from the outset. That environment accelerates the need for cost attribution, utilisation efficiency and governance early in the AI journey,” he said.
This means Indian organisations are prioritising full-stack visibility across data pipelines, model execution and infrastructure earlier, rather than overprovisioning capacity in the initial stages.
He added that GPU underutilisation is often not the root problem, but a symptom of inefficiencies elsewhere in the stack.
“Optimising AI costs is not just about monitoring GPU usage in isolation, but understanding the full chain, from data pipelines and model orchestration to application performance and infrastructure dependencies,” he said.
As AI adoption matures, these approaches are beginning to converge globally. The need for visibility, accountability and efficient utilisation becomes universal, but in markets like India, that discipline is being built in from day one.
For the channel ecosystem, this shift is creating a new mandate.
The next phase of enterprise AI adoption is centred on visibility and operational control rather than infrastructure deployment alone. For partners, that is translating into demand for services around AI observability, utilisation optimisation and governance across complex GPU environments.