Wednesday, September 23, 2026

Dell and Nvidia take agentic AI from cloud to desk

Organisations that design workloads efficiently could create a cost advantage, while others could create what Dell's CSO termed “a structural cost problem at scale.”

A team inside Dell Technologies used 25 million tokens to tackle an agentic AI problem by loading information into a large context window. Another team addressed the same task with a coordination agent that orchestrated tools, using 20,000 tokens.

The difference is roughly 1,250 times fewer tokens. Dell Technologies’ chief strategy officer Sam Burd told attendees at Dell Technologies Forum in Singapore that the example illustrates a growing enterprise risk – agentic AI can become an expensive form of automation.

Sam Burd Dell Tech
Sam Burd at Dell Tech Forum 2026 in Singapore

“Same problem, two architectures, a thousand-fold difference in cost,” Sam said.

The lesson is not simply that businesses should use fewer tokens. It is that agent architecture will increasingly determine the economics of enterprise AI. 

Systems that are persistent, context-heavy and capable of executing multiple steps can consume far more compute than a one-off chatbot interaction. Without clear guardrails, model routing and workflow design, AI agents may add a significant and unpredictable operating cost.

“Even if cost per token falls, aggregate AI spend keeps climbing,” Sam said, pointing to AI reasoning usage that he said has risen 320 times. 

Tokens, he said, are becoming a line item on the P&L. Organisations that design workloads efficiently could create a cost advantage, while others could create what he termed “a structural cost problem at scale.”

From experiments to infrastructure

That concern sits against an AI investment cycle that Sam believes has repeatedly exceeded expectations. He said estimates from around two years ago put hyperscaler AI capital expenditure at approximately USD250 billion. Current spending, he said, has since climbed past USD800 billion this year.

“We’ve been undercalling that forever,” Sam said. “Probably a good chance that that number will be undercalled, especially as we look at the future.”

Even if cost per token falls, aggregate AI spend keeps climbing

Sam Burd

The shift reflects the accelerating capabilities of AI systems and expanding expectations around their use. For enterprises, Sam said, experimentation is no longer sufficient: AI “absolutely has to scale.”

However, the infrastructure foundation remains incomplete. Dell said that 76-percent of organisations in Singapore reported that their data centres were not ready for demanding AI-era workloads. Dell’s proposed response is an integrated technology stack intended to simplify operations and standardise the data centre as AI deployment expands.

Dell’s AI proposition

Dell is positioning local AI as part of the answer to rising token consumption. Its Deskside Agentic AI approach is intended to let organisations build, test and run selected agents on premises, rather than automatically sending every model call to a public-cloud API.

the Dell Pro Max with GB10, an AI developer workstation powered by NVIDIA’s GB10 Grace Blackwell Superchip

The company’s work with NVIDIA centres on the NemoClaw reference stack and NVIDIA OpenShell. NemoClaw is built on OpenClaw and incorporates OpenShell, a secure runtime designed to isolate and govern agent activity, alongside NVIDIA’s open models and tooling. Dell says the stack can support agent deployment from deskside systems to data-centre infrastructure.

During Dell Technologies World in Las Vegas, NVIDIA chief executive Jensen Huang described the architecture as hybrid AI: open models can be run locally, while frontier models can be accessed through the cloud when a task requires them. 

According to Sam, the same workload that might generate substantial cloud-model charges could have close to zero marginal token cost when it is run locally on a device.

That does not mean local AI is free. But, its advantage is a shift away from recurring, metered public-cloud token charges towards more predictable infrastructure spending – an approach likely viable for sustained, high-volume or data-sensitive workloads.

Deskside AI

Defining the AI agent’s identity/role and tasks, or Soul and Skills, respectively.

At the Singapore forum, attendees were able to use the Dell Pro Max with GB10, an AI developer workstation powered by NVIDIA’s GB10 Grace Blackwell Superchip. The system is designed to run AI development and inference workloads locally, with 128GB of unified memory enabling the use of larger open-weight models than conventional desktop GPUs typically accommodate.

The event workshop used NVIDIA’s NemoClaw/OpenShell architecture and pre-loaded large language models to show how enterprises can create governed agentic systems. Dell framed that process around three operational elements – an agent’s identity and role, the task it is permitted to do, and policies that govern its access to data, systems and actions.

For Dell, the message is that enterprise AI will not be defined only by access to the largest models. It will depend on whether businesses can make agents secure, governable and economical enough to operate at scale.

Powered byspot_img

Read more

News

Powered byspot_img