Tag: time to first token

  • The New AI Math: Time-to-Token and Cost-per-Token Gets Highlighted at Dell Technologies World

    The New AI Math: Time-to-Token and Cost-per-Token Gets Highlighted at Dell Technologies World

    At Dell Technologies World this morning, Michael Dell introduced new metrics for measuring whether enterprise AI infrastructure is actually delivering. The AI infrastructure conversation has been dominated by GPU counts, cloud-versus-on-premises debates, and model benchmarks. Dell’s opening keynote added two measures that tie those inputs to outcomes. The first is  how quickly your infrastructure generates tokens, and the second is at what cost.

    Uptime, GPU capacity, and model advances are all foundational. GPUs are what get you to tokens in the first place. But they are not enough on their own to make infrastructure decisions. Organizations are now facing a more challenging optimization problem. Companies face a seemingly endless demand for AI compute and the energy required to support it. Businesses must also balance the performance of these evolving AI workloads within budgetary and location constraints.

     Time-to-token and cost per token are the metrics that tie these competing priorities together. They measure how quickly and how cheaply your infrastructure converts data and compute into intelligence that agents and models can act on. Jensen Huang reinforced this on stage alongside Dell, and OpenAI’s Greg Brockman made the same point independently on X today: “tokens are rapidly becoming the universal input for solving problems.” When the infrastructure providers and the model providers converge on the same metrics within weeks of each other, that is a directional signal worth acting on.

    What Michael Dell described was a two-year refinement of the Dell AI Factory with NVIDIA, informed by 5,000 enterprise customers now running production AI workloads on it. The announcements were substantial. What they mean for enterprise buyers is worth unpacking.

    The Data Bottleneck Is the Real Constraint

    Michael Dell said something on stage that every CIO needs to hear: “If your data is siloed, your agents are blind.” That is a concise description of why so many enterprise AI programs stall after the pilot. It also echoes what Irfan Khan and Muhammed Alam described as the need for business context at SAP’s Sapphire and in an online event about business data

    Most organizations are simultaneously trying to prepare data for AI and reengineer the data infrastructure needed to support it. Those are two hard problems happening at the same time, on top of each other. Dell’s announcements around the Dell AI Data Platform addressed both directly.

    One of the things that has changed is the data orchestration engine, the intelligent control center within the Dell AI Data Platform that turns raw, fragmented enterprise data into production-ready AI fuel. It indexes billions of unstructured files of all types, builds governed data pipelines, connects them to the models and agents that need them, and delivers structured outputs at speeds that make agentic workflows viable. Dell claims the platform now delivers 12 times faster vector indexing, 6 times faster data querying, and 19 times faster time to first token than prior generations. While the claims still need to be verified, the direction is exactly what enterprise buyers need.

    Why does this matter? AI agents need business context to be useful. An agent that can reason brilliantly but cannot reliably access your CRM, internal knowledge bases, operational systems, or proprietary data is not doing useful work. The data orchestration engine is what connects the model to the context. Without it, you have a powerful system with nothing meaningful to act on.

    Dell’s approach integrates orchestration, search, and governed pipelines natively into the platform. Getting that platform connected to your actual data sources still requires integration work. Budget for it before you buy the hardware.

    AI Infrastructure Is a Team Sport

    The broader lesson from the Dell Technologies keynote is not about any specific product. It is about what has changed in two years.

    When Dell announced the Dell AI Factory with NVIDIA in 2024, it was largely a hardware and partnership story. Today, with 5,000 enterprise customers running production workloads on it, the conversation has shifted to execution. How do you get from pilot to production? How do you keep agents from being blind to your actual business data? How do you manage cost curves as token consumption scales? How do you maintain security and governance when agents are operating autonomously at machine speed?

    Part of Dell’s answer is its ecosystem. The new Dell AI Ecosystem Program gives AI software providers a validated path to certify solutions on Dell infrastructure. For enterprise buyers, this speeds AI deployments by reducing integration some of the integration burden. Rather than assembling a custom stack from scratch, you get pre-validated blueprints that automate the deployment of software, services, and models together. That automation is a direct lever for reducing time-to-token at the program level. Dell claims it can deliver hundreds of AI racks a week to a given customer and have them generating outcomes within hours. Part of this is also achieved with ecosystem partner blueprints for automation. 

    The ecosystem also extends to the agent layer itself. Jensen Huang described on stage how agents do not run directly on the large language model. They run on a harness. The harness sits in a secure, governed container called a sandbox. It manages the agent’s reasoning loop, handles tool use, controls what data and systems the agent can access, and determines when to call the larger model and when to use a smaller local model instead. NVIDIA’s OpenShell is the open-source sandbox now supported across the entire Dell AI Factory. For enterprise buyers evaluating AI infrastructure, the harness is not a detail. It is a primary evaluation criterion. An infrastructure stack that does not clearly define how agent harnesses are deployed, secured, and governed is not production-ready.

    The ecosystem partners Dell named today include Google, Hugging Face, OpenAI, Palantir, ServiceNow, and SpaceXAI, among others. AI is not a solo deployment. The strength of the ecosystem around the infrastructure determines how fast you can actually move.

    Michael Dell put the security dimension plainly: “You can’t protect what you can’t see, and you can’t manage what you can’t see.” That applies to agents as much as it applies to data. Agents have credentials, memory, and access to systems. When they operate autonomously at machine speed, the blast radius of a security failure is no longer contained to one system. It can propagate across workflows and infrastructure. There are also tech tools such as X that help with confidential computing. 

    The Token Economics of Hybrid AI

    Sixty-seven percent of AI workloads already run outside the public cloud, on-premises, at the edge, or in co-location environments, according to Dell’s own survey data. Eighty-eight percent of organizations are running at least one AI workload on-premises. But, that does not mean cloud is going away. It means the real question for enterprise infrastructure leaders is not cloud versus on-premises. It is how to run both well.

    Hybrid AI is not a compromise. It is the architectural reality for most large enterprises. Some workloads need the cloud, which offers speed, training, high capacity, and flexibility. Others belong on-premises because it may access sensitive data that a company doesn’t want in the cloud or regulations require to be in a certain place. There may also be high-volume, continuous inference, where unpredictable cloud token costs create real budget exposure. The strategic challenge is matching the workload to the right environment and doing it consistently at scale.

    Jensen Huang described on stage why the compute requirements have shifted so dramatically. Agentic systems require 100x to 1,000x more computation than responding to a simple query because the agent has to reason, plan, use tools, evaluate results, and iterate. At that scale of consumption, every infrastructure decision has a direct cost consequence.

    Dell’s answer for high-volume on-premises workloads is what it calls “unmetered intelligence.” The idea is that owning infrastructure converts variable cloud API spend into a fixed infrastructure cost. Dell claims organizations can break even on API costs compared to the public cloud in as little as 3 months with desk-side agentic AI configurations.

    Balancing this correctly requires thinking about four variables simultaneously. Latency measures the time it takes for a system to process a request and return a response. Performance refers to the overall capability, accuracy, and capacity of the AI model to handle complex tasks. Cost is what you pay to get those outputs at the required latency and performance level. Energy is the fourth variable, and it is no longer theoretical. A single rack of NVIDIA Rubin GPUs can draw over 130 kilowatts.

    Granted, the average enterprise won’t be running a rack of Vera Rubin’s, but energy availability is becoming a real constraint regardless of sustainability goals. And if you’re using the cloud, you pay one way or the other for that energy. The right infrastructure solution varies by workload type. The time to token and the cost per token let you compare options on the same terms.

    Sovereign AI Is Becoming a Procurement Reality

    Two years ago, sovereign AI was a concept mostly discussed in European regulatory contexts and by a small number of governments building national AI infrastructure. Today, it shows up in enterprise procurement conversations across regulated and unregulated industries.

    Sovereign AI means the ability to independently develop, deploy, and govern AI systems entirely within an organization’s strategic, legal, and jurisdictional boundaries. For enterprises, this means your data does not leave your environment, your model choices are not constrained by a hyperscaler’s catalog, and your AI outputs are not subject to external policy changes.

    The ecosystem Dell announced is designed to both speed AI deployments and address sovereign AI requirements. Google’s Gemini 3 Flash models running on-premises via Google Distributed Cloud on Dell PowerEdge servers. OpenAI’s Codex is connected to the Dell AI Data Platform for agentic workflows on enterprise data. Palantir’s Foundry and AIP platform is deployed on-premises with Dell ObjectScale and PowerFlex as the data layer. SpaceXAI’s Grok is available in on-premises or hybrid enterprise deployments. Reflection’s open-source frontier models for regulated industries and sovereign entities.

    The pattern is consistent: bring the model to the data rather than the data to the model. For organizations in healthcare, financial services, defense, and government, this is not a preference. It is often a compliance requirement.

    Planning for Hybrid AI

    Today’s AI question is how to architect a hybrid AI solution that aligns with our organization’s specific workloads, data environment, cost constraints, and governance requirements. Some of that runs on-premises. Some runs in the cloud. The mix differs across organizations and will shift as workloads evolve and model costs change.

    The questions worth asking now: What is your time to first token across your most important workloads? What is your cost per token at scale? Does your data orchestration layer connect your proprietary data to the models that need it? How are your agent harnesses deployed and governed? And do you have the security architecture in place before your agents start making autonomous decisions?

    It’s not easy but nothing worthwhile ever is. 

     

  • Dell Shares AI Advances And New Metrics To Evaluate Infrastructure

    Dell Shares AI Advances And New Metrics To Evaluate Infrastructure

    At Dell Technologies World in Las Vegas, Dell Technologies chairman and CEO Michael Dell made a pointed argument to a room full of enterprise technology leaders: the metrics organizations use to evaluate infrastructure are evolving.

    GPU counts, cloud versus on-premises comparisons, and model benchmarks have dominated the conversation. Michael Dell’s day one keynote introduced two additional measures aimed at tying infrastructure decisions to actual outcomes: time to token, which measures how quickly a system processes a request and returns a usable AI output, and cost per token, which measures how cheaply that output is produced at scale.

    “Time to first token is incredibly important with investments of this scale,” Dell said on stage, noting that the company now has 5,000 enterprise customers running production AI workloads on its Dell AI Factory with NVIDIA platform. The figure represents a significant increase from the program’s launch two years ago.

    NVIDIA founder and CEO Jensen Huang, appearing alongside Dell, described why those metrics have taken on new urgency at both its NVIDIA GTC conference and at Dell Technologies World. Agentic AI systems, which reason, plan, and execute tasks autonomously over extended periods, require anywhere from 100 to 1,000 times more computation than a system simply responding to a query. “What took months now takes weeks, what took weeks now takes days, and what takes days now takes hours,” Huang said, describing the productivity transformation already underway at companies running agentic workflows. The demand implications for infrastructure are substantial.

    OpenAI president Greg Brockman echoed the framing independently on X.com the same day, writing that “tokens are rapidly becoming the universal input for solving problems.” The convergence of infrastructure vendors and model providers on the same metrics within weeks of each other signals a broader shift in how enterprise AI spending will be evaluated.

    The Data Problem Underneath the Infrastructure Problem

    One of Dell’s significant AI product announcements centered on a new data orchestration engine in the Dell AI Data Platform, which the company positioned as the missing layer between enterprise data and production-ready AI agents.

    The data orchestration engine is the platform’s intelligent control center. It indexes billions of unstructured files, builds governed data pipelines, and connects them to the models and agents that need them at speeds designed to make agentic workflows viable. Dell claims the updated platform delivers 12 times faster vector indexing, six times faster data querying, and 19 times faster time to first token compared to prior generations. While these claims still need to be validated, the proposed increase in performance is good news for enterprises looking to scale AI.

    The underlying problem the engine addresses is one most large organizations know well. Enterprises are simultaneously preparing existing data for AI use and reengineering the data infrastructure required to support AI workloads at scale. Those two efforts compete for the same resources and skills simultaneously.

    “If your data is siloed, your agents are blind,” Dell said. The statement is a precise description of why many enterprise AI pilots have not reached production. An agent operating without access to an organization’s proprietary data, internal knowledge bases, and operational systems cannot deliver the business context that makes agentic AI useful.

    Dell also announced GPU-accelerated SQL analytics through the Dell Data Analytics Engine, powered by Starburst, delivering up to six times faster query performance on NVIDIA Blackwell GPUs. Bank of America, which already has a partnership with Starburst, NVIDIA, and Dell, is among the institutions expected to use the capability.

    A Broad Ecosystem Built to Reduce Time to Production

    Dell announced a new Dell AI Ecosystem Program alongside a significant expansion of frontier model partnerships, positioning both as mechanisms for reducing the time between infrastructure procurement and production AI deployment.

    On the model side, Dell announced collaborations bringing several major AI providers on-premises to the Dell AI Factory. Google and Dell are collaborating to run Gemini 3 Flash models via Google Distributed Cloud on Dell PowerEdge XE9780 servers, enabling enterprises to run advanced generative AI workloads in a confidential computing environment that meets data residency and sovereignty requirements. OpenAI’s Codex will connect with the Dell AI Data Platform, giving enterprises a path to deploy agentic coding capabilities against their internal codebases, documentation, and business systems. SpaceXAI’s Grok is available in on-premises or hybrid enterprise deployments. Palantir’s Foundry and AIP platform is coming on-premises with its Ontology layer deployed on Dell ObjectScale and PowerFlex, allowing organizations to connect data sources and automate business workflows within their own environment.

    The Dell Enterprise Hub on Hugging Face gives enterprises on-premises access to a curated collection of open-weight models including MiniMax-M2.7, DeepSeek Pro, DeepSeek-V4, GLM 5.1, and Kimi K2.6, optimized for Dell AI Factory infrastructure.

    The Dell AI Ecosystem Program formalizes the partner relationship by providing software providers with a validated path to certify their solutions on Dell infrastructure. For enterprise buyers, the practical benefit is pre-validated deployment blueprints that automate the configuration of a specific software, service, or model, reducing integration work that has historically extended timelines from procurement to production.

    The Agent Harness: An Evaluation Criterion Enterprises Are Not Yet Asking About

    One of the more technically substantive moments in the keynote came from Huang’s description of how agents actually operate in production. Agents, he explained, do not run directly on the large language model. They run on a harness, a software layer that sits in a secure, governed sandbox. The harness manages the agent’s reasoning loop, controls tool access, determines when to call a large external model and when to use a smaller local model, and handles memory and context across multi-step tasks.

    NVIDIA’s OpenShell, the open-source sandbox, is now supported across the entire Dell AI Factory from deskside workstations through PowerEdge data center servers.  Dell also announced support for NVIDIA AIQ i.0 blueprints, which provide tested foundations for deploying multi-agent workflows.

    For CIOs evaluating AI infrastructure, the harness architecture is a meaningful addition to the evaluation checklist. Infrastructure that does not clearly define how agent harnesses are deployed, governed, and secured leaves a significant operational and security gap, particularly as agents acquire credentials, access enterprise systems, and take autonomous actions at machine speed.

    For example, “You can’t protect what you can’t see, and you can’t manage what you can’t see,” Dell said, framing the security challenge in terms that apply as directly to agents as to human users. An agent with compromised access or misconfigured permissions can propagate errors or security failures across workflows in ways that a single human user cannot.

    Hybrid AI Infrastructure and the Energy Constraint

    Dell’s survey data shows that 67% of AI workloads are already running outside the public cloud, and 88% of organizations are running at least one AI workload on-premises. The company positioned hybrid AI not as a transitional state but as the long-term architecture reality for most large enterprises.

    The new Dell PowerRack, announced Monday, is a fully integrated rack-scale system that combines compute, networking, and storage, engineered and validated as a single unit. It is designed to reduce the integration overhead of assembling AI infrastructure from components while supporting thermal management and power optimization at rack scale.

    Dell also introduced the Dell PowerCool CDU C7000, the first rack-mount cooling distribution unit designed to meet the cooling requirements of the NVIDIA Vera Rubin NVL72 platform, delivering more than 220 kilowatts of cooling capacity in a 4U form factor. A single rack of NVIDIA Rubin GPUs can draw over 130 kilowatts of power, and Dell noted that energy availability is an increasingly real constraint on AI deployment timelines, independent of sustainability considerations.

    For high-volume on-premises workloads, Dell introduced Dell Deskside Agentic AI, pairing high-performance Dell Pro Precision workstations with NVIDIA NemoClaw. The company claims the configuration enables enterprises to break even against public cloud API costs in as little as 3 months, converting variable token costs into a fixed infrastructure investment.

    What Changes for Enterprise Buyers

    The announcements from Dell Technologies World day one collectively continue to move the enterprise AI infrastructure conversation from capability to faster execution. The core questions are how quickly a given infrastructure configuration can reach first token on a production workload, at what cost per token, and with what governance architecture underpinning the agents running on it.

    The organizations best positioned to answer those questions are the ones that have already started rationalizing their data architecture, defined their hybrid workload placement strategy, and begun evaluating how agent harnesses will be secured and governed. The infrastructure improves almost daily, but the execution discipline required to use it remains the variable that separates AI programs that reach production from those that stay in pilot.

    Maribel Lopez is the founder and principal analyst at Lopez Research, a market research and strategy consulting firm specializing in enterprise AI, AI infrastructure, agentic systems, AI governance, and AI-driven customer experience. I version of this article of originally posted on Forbes.com.