Tag: sovereign AI

  • HPE Updates Hardware, Private Cloud And Networking For  Agentic AI Era

    HPE Updates Hardware, Private Cloud And Networking For Agentic AI Era

    When public cloud computing emerged in the late 2000s and began scaling through the early 2010s, the prevailing theory was straightforward. Cloud computing theory claimed nearly every workload would eventually migrate off-premises. The economics were compelling, the convenience undeniable, and the momentum felt unstoppable. More than 15 years later, enterprise IT leaders are still managing substantial on-premises infrastructure — and many are actively investing in upgrades.

    Today, various research reports estimate that between 35 and 50 percent of workloads have moved to the cloud. Whether that figure is above or below 50 percent, it’s clear that workloads remain distributed. Cost, data sensitivity, regulatory compliance, latency requirements, and operational control all shape where a given workload belongs. The public cloud became one option in a complex portfolio, and the same will be true for AI workload placement. 

    Organizations have absorbed that lesson. In the early days, everyone ran AI proofs of concept in the cloud, but enterprise leaders are asking more specific questions on how to design scalable AI architecture. Today’s discussion centers on which AI workloads belong where and under what conditions? Cost recently surfaced as a major concern as AI token use skyrocketed. In many cases, AI costs escalated due to flawed policies that incentivized employees to consume as many tokens as possible to prove they were using AI.

    Today, organizations are considering the rationale for keeping workloads on-premises and whether upgrading their on-premises technology will be a cost-benefit or a disadvantage. It is not an easy question to answer. No technology vendor has delivered a definitive framework or spreadsheet that simplifies that decision. What has happened is that every major cloud and hardware vendor now offers some combination of managed AI services and pre-validated AI factory reference designs intended to help organizations scale AI deployments beyond proof of concept.

    The Hybrid AI Reality Is Already Here and Continues to Gain Momentum

    Walk into a strategy conversation with a large enterprise today, and you will rarely encounter a pure cloud or pure on-premises AI strategy. Hybrid deployments are the operating assumption. The more substantive discussion is about placement logic — what drives a workload toward private infrastructure, what drives it toward a public cloud, and what governance and connectivity model bridges the two. It is worth noting up front that HPE, as a hardware and infrastructure company, has a clear commercial interest in upgrading on-premises deployments. That context does not invalidate the market need, but it is relevant when evaluating how the company frames its positioning.

    Regulated financial institutions, healthcare systems, defense contractors, and national governments all have data residency requirements, sovereignty mandates, and security classifications must consider partitioning workloads between cloud and updated on-premises infrastructure.

    The AI Infrastructure Market Responds to Private and Sovereign Demand

    In 2026, most vendors are discussing which infrastructure advancements are needed to support agentic AI. While there were many hardware announcements at HPE’s annual Discover conference, the company also discussed updates to private cloud and sovereign AI infrastructure for agentic AI. HPE launched its first private cloud offerings roughly 2 years ago. 

    HPE is not alone in this space. Dell, Lenovo, and others have announced comparable on-premises AI infrastructure products built around Nvidia accelerators. Google Cloud was among the early movers in offering air-gapped, disconnected cloud offerings for regulated industries. The differentiation between these offerings — at the architecture, software, and services layer — is still being established in the market and warrants scrutiny from buyers evaluating alternatives. 

    HPE CEO Antonio Neri described the AI infrastructure decision as inseparable from data governance and sovereignty. Lopez Research has found this framing is consistent with what enterprise buyers in regulated industries report when asked about deployment constraints. HPE is delivering a pre-validated, purpose-built on-premises environment for AI workloads that reduces integration complexity and accelerates deployment timelines relative to assembling components independently. HPE also announced a Sovereign AI Factory configuration targeting governments and regulated industries, with built-in defense-grade security hardening, federal compliance readiness, and air-gapped operation. Let’s talk specifically about some of the ways HPE is addressing the agentic AI challenge. 

    Agentic AI Adds New Complexity to the Security and Governance Problem

    Neri articulated a theme that was consistent across the 2026 enterprise technology conference season when he discussed how AI has moved beyond generative assistants to autonomous AI agents. Technology vendors have discussed AI agents for more than a year. Many enterprises are already encountering agentic AI through their existing SaaS platforms. More advanced organizations are building and deploying their own agents. Lopez Research’s conversations with early adopters consistently surface the same operational pain points, including multi-agent orchestration across enterprise applications, securing and permissioning AI agents, governance, and company-wide observability into what agents are doing.

    Agentic AI introduces a security and governance surface for every organization that most existing enterprise security stacks were not designed to handle. Today’s security products were designed for individuals, not AI agents. These solutions are anchored on user credentials, access policies, and behavioral patterns. An AI agent, if allowed to do so, can operate autonomously, continuously, and at machine speed across multiple systems simultaneously. Bad situations propagate across workflows fast before any human reviewer is aware that something has gone wrong.

    HPE’s response at Discover included a three-tier identity model for agentic workloads that includes user verification, agent-level governance, and human approval gates for sensitive actions. It also offers the ability to wrap agents built in any framework with security controls, including API protection, identity management, and encryption, without requiring code changes. Integration with Nvidia OpenShell provides isolated execution environments per agent. NeMo Guardrails enforce policy at the model level. Zerto integration enables rollback to a clean state if an agent executes incorrectly.

    Private Cloud AI now includes a governed data layer with deep integration into the Nvidia AI Data Platform, giving enterprises a unified way to access, prepare, and manage data across their existing environments — no custom pipelines required. The HPE Alletra Storage MP Extend 1000 serves as the storage foundation, purpose-built for the performance demands of modern AI workloads. It adds real-time metadata enrichment and native MCP support, so agents and applications can retrieve the right data and context faster across both structured and unstructured data. HPE claims the result is a 7–12x faster time to value compared to building the environment yourself.

    Once your data is governed and ready, Private Cloud AI delivers the infrastructure to scale inference. Multi-node inference allows larger models to be served across multiple systems, so capacity grows naturally with demand. A new unified gateway gives teams a single API for accessing both frontier and open-source models, with centralized credentials, budgets, and policies built in. New configurations now scale up to 256 GPUs, which includes the new ProLiant DL394 with Nvidia GPUs optimized specifically for inferencing and long-context workloads. Additionally, shared KV cache capabilities eliminate the need to repeatedly recompute context, reducing cost per first token and delivering significant performance gains across the board.

    No architecture from any specific vendor will ever fully address a company’s governance or security problems. Still, buyers need to ensure that their vendors are addressing the problem and that they are willing to work with others to support a holistic approach. What HPE’s offering does represent is a concrete, specific engineering response to a recognized gap. Rami Rahim, HPE’s EVP and President of Networking, expanded on this in a separate day 2 session, arguing that the network itself must become an active enforcement layer for agentic security through zero-trust architecture, AI-driven anomaly detection, and automated policy enforcement. 

    Familiar AI Themes, With A Focus on Execution

    Assessed across the first half of the 2026 enterprise technology conference season, HPE Discover did not introduce themes that were new to the industry conversation. Sovereign AI requirements, the governance gap in agentic systems, and hybrid placement logic have all been visible in analyst briefings, vendor roadmaps, and customer conversations for some time. 

    That observation, however, should not diminish the significance of what Neri and Rahim presented. Identifying a trend early is necessary but not sufficient. The value of HPE Discover lies not in the novelty of the concept but in the focus of execution. What HPE customers were looking for was the degree to which the company has translated an accurate read of market direction into deployable infrastructure that enterprises and governments can procure and operate today. Whether HPE’s implementation proves durable competitive differentiation in a space where competitors are moving quickly is a question the next twelve months will answer. 

    HPE’s position is architecturally sound and consistent with broader industry direction. The customer deployments already underway illustrate there’s real demand in this space from various industry segments. The U.S. Defense Information Systems Agency (DISA) awarded HPE a ten-year contract to modernize its digital and AI platform capabilities, requiring a NIST-compliant private cloud environment that meets federal security classifications. In Europe, HPE is building the HammerHAI system at the High-Performance Computing Center Stuttgart (HLRS) in Germany. It is a sovereign AI installation delivering more than 15 exaflops of peak AI inference performance for research institutions and industrial organizations that must comply with European data residency requirements. In healthcare, St. Jude Children’s Research Hospital is using HPE Private Cloud AI to bring AI capabilities to its clinical and research teams while protecting sensitive pediatric oncology data. These three deployments — federal defense, national research infrastructure, and regulated healthcare — represent the segment of buyers for whom private and sovereign AI is a requirement, not a preference.

    Scaling AI Requires A Portfolio Approach

    The cloud never won all of the workloads. The economics, the regulations, and the operational realities of enterprise IT ensured that on-premises infrastructure remained relevant long after public cloud momentum suggested otherwise. The same dynamics are now shaping AI infrastructure decisions. 

    For organizations in regulated industries, private and sovereign AI is not a conservative hedge against innovation. For many, it is the enabling condition for AI adoption at all. But private AI infrastructure is not a complete exit from cloud services, and buyers should be cautious about treating it as such. 

    Even organizations running heavily on-premises AI environments (regulated or not) will almost certainly rely on the public cloud for a portion of their AI workloads. The question is how much. AI model training is the most frequently cited example — training large foundation models requires burst compute capacity that few enterprises can justify owning outright. But the dependencies extend further. AI experimentation and prototyping typically benefit from the speed and low commitment of cloud environments before workloads are validated for on-premises production. And accessing specialized frontier models — from providers like Anthropic, Google, OpenAI, or open-source alternatives — will often happen via cloud APIs, particularly for capabilities that do not require fine-tuning on proprietary data.

    In some cases, even those cloud touchpoints will need to meet sovereign requirements. Major hyperscalers have developed sovereign cloud offerings for regulated industries. A growing tier of neoclouds is positioning explicitly around sovereign infrastructure for organizations requiring local data residency, compliance certification, and jurisdictional control within a managed cloud model. The decision, in most cases, is not binary — it is a portfolio question about which workloads run where and what governance conditions each environment must satisfy.

    Regardless of the type of organization you are, validated reference designs for private deployment that can scale in a reasonable timeframe with minimal execution risk serve a real purpose. AI is not trivial to engineer independently. Pre-validated stacks lower the barrier for organizations that need private AI the most but have the least tolerance for deployment risk. HPE has made a substantive set of announcements in this space. So have others. Organizations keep getting better solutions to the “How do I build AI?” question nearly every month. Perhaps, the bigger question businesses need to focus on is “What do we need to build?” 

  • Dell Shares AI Advances And New Metrics To Evaluate Infrastructure

    Dell Shares AI Advances And New Metrics To Evaluate Infrastructure

    At Dell Technologies World in Las Vegas, Dell Technologies chairman and CEO Michael Dell made a pointed argument to a room full of enterprise technology leaders: the metrics organizations use to evaluate infrastructure are evolving.

    GPU counts, cloud versus on-premises comparisons, and model benchmarks have dominated the conversation. Michael Dell’s day one keynote introduced two additional measures aimed at tying infrastructure decisions to actual outcomes: time to token, which measures how quickly a system processes a request and returns a usable AI output, and cost per token, which measures how cheaply that output is produced at scale.

    “Time to first token is incredibly important with investments of this scale,” Dell said on stage, noting that the company now has 5,000 enterprise customers running production AI workloads on its Dell AI Factory with NVIDIA platform. The figure represents a significant increase from the program’s launch two years ago.

    NVIDIA founder and CEO Jensen Huang, appearing alongside Dell, described why those metrics have taken on new urgency at both its NVIDIA GTC conference and at Dell Technologies World. Agentic AI systems, which reason, plan, and execute tasks autonomously over extended periods, require anywhere from 100 to 1,000 times more computation than a system simply responding to a query. “What took months now takes weeks, what took weeks now takes days, and what takes days now takes hours,” Huang said, describing the productivity transformation already underway at companies running agentic workflows. The demand implications for infrastructure are substantial.

    OpenAI president Greg Brockman echoed the framing independently on X.com the same day, writing that “tokens are rapidly becoming the universal input for solving problems.” The convergence of infrastructure vendors and model providers on the same metrics within weeks of each other signals a broader shift in how enterprise AI spending will be evaluated.

    The Data Problem Underneath the Infrastructure Problem

    One of Dell’s significant AI product announcements centered on a new data orchestration engine in the Dell AI Data Platform, which the company positioned as the missing layer between enterprise data and production-ready AI agents.

    The data orchestration engine is the platform’s intelligent control center. It indexes billions of unstructured files, builds governed data pipelines, and connects them to the models and agents that need them at speeds designed to make agentic workflows viable. Dell claims the updated platform delivers 12 times faster vector indexing, six times faster data querying, and 19 times faster time to first token compared to prior generations. While these claims still need to be validated, the proposed increase in performance is good news for enterprises looking to scale AI.

    The underlying problem the engine addresses is one most large organizations know well. Enterprises are simultaneously preparing existing data for AI use and reengineering the data infrastructure required to support AI workloads at scale. Those two efforts compete for the same resources and skills simultaneously.

    “If your data is siloed, your agents are blind,” Dell said. The statement is a precise description of why many enterprise AI pilots have not reached production. An agent operating without access to an organization’s proprietary data, internal knowledge bases, and operational systems cannot deliver the business context that makes agentic AI useful.

    Dell also announced GPU-accelerated SQL analytics through the Dell Data Analytics Engine, powered by Starburst, delivering up to six times faster query performance on NVIDIA Blackwell GPUs. Bank of America, which already has a partnership with Starburst, NVIDIA, and Dell, is among the institutions expected to use the capability.

    A Broad Ecosystem Built to Reduce Time to Production

    Dell announced a new Dell AI Ecosystem Program alongside a significant expansion of frontier model partnerships, positioning both as mechanisms for reducing the time between infrastructure procurement and production AI deployment.

    On the model side, Dell announced collaborations bringing several major AI providers on-premises to the Dell AI Factory. Google and Dell are collaborating to run Gemini 3 Flash models via Google Distributed Cloud on Dell PowerEdge XE9780 servers, enabling enterprises to run advanced generative AI workloads in a confidential computing environment that meets data residency and sovereignty requirements. OpenAI’s Codex will connect with the Dell AI Data Platform, giving enterprises a path to deploy agentic coding capabilities against their internal codebases, documentation, and business systems. SpaceXAI’s Grok is available in on-premises or hybrid enterprise deployments. Palantir’s Foundry and AIP platform is coming on-premises with its Ontology layer deployed on Dell ObjectScale and PowerFlex, allowing organizations to connect data sources and automate business workflows within their own environment.

    The Dell Enterprise Hub on Hugging Face gives enterprises on-premises access to a curated collection of open-weight models including MiniMax-M2.7, DeepSeek Pro, DeepSeek-V4, GLM 5.1, and Kimi K2.6, optimized for Dell AI Factory infrastructure.

    The Dell AI Ecosystem Program formalizes the partner relationship by providing software providers with a validated path to certify their solutions on Dell infrastructure. For enterprise buyers, the practical benefit is pre-validated deployment blueprints that automate the configuration of a specific software, service, or model, reducing integration work that has historically extended timelines from procurement to production.

    The Agent Harness: An Evaluation Criterion Enterprises Are Not Yet Asking About

    One of the more technically substantive moments in the keynote came from Huang’s description of how agents actually operate in production. Agents, he explained, do not run directly on the large language model. They run on a harness, a software layer that sits in a secure, governed sandbox. The harness manages the agent’s reasoning loop, controls tool access, determines when to call a large external model and when to use a smaller local model, and handles memory and context across multi-step tasks.

    NVIDIA’s OpenShell, the open-source sandbox, is now supported across the entire Dell AI Factory from deskside workstations through PowerEdge data center servers.  Dell also announced support for NVIDIA AIQ i.0 blueprints, which provide tested foundations for deploying multi-agent workflows.

    For CIOs evaluating AI infrastructure, the harness architecture is a meaningful addition to the evaluation checklist. Infrastructure that does not clearly define how agent harnesses are deployed, governed, and secured leaves a significant operational and security gap, particularly as agents acquire credentials, access enterprise systems, and take autonomous actions at machine speed.

    For example, “You can’t protect what you can’t see, and you can’t manage what you can’t see,” Dell said, framing the security challenge in terms that apply as directly to agents as to human users. An agent with compromised access or misconfigured permissions can propagate errors or security failures across workflows in ways that a single human user cannot.

    Hybrid AI Infrastructure and the Energy Constraint

    Dell’s survey data shows that 67% of AI workloads are already running outside the public cloud, and 88% of organizations are running at least one AI workload on-premises. The company positioned hybrid AI not as a transitional state but as the long-term architecture reality for most large enterprises.

    The new Dell PowerRack, announced Monday, is a fully integrated rack-scale system that combines compute, networking, and storage, engineered and validated as a single unit. It is designed to reduce the integration overhead of assembling AI infrastructure from components while supporting thermal management and power optimization at rack scale.

    Dell also introduced the Dell PowerCool CDU C7000, the first rack-mount cooling distribution unit designed to meet the cooling requirements of the NVIDIA Vera Rubin NVL72 platform, delivering more than 220 kilowatts of cooling capacity in a 4U form factor. A single rack of NVIDIA Rubin GPUs can draw over 130 kilowatts of power, and Dell noted that energy availability is an increasingly real constraint on AI deployment timelines, independent of sustainability considerations.

    For high-volume on-premises workloads, Dell introduced Dell Deskside Agentic AI, pairing high-performance Dell Pro Precision workstations with NVIDIA NemoClaw. The company claims the configuration enables enterprises to break even against public cloud API costs in as little as 3 months, converting variable token costs into a fixed infrastructure investment.

    What Changes for Enterprise Buyers

    The announcements from Dell Technologies World day one collectively continue to move the enterprise AI infrastructure conversation from capability to faster execution. The core questions are how quickly a given infrastructure configuration can reach first token on a production workload, at what cost per token, and with what governance architecture underpinning the agents running on it.

    The organizations best positioned to answer those questions are the ones that have already started rationalizing their data architecture, defined their hybrid workload placement strategy, and begun evaluating how agent harnesses will be secured and governed. The infrastructure improves almost daily, but the execution discipline required to use it remains the variable that separates AI programs that reach production from those that stay in pilot.

    Maribel Lopez is the founder and principal analyst at Lopez Research, a market research and strategy consulting firm specializing in enterprise AI, AI infrastructure, agentic systems, AI governance, and AI-driven customer experience. I version of this article of originally posted on Forbes.com.