STOCK TITAN

QumulusAI Completes Deployment of All 616 NVIDIA RTX PRO 6000 Blackwell GPUs for Runpod, Weeks Ahead of Schedule

QumulusAI (QMLS) has completed deployment of 616 NVIDIA RTX PRO 6000 Blackwell GPUs for Runpod under one- and two-year reserved-capacity agreements, with all 77 GPU nodes now active as of August 2026, ahead of the original September 1, 2026 target.

(Moderate)
(Neutral)
Tags
AI
See more from StockTitan in Google Search and AI answers. Adds StockTitan as a preferred source · opens Google
Add on Google

Full activation across 77 GPU nodes — combining one- and two-year reserved-capacity agreements — was reached in August 2026, ahead of the original Sept. 1, 2026 target

ATLANTA--(BUSINESS WIRE)-- QumulusAI (Nasdaq: QMLS), a neocloud infrastructure provider purpose-built for the AI computing era, today announced that it has completed deployment of 616 NVIDIA RTX PRO 6000 Blackwell GPUs for Runpod under two reserved-capacity agreements combining one- and two-year terms.

All 77 GPU nodes under the agreements are now active. Full deployment was reached in August 2026, ahead of the agreements' original Sept. 1, 2026 target. The GPUs are being served from QumulusAI's existing U.S. data center footprint.

"Teams building on Runpod need inference capacity that lands fast and prices well. QumulusAI got all 616 of these GPUs live ahead of schedule, which means developers get to use them sooner," said Bill Sehmel, Manager of Datacenter Infrastructure at Runpod.

Runpod, a GPU cloud marketplace for AI developers, will offer the RTX PRO 6000 Blackwell capacity to its customers for production AI inference and agentic workloads. RTX PRO 6000 Blackwell GPUs extend QumulusAI's inference- and visualization-class compute alongside its NVIDIA Blackwell B300 and B200 fleet, consistent with the company's approach of matching each workload to the right architecture.

"Bringing all 616 of these RTX PRO 6000 GPUs live for Runpod weeks ahead of schedule is exactly what our demand-led model is built to do," said Ryan DiRocco, CTO of QumulusAI.

About QumulusAI

QumulusAI is a distributed AI cloud platform that delivers accelerated access to high-performance GPU compute. Through an inference-first, demand-led deployment model across a network of data center sites, QumulusAI brings compute closer to customer demand, helping AI teams and enterprises scale production AI workloads with speed, flexibility and control. By combining rapid deployment with flexible private cloud infrastructure, QumulusAI gives customers a faster, more adaptable path beyond the capacity constraints of traditional centralized and hyperscale cloud models. Learn more at QumulusAI.com.

Follow us on LinkedIn and X @QumulusAI.

About Runpod

Runpod is the AI developer cloud. The platform provides the infrastructure AI developers need across the full lifecycle: experiment, train, fine-tune, deploy and scale. Over 1 million developers build on Runpod. Specifically for AI workloads, Runpod is the fastest path from AI experiment to production. For more information, visit runpod.io.

Forward-Looking Statements

This press release contains forward-looking statements within the meaning of Section 27A of the Securities Act of 1933, as amended, and Section 21E of the Securities Exchange Act of 1934, as amended, including statements regarding the company’s reserved-capacity agreement with Runpod for NVIDIA RTX PRO 6000 Blackwell GPUs, the anticipated term of the agreement, the company’s ability to serve the agreement from its existing data center footprint, and Runpod’s offer and use of the capacity to serve its customers for production AI inference and agentic workloads. Words such as “anticipate,” “believe,” “estimate,” “expect,” “guidance,” “intend,” “can,” “may,” “on track,” “plan,” “project,” “target,” “will” and similar expressions are intended to identify forward-looking statements. These statements are based on management's current expectations and assumptions as of the date of this release and are subject to risks and uncertainties that could cause actual results to differ materially, including, among others, the company's dependence on a limited number of large customers; the availability and cost of power, network connectivity and specialized hardware such as graphics processing units; the company's substantial capital requirements and access to financing; competition and rapid technological change in the high-performance computing and AI markets; the company's limited operating history and history of net losses; and those described in the “Risk Factors” section of the company's registration statement on Form S-1, as amended (File No. 333-292514), filed with the U.S. Securities and Exchange Commission (SEC), and the company’s quarterly report on Form 10-Q for the quarter ended June 30, 2026, as amended, as such factors may be updated in the company's subsequent filings with the SEC. QumulusAI undertakes no obligation to update or revise any forward-looking statement, whether as a result of new information, future developments or otherwise, except as required by applicable law.

Investor Contact
investors@qumulusai.com

Media Contact
media@qumulusai.com

Source: QumulusAI

Key Terms

reserved-capacity agreements financial
Contracts where a buyer pays to reserve a portion of a seller’s production, transportation, storage or service capacity for future use, often including minimum payments or "take-or-pay" clauses that guarantee availability even if the buyer does not fully use the reserved capacity. For investors, these agreements matter because they create predictable revenue and cash flow for the seller while also creating utilization metrics and potential fixed-payment obligations that affect a company’s financial stability and capacity-related risk—similar to booking and paying for a hotel room in advance whether you stay or not.
GPU nodes technical
Compute servers or cluster machines that include one or more graphics processing units (GPUs) alongside central processors, built to run highly parallel, compute‑intensive tasks such as machine learning, data analytics, scientific simulations, and video rendering. Think of a GPU node as a high‑performance engine added to a standard server: it speeds up workloads that can be split into many small pieces, so its presence affects a company’s product capabilities, cloud and hardware costs, and the scalability of AI‑related services.
production AI inference technical
The process of running trained artificial intelligence models to generate predictions, classifications, or outputs for real users or business processes on an ongoing basis. It covers the software and hardware systems that deliver model results with required speed, scale, reliability and cost, like a factory line turning raw inputs into finished products. Investors care because production inference affects a company’s operating costs, product performance, user experience and ability to scale AI-driven revenue or efficiency gains.
agentic workloads technical
Agentic workloads are sets of tasks that are carried out by autonomous software agents rather than people, such as routing customer requests, making routine decisions, or executing repeatable processes. For investors, they matter because shifting work from humans to these self-directed programs can cut labor costs, speed operations, and scale services like a factory robot — changes that can boost productivity, alter capital needs, and affect a company’s future earnings potential.