STOCK TITAN

AMD and Cerebras Announce Industry-Leading Ultra-Low-Latency and High Throughput AI Inference Solution

(Moderate)
(Very Positive)
Tags
AI

AMD (NASDAQ: AMD) and Cerebras Systems (NASDAQ: CBRS) announced a technical partnership to create a disaggregated AI inference solution that links AMD Helios rackscale systems and Cerebras Wafer-Scale Engine into a single workflow for ultra-low-latency, high-throughput AI inference.

AMD Helios handles high-throughput prompt and large-context processing, while Cerebras technology focuses on ultra-fast, low-latency decode and token generation. According to the companies, this combined approach is modeled to deliver up to 5x higher tokens per second per watt versus a Cerebras-only configuration. Cerebras plans to deploy AMD Helios in its data centers, with initial access to the joint solution expected via Cerebras Cloud in the second half of 2026.

Loading...
Loading translation...

Positive

  • Modeled 5x efficiency gain in tokens per second per watt versus Cerebras-only setup
  • Joint AMD Helios and Cerebras Wafer-Scale Engine workflow targets ultra-low-latency inference segment
  • Planned deployment of AMD Helios in Cerebras data centers expands AMD’s AI footprint
  • Initial commercial availability targeted via Cerebras Cloud in H2 2026

Negative

  • None.

News Market Reaction – AMD

-2.29%
4 alerts
-2.29% Session close to close
$900.63B Market Cap
0.5x Rel. Volume

In the Jul 23 session, AMD declined 2.29%, reflecting a moderate negative market reaction. Our momentum scanner triggered 4 alerts that day, indicating moderate trading interest and price volatility.

Data tracked by StockTitan Argus on the day of publication.

Market Context

AMD's five AI-tagged historical events averaged a 0.26% move. That record adds mixed precedent to th...
Analysis

AMD's five AI-tagged historical events averaged a 0.26% move. That record adds mixed precedent to this technical partnership. Low short positioning limits one disclosed volatility factor, while recent insider activity was Net Selling.

Key Figures

Tokens per second per watt: up to 5x higher T/s/W Initial availability: second half of 2026 Benchmark model: Kimi 2.6 1T Model +1 more
4 metrics
Tokens per second per watt up to 5x higher T/s/W Joint AMD Helios and Cerebras Wafer-Scale Engine solution
Initial availability second half of 2026 Expected availability through Cerebras Cloud
Benchmark model Kimi 2.6 1T Model Modelling comparison cited in the article footnote
Announcement date July 23, 2026 AMD and Cerebras technical partnership announcement

Previous AI Reports

5 past events · Latest: Jul 22 (Positive)
Same Type Pattern 5 events
Date Event Sentiment 24h Move Catalyst
Jul 22 AI project support Positive +1.4% AMD and Vultr supported Cambridge's planetary-scale environmental monitoring AI project
Jun 09 AI infrastructure architecture Positive -0.5% BlueRock added AMD platform DMA remapping support for isolated AI workloads
Jun 08 AI investment commitment Positive +4.6% AMD committed up to £2 billion over five years to UK AI innovation
Apr 28 AI event announcement Positive -3.4% AMD announced its Advancing AI 2026 event scheduled for July 23
Mar 02 AI processor expansion Positive -0.8% AMD expanded Ryzen AI 400 and Ryzen AI PRO 400 processor portfolios

24h Move is the share-price change in the day after each event; other market factors may also have contributed.

Pattern Detected

AMD's AI-tagged announcements produced mixed historical reactions, with two aligned positive outcomes and three divergences.

Key Terms

disaggregated inference, tokens per second per watt, rackscale solution
3 terms
disaggregated inference technical
"The AMD and Cerebras solution addresses this challenge through disaggregated inference"
Disaggregated inference is the practice of breaking down data-driven conclusions into smaller groups or categories—such as customer segments, regions, age bands, or patient subgroups—rather than treating everyone as one average. For investors, it matters because it can reveal hidden strengths, weaknesses, or risks that aggregate numbers hide, much like checking individual bulbs on a string of lights instead of assuming the whole string works; this supports clearer due diligence, pricing, and regulatory assessment.
tokens per second per watt technical
"up to 5x higher tokens per second per watt (T/s/W)"
A metric that measures how many text units (tokens) a machine learning system can process each second for every watt of electrical power it consumes. Think of it like how many pages a printer can print per minute for each gallon of fuel: it captures processing speed divided by energy use, so it helps compare the energy efficiency and operating cost of different AI models or hardware setups.
rackscale solution technical
"AMD Helios™ rackscale solutions with the Cerebras Wafer-Scale Engine"
A rackscale solution is a pre-integrated set of data center hardware and software packaged to fit and operate within a single equipment rack, typically combining servers, storage, networking, and management tools as one coordinated unit. It matters to investors because it can lower deployment time and operational complexity—like buying a prebuilt kitchen instead of assembling appliances separately—affecting capital spending, scalability, and how quickly a business can expand or update its IT capacity.

AI-generated analysis. How Rhea-AI works. Not financial advice.

See more from StockTitan in Google Search and AI answers. Adds StockTitan as a preferred source · opens Google
Add on Google

News Highlights 

  • AMD and Cerebras are collaborating to advance a workload-optimized approach to ultra-low-latency AI inference infrastructure.  
  • AMD Helios™ and the Cerebras Wafer-Scale Engine will operate as a single disaggregated inference workflow, combining ultra-high-throughput from AMD Instinct™ GPUs, with ultra-fast token generation of Cerebras Wafer-Scale Engine.
  • Cerebras plans to deploy AMD Helios in its data centers, with the joint solution expected to be available first through Cerebras Cloud in the second half of 2026. 

SAN FRANCISCO and SUNNYVALE, Calif., July 23, 2026 (GLOBE NEWSWIRE) -- AMD (NASDAQ: AMD) and Cerebras Systems (NASDAQ: CBRS) announced a technical partnership to deliver a new disaggregated AI inference solution that combines AMD Helios™ rackscale solutions with the Cerebras Wafer-Scale Engine. Unveiled at Advancing AI 2026, the solution is designed to deliver the ultra-low latency required for the most advanced AI applications while dramatically increasing the throughput and efficiency. 

The joint AMD and Cerebras solution will deploy AMD Helios alongside Cerebras Wafer-Scale Engine technology integrated in a single inference workflow for maximum performance and efficiency. AMD Helios will provide a high-performance, scalable throughput engine. Cerebras Wafer-Scale Engine technology will provide ultra-fast, ultra-low latency decode and token generation. Together, the two compute engines are expected to deliver up to 5x higher tokens per second per watt (T/s/W) i

AI inference workloads increasingly have different requirements across latency, throughput, token capacity, cost and scale. High-volume workloads prioritize maximizing token generation, while coding, real-time copilots, live agents and agentic workflows demand faster response times. These differences are driving demand for heterogeneous infrastructure that matches compute technologies to specific workload requirements.  

The AMD and Cerebras solution addresses this challenge through disaggregated inference, optimizing the two primary stages of the workflow independently. AMD Helios provides ultra-high throughput, processing prompts and large context windows. The Cerebras Wafer-Scale Engine accelerates the memory-bandwidth-intensive token generation, with ultra-low latency. By connecting these best-in-class engines through one integrated workflow, the companies are creating a differentiated platform for ultra-low-latency inference without sacrificing throughput or scale. 

“AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach,” said Dr. Lisa Su, chair and CEO, AMD. “AMD Helios delivers leadership performance and scale for the broadest range of inference workloads. Together with Cerebras, we are extending that leadership into the most latency-sensitive applications and creating a powerful new platform for real-time agentic AI.”   

 “The demand for ultra-fast inference is growing at an unprecedented pace. Cerebras delivers the world’s fastest, ultra-low-latency inference,” said Andrew Feldman, CEO and co-founder, Cerebras. “Partnering with AMD gives us an incredible opportunity to bring that performance to even more customers.” 

Fast token generation is becoming increasingly important as AI moves into software development, autonomous agents, robotics, scientific discovery and other applications where response time directly shapes the user experience and the usefulness of the system. The joint solution brings together complementary architectures purpose-built for these demands.  

AMD Helios provides the high-throughput prompt engine, rack-scale efficiency and deployment scale required to process large numbers of complex requests. Cerebras Wafer-Scale Engine technology provides the ultra-low-latency and decode performance needed to return tokens in real time. The result is a solution designed specifically for the ultra-low-latency segment of the inference market, with AMD Helios as the foundation for high-throughput and balanced inference workloads across the data center. 

Cerebras plans to deploy AMD Helios systems in its data centers, with the joint solution expected to become available initially through Cerebras Cloud in the second half of 2026.  

Supporting Resources

About AMD
AMD (NASDAQ: AMD) drives innovation in high-performance and AI computing to solve the world’s most important challenges. Today, AMD technology powers billions of experiences across cloud and AI infrastructure, embedded systems, AI PCs and gaming. With a broad portfolio of AI-optimized CPUs, GPUs, networking and software, AMD delivers full-stack AI solutions that provide the performance and scalability needed for a new era of intelligent computing. Learn more at www.amd.com.

About Cerebras Systems

Cerebras Systems (NASDAQ: CBRS) builds the world’s fastest AI infrastructure. The Cerebras team of pioneering computer architects, computer scientists, AI researchers, and engineers of all types came together to make AI blisteringly fast through innovation and invention. We believe that when AI is fast, it will change the world. Leading global corporations, research institutes, and governments choose Cerebras to run their AI workloads. Cerebras solutions are available on premises and in the cloud. Visit cerebras.ai for more.

AMD CAUTIONARY STATEMENT

This press release contains forward-looking statements concerning Advanced Micro Devices, Inc. (AMD) such as the features, functionality, performance, availability, scalability, deployment, timing and expected benefits of AMD’s collaboration and joint solution with Cerebras, which are made pursuant to the Safe Harbor provisions of the Private Securities Litigation Reform Act of 1995. Forward-looking statements are commonly identified by words such as "would," "may," "expects," "believes," "plans," "intends," "projects" and other terms with similar meaning. Investors are cautioned that the forward-looking statements in this press release are based on current beliefs, assumptions and expectations, speak only as of the date of this press release and involve risks and uncertainties that could cause actual results to differ materially from current expectations. Such statements are subject to certain known and unknown risks and uncertainties, many of which are difficult to predict and are generally beyond AMD's control, that could cause actual results and other future events to differ materially from those expressed in, or implied or projected by, the forward-looking information and statements. Material factors that could cause actual results to differ materially from current expectations include, without limitation, the following: impact of government actions and regulations such as export regulations, import tariffs, trade protection measures, and licensing requirements; competitive markets in which AMD’s products are sold; the cyclical nature of the semiconductor industry; market conditions of the industries in which AMD products are sold; AMD’s ability to introduce products on a timely basis with expected features and performance levels; loss of a significant customer; economic and market uncertainty; quarterly and seasonal sales patterns; AMD's ability to adequately protect its technology or other intellectual property; unfavorable currency exchange rate fluctuations; ability of third party manufacturers to manufacture AMD's products on a timely basis in sufficient quantities and using competitive technologies; availability of essential equipment, materials, components (such as memory supply), substrates or manufacturing processes; ability to achieve expected manufacturing yields for AMD’s products; AMD's ability to generate revenue from its semi-custom SoC products; potential security vulnerabilities; potential security incidents including IT outages, data loss, data breaches and cyberattacks; uncertainties involving the ordering and shipment of AMD’s products; AMD’s reliance on third-party intellectual property to design and introduce new products; AMD's reliance on third-party companies for design, manufacture and supply of motherboards, software, memory and other computer platform components; AMD's reliance on Microsoft and other software vendors' support to design and develop software to run on AMD’s products; AMD’s reliance on third-party distributors and add-in-board partners; impact of modification or interruption of AMD’s internal business processes and information systems; compatibility of AMD’s products with some or all industry-standard software and hardware; costs related to defective products; failure to maintain an efficient supply chain as customer demand changes; AMD's ability to rely on third party supply-chain logistics functions; AMD’s ability to effectively control sales of its products on the gray market; impact of climate change on AMD’s business; AMD’s ability to realize its deferred tax assets; potential tax liabilities; current and future claims and litigation; impact of environmental laws, conflict minerals related provisions and other laws or regulations; evolving expectations from governments, investors, customers and other stakeholders regarding corporate responsibility matters; issues related to the responsible use of AI; restrictions imposed by agreements governing AMD’s notes, the guarantees of Xilinx’s notes and the revolving credit agreement; AMD’s ability to satisfy financial obligations under guarantees, leases and other commercial commitments; impact of acquisitions, joint ventures and/or investments on AMD’s business and AMD’s ability to integrate acquired businesses; impact of any impairment of the combined company’s assets; political, legal and economic risks and natural disasters; future impairments of technology license purchases; AMD’s ability to attract and retain key employees; and AMD’s stock price volatility. Investors are urged to review in detail the risks and uncertainties in AMD’s Securities and Exchange Commission filings, including but not limited to AMD’s most recent reports on Forms 10-K and 10-Q.   

CEREBRAS DISCLOSURE INFORMATION

Cerebras uses its investor relations page (investors.cerebras.ai), its X account (@cerebras), and its LinkedIn page (linkedin.com/company/cerebras-systems/) to disclose material non-public information and for complying with its disclosure obligations under Regulation FD. Accordingly, investors should monitor these channels, in addition to following Cerebras’ press releases, Securities and Exchange Commission (SEC) filings, public conference calls and public webcasts.

Forward-Looking Statements

This press release contains “forward-looking statements” within the meaning of applicable securities laws. All statements other than statements of historical fact could be deemed to be forward-looking, including, but not limited to, statements regarding the features, capacity, scalability, performance, timing, data center deployment and implementation, costs and expected benefits and opportunities associated with Cerebras' collaboration and joint solution with AMD, and any assumptions relating to the foregoing. The words “may,” “will,” “shall,” “should,” “expects,” “plans,” “anticipates,” “could,” “intends,” “target,” “projects,” “contemplates,” “believes,” “estimates,” “predicts,” “potential,” “objective,” or “continue,” or the negative of these words or other similar terms or expressions that concern our expectations, strategy, plans, or intentions are intended to identify forward-looking statements, although not all forward-looking statements contain these identifying words. These forward-looking statements are subject to a number of risks and uncertainties, many of which involve factors or circumstances that are beyond Cerebras’ control. These risks and uncertainties include, but are not limited to: Cerebras’ ability to sustain and manage its growth, access borrowings and other sources of capital on acceptable terms, and deploy available capital to support growth; its history of net losses and ability to achieve and maintain profitability; its limited operating history at its current scale and ability to accurately forecast revenue and appropriately budget and manage expenses; its dependence on a limited number of significant customers, including OpenAI, Group 42 Holding Ltd, Mohamed bin Zayed University of Artificial Intelligence, and AWS, and the potential impact of any reduction in demand from, material adverse development in its relationships with, or failure to meet its obligations to, such customers, including under its Master Relationship Agreement with OpenAI; the timing, execution and expected benefits of its strategic customer, partner and financing arrangements; its historical reliance on sales of hardware systems and the early-stage, rapidly evolving market for its cloud-based offerings and AI infrastructure; its ability to secure sufficient data center capacity and capital to support its cloud-based offerings; its ability to launch new offerings and add new product capabilities; and its ability to compete effectively in the rapidly evolving and competitive market for AI computing solutions.

Cerebras’ actual results could differ materially from those stated or implied in forward-looking statements due to a number of factors. Accordingly, undue reliance should not be placed on such statements. These forward-looking statements are made as of the date they were first issued and are based on information available to Cerebras together with Cerebras’ expectations, estimates, forecasts, projections, beliefs, and assumptions as of such date. These forward-looking statements should not be relied upon as representing Cerebras’ views as of any date subsequent to the date of this press release. Past performance is not necessarily indicative of future results. Cerebras undertakes no intention or obligation to update or revise any forward-looking statements, whether as a result of new information, future events, or otherwise, except as required by law.

Further information on potential risks that could affect actual results is included in Cerebras’ most recent filings with the Securities and Exchange Commission (the “SEC”), including in Cerebras’ most recent Quarterly Report on Form 10-Q, copies of which may be obtained by visiting Cerebras’ Investor Relations website at investors.cerebras.ai or the SEC’s website at www.sec.gov.

Contacts:
Aaron Grabein
AMD Communications
737-256-9518
aaron.grabein@amd.com

Liz Stine
AMD Investor Relations
720-652-3965
liz.stine@amd.com

Kriselle Laran
Cerebras
pr@cerebras.ai

Sean Dorsey
Cerebras Investor Relations
investors@cerebras.ai

_______________
i Based on modelling by AMD Performance Labs and Cerebras in July 2026 to determine tokens per second per kilowatt (TPS/kW) at a comparable interactivity point with Kimi 2.6 1T Model comparing an AMD Helios rackscale solution with Cerebras WSE to a Cerebras WSE-only configuration. System manufacturers may vary configurations, yielding different results. MI400-021


FAQ

What did AMD (AMD) and Cerebras announce on July 23, 2026 about AI inference?

AMD and Cerebras announced a technical partnership to build a disaggregated AI inference solution combining AMD Helios and the Cerebras Wafer-Scale Engine. According to AMD and Cerebras, this integrated workflow targets ultra-low-latency, high-throughput inference for advanced applications like copilots, agents and coding workloads.

How will the AMD Helios and Cerebras Wafer-Scale Engine solution improve AI inference performance for AMD (AMD)?

The joint solution is modeled to deliver up to 5x higher tokens per second per watt versus a Cerebras-only configuration. According to AMD and Cerebras, AMD Helios handles high-throughput prompt processing while the Wafer-Scale Engine accelerates ultra-low-latency token generation in a single workflow.

When will the AMD and Cerebras ultra-low-latency inference solution be available to customers?

The combined AMD Helios and Cerebras Wafer-Scale Engine solution is expected to be available first through Cerebras Cloud in the second half of 2026. According to Cerebras, AMD Helios systems are planned for deployment in its data centers supporting this rollout.

What role does AMD Helios play in the new AMD (AMD) and Cerebras AI infrastructure?

AMD Helios acts as a high-throughput prompt engine, processing prompts and large context windows in the disaggregated inference workflow. According to AMD, Helios provides rack-scale efficiency and deployment scale while Cerebras technology focuses on ultra-low-latency decode and token generation for real-time responses.

How does the AMD and Cerebras solution address ultra-low-latency AI use cases?

The solution separates prompt processing and token generation into optimized stages using AMD Helios and Cerebras Wafer-Scale Engine. According to AMD and Cerebras, this disaggregated design targets latency-sensitive workloads such as real-time copilots, live agents, software development and robotics while maintaining high throughput.

What metric supports the efficiency claims of the AMD and Cerebras AI inference partnership?

AMD and Cerebras cite modeled gains of up to 5x tokens per second per watt compared with a Cerebras-only configuration. According to both companies, this tokens-per-second-per-watt metric was modeled on an AMD Helios rackscale solution with Cerebras WSE at a comparable interactivity point.