Hippocratic AI Scales to 10 Million Patient Calls at 99.9% Clinical Safety on DigitalOcean's AI-Native Cloud, powered by NVIDIA Blackwell Ultra GPUs
Key Terms
inference technical
latency technical
clinical safety medical
quantization technical
kv-cache technical
p99 latency technical
DigitalOcean’s AI-Native Cloud, optimized from infrastructure to inference, delivers 2× prefill speedup and ~
Hippocratic AI's Polaris system has reported a
"Polaris is built for the realities of clinical care: long sessions, real human conversations, zero room for error. With DigitalOcean and NVIDIA, we have early access to NVIDIA HGX™ B300 and the optimization techniques it unlocks, including NVFP4 quantization,” said Debajyoti Datta, Co-Founder, Hippocratic AI. “That is what allows us to hold a 400-millisecond time-to-first-token at production scale, on the clinical conversations our patients depend on."
Engineered to Support Safety-Critical Inference
Production healthcare AI breaks the assumptions most inference stacks are built on. Sessions are long. Tokens are time-sensitive. A dropped connection in the middle of a care plan retrieval is not a UX bug. It is a clinical interruption. Meeting that bar requires deep platform engineering and reliability at scale, the kind that off-the-shelf GPU access cannot provide and that only a purpose-built inference cloud can deliver.
Over the past year, the engineering teams at DigitalOcean worked in close collaboration with Hippocratic AI to optimize every layer of the inference stack. DigitalOcean engineered its AI-Native Cloud with hardware-aware scheduling, optimized inference runtimes, and platform-level scaling tuned for sustained high-concurrency workloads. Hippocratic AI's model team contributed proprietary inference work, including FP8 and NVFP4 quantization, KV-cache optimization, custom MoE kernels, and a cache-aware routing architecture that maximizes KV-cache hit rate and context reuse across long-horizon clinical sessions. NVIDIA provided early access to next-generation HGX™ B300 hardware, alongside engineering collaboration on Hopper and Blackwell architecture.
The combined result, on long-context clinical sessions, is approximately
"What Hippocratic AI has built in healthcare AI is remarkable, hundreds of millions of real patient interactions across some of the most complex and sensitive moments in people's lives,” said Paddy Srinivasan, Chief Executive Officer, DigitalOcean. “Delivering that at
Among the First Production Customers on NVIDIA HGX™ B300
Having Hippocratic AI among the first production customers on NVIDIA HGX™ B300 GPUs, made available through DigitalOcean's early work with NVIDIA, means DigitalOcean is validating its inference platform against one of the most demanding real-world workloads, not synthetic benchmarks. For workloads where every token affects clinical experience, Blackwell Ultra unlocks a step-change in capacity per node, allowing Hippocratic AI to support more concurrent sessions at the same latency targets and to extend context windows on long-horizon clinical conversations.
"The demands of safety-critical AI workloads are fundamentally different from consumer applications,” said Dave Salvator, Director of Accelerated Computing Products, NVIDIA. “DigitalOcean and Hippocratic AI are demonstrating how tightly integrated infrastructure and inference optimization, built on NVIDIA Hopper and Blackwell architecture, can deliver both performance and reliability at scale."
A Different Bar for Healthcare AI Infrastructure
The infrastructure requirements of safety-critical AI are not the requirements of consumer or enterprise AI scaled up. They are different in kind. Latency translates directly into clinical workflow quality. Reliability is measured in successful patient interactions, not nine-fives uptime. Cost efficiency determines whether a healthcare AI workload can scale to serve a population, not just a pilot.
In healthcare AI, infrastructure is not just about performance. It is foundational to patient safety. The Hippocratic AI deployment on the DigitalOcean AI-Native Cloud reflects this shift, and the platform engineering behind it shows what production AI looks like when infrastructure, model optimization, and hardware are designed together for outcomes that matter.
Read the full customer case study, including a video interview with Hippocratic AI Co-Founder Debajyoti Datta, at digitalocean.com/customers/hippocratic-ai.
About DigitalOcean
DigitalOcean is the AI-Native Cloud purpose-built for the inference and agentic era. Its five-layer integrated platform - spanning infrastructure, core cloud, inference, data, and managed agents - is open throughout with no vendor lock-in, giving builders everything they need to start fast, scale production AI workloads, and improve unit economics. More than 650,000 customers globally trust DigitalOcean to build, ship, and scale their applications. Learn more at digitalocean.com.
About Hippocratic AI
Hippocratic AI has developed the safest generative AI Agents for healthcare. The company believes that generative AI has the ability to bring healthcare abundance to every person in the world. The company focuses on building non-diagnostic patient-facing clinical AI agents and does not allow its agents to be used to prescribe or diagnose. Hippocratic AI has received a total of
View source version on businesswire.com: https://www.businesswire.com/news/home/20260527711308/en/
Media Relations
Meghan Grady
press@digitalocean.com
Investor Relations
Radu Patrichi, CFA
investors@digitalocean.com
Source: DigitalOcean