惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Google DeepMind News
Google DeepMind News
C
Check Point Blog
J
Java Code Geeks
腾讯CDC
Martin Fowler
Martin Fowler
MongoDB | Blog
MongoDB | Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
博客园 - 三生石上(FineUI控件)
Apple Machine Learning Research
Apple Machine Learning Research
大猫的无限游戏
大猫的无限游戏
Engineering at Meta
Engineering at Meta
罗磊的独立博客
Last Week in AI
Last Week in AI
B
Blog
IT之家
IT之家
S
SegmentFault 最新的问题
D
DataBreaches.Net
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
GbyAI
GbyAI
博客园 - 聂微东
U
Unit 42
有赞技术团队
有赞技术团队
Y
Y Combinator Blog
MyScale Blog
MyScale Blog

Runpod Blog.

New Runpod datacenter now live: AP-IN-1 Track GPU spend across your team with Cost Centers The GPU supply supercycle is here. Here’s what AI builders need to know. Community Spotlight: One-click AI image and video generation on Runpod with SwarmUI | Runpod Blog Community Spotlight: LoRA Pilot Data Prep to Inference Introducing the Runpod Assistant: Manage Your Cloud GPU Resources with Natural Language OpenAI's Parameter Golf: Train the Best Language Model That Fits in 16MB on Runpod LLM inference optimization: techniques that actually reduce latency and cost Pruna P-Video and Vidu Q3 public endpoints now available on Runpod Runpod brand spelling guide Quickstart - Runpod Documentation The AI market looks nothing like the narrative Training StyleGAN3 with Vision-Aided GAN on Runpod KoboldAI – The Other Roleplay Front End, And Why You May Want to Use It How to Connect Cursor to LLM Pods on Runpod for Seamless AI Dev Community Spotlight: How AnonAI Scaled Its Private Chatbot Platform with Runpod Prompt Scheduling with Disco Diffusion on Runpod Runpod's Latest Innovation: Dockerless CLI for Streamlined AI Development Run Your Own AI from Your iPhone Using Runpod Introducing Flash: Run GPU workloads on Runpod Serverless: No Docker required Use Claude Code with your own model on Runpod: No Anthropic account required Avoid Errors by Selecting the Proper Resources for Your Pod What hackers built on Runpod at TreeHacks 2026 Easily Back Up and Restore Your Pod with Cloud Sync + Backblaze B2 The Complete Guide to GPU Requirements for LLM Fine-Tuning AI Guides, Tutorials & GPU Infrastructure Insights | Runpod Your first Claude Code project within Runpod: a complete setup guide 10 billion Serverless requests and counting Building for resilience: Runpod’s response to the AWS us-east-1 outage How to Connect Google Colab to Runpod
Groundbreaking H100 NVidia GPUs Now Available On Runpod
Brendan McKeag · 2023-06-08 · via Runpod Blog.

The demand for generative AI models continues to explode, as does the need for hardware capable of harnessing their ever-escalating performance requirements. While consumer-grade GPUs typically used for gaming are great for learning, tinkering, or hobbyist pursuits, as noted in our previous blog the computational demands can quickly outstrip the abilities of these cards, and enterprise-grade solutions are required to produce high-end solutions. In particular, VRAM in these consumer cards is often a limiting factor, and the cost per gigabyte can quickly spiral out of control which can make completing these tasks untenable. While Runpod already offers access to high-end solutions that are often out of the reach of many consumers such as the RTX 6000, A40, and A100, we are now pleased to also offer the H100, the latest generation in NVidia's large-scale AI GPU solutions. We're delighted to announce the launch of the H100 on Runpod—the most powerful GPU ever created for accelerating AI workloads. This game-changing addition to our lineup will revolutionize the way you leverage AI as an unbelievable tool to drive success in your organization.

Why the H100?

The H100 is already in use or planned to be used by some of the biggest names in AI for their solutions, including OpenAI, Meta, Stability AI, Mitsui, Johns Hopkins, and many others. The trend for AI models is that they are going to become larger and larger for the foreseeable future. For example, most commonly available consumer text generation models are between 6 and 50 billion parameters, and GPT-3 is 175, but GPT-4 is conservatively estimated to be up to 1 trillion. This requires hardware that not only has the raw computational power to shoulder the load, but is also able to do so in a responsible, sustainable manner. This means that not only will the job get done faster, but you'll also pay less per unit of compute in the process. The cost of training AI models continues to rise, and this is only going to be exacerbated by the parameter space growing at the pace that it is – so it's important to not only have solutions that get the job done faster, but also smarter.

NVIDIA chart comparing GPT-3 training on HGX A100 vs HGX H100 at fixed budget and fixed node count

Compare the H100 performance to the A100 currently offered on Runpod; while the pricing for the H100 is to be determined, you should see an increase in performance several times over for what will likely be a nominal price increase in line with previous generational increases.

NVIDIA bar charts showing H100 inference performance versus A100 and other accelerators in MLPerf

The H100 is also specifically optimized for generative AI processes on a level that has not yet been seen before, including large language models, recommendation models, and video and image generation. Depending on the use case,  the H100 can have a performance gain of 7 to 12 times over the A100. Further, there are software optimizations available to only the H100 combined with the new Hopper architecture specifically built for AI inference that will further put it head and shoulders above the previous generations.

NVIDIA chart of H100 inference speedups from MLPerf v2.1 to v3.0, up to 54 percent from software gains

As of the launch of the H100, Runpod will offer the following server loadouts:

Down the road, plans for the H100 NVL are also in motion with an estimated launch by year-end of 2023.

Questions?

Due to anticipated high demand, Runpod is currently planning to offer the H100 on a reservation system. If you are interested in potentially accessing one of these pods, please reach out to us and submit your use case for consideration.

Author profile: Brendan McKeag