惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Cyberwarzone
Cyberwarzone
F
Fortinet All Blogs
Y
Y Combinator Blog
C
Check Point Blog
Latest news
Latest news
A
About on SuperTechFans
Spread Privacy
Spread Privacy
W
WeLiveSecurity
Know Your Adversary
Know Your Adversary
Stack Overflow Blog
Stack Overflow Blog
云风的 BLOG
云风的 BLOG
Recent Announcements
Recent Announcements
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Microsoft Azure Blog
Microsoft Azure Blog
博客园 - 叶小钗
Last Week in AI
Last Week in AI
S
SegmentFault 最新的问题
T
Troy Hunt's Blog
T
Threatpost
Recent Commits to openclaw:main
Recent Commits to openclaw:main
博客园 - 司徒正美
Cloudbric
Cloudbric
J
Java Code Geeks
N
News | PayPal Newsroom
雷峰网
雷峰网
N
News and Events Feed by Topic
罗磊的独立博客
博客园 - 三生石上(FineUI控件)
Recorded Future
Recorded Future
爱范儿
爱范儿
C
Cisco Blogs
P
Proofpoint News Feed
Hacker News: Ask HN
Hacker News: Ask HN
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Hugging Face - Blog
Hugging Face - Blog
O
OpenAI News
大猫的无限游戏
大猫的无限游戏
Webroot Blog
Webroot Blog
T
The Blog of Author Tim Ferriss
宝玉的分享
宝玉的分享
S
Secure Thoughts
博客园 - 【当耐特】
人人都是产品经理
人人都是产品经理
V
Vulnerabilities – Threatpost
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
D
Darknet – Hacking Tools, Hacker News & Cyber Security
SecWiki News
SecWiki News
Martin Fowler
Martin Fowler
阮一峰的网络日志
阮一峰的网络日志
P
Privacy & Cybersecurity Law Blog

Runpod Blog.

DeepSeek V4 in the wild, and how to run it on Runpod New Runpod datacenter now live: AP-IN-1 Track GPU spend across your team with Cost Centers The GPU supply supercycle is here. Here’s what AI builders need to know. Community Spotlight: One-click AI image and video generation on Runpod with SwarmUI | Runpod Blog Community Spotlight: LoRA Pilot Data Prep to Inference Introducing the Runpod Assistant: Manage Your Cloud GPU Resources with Natural Language OpenAI's Parameter Golf: Train the Best Language Model That Fits in 16MB on Runpod LLM inference optimization: techniques that actually reduce latency and cost Pruna P-Video and Vidu Q3 public endpoints now available on Runpod Runpod brand spelling guide Quickstart - Runpod Documentation The AI market looks nothing like the narrative Training StyleGAN3 with Vision-Aided GAN on Runpod KoboldAI – The Other Roleplay Front End, And Why You May Want to Use It How to Connect Cursor to LLM Pods on Runpod for Seamless AI Dev Community Spotlight: How AnonAI Scaled Its Private Chatbot Platform with Runpod Prompt Scheduling with Disco Diffusion on Runpod Runpod's Latest Innovation: Dockerless CLI for Streamlined AI Development Run Your Own AI from Your iPhone Using Runpod Introducing Flash: Run GPU workloads on Runpod Serverless: No Docker required Use Claude Code with your own model on Runpod: No Anthropic account required Avoid Errors by Selecting the Proper Resources for Your Pod What hackers built on Runpod at TreeHacks 2026 Easily Back Up and Restore Your Pod with Cloud Sync + Backblaze B2 The Complete Guide to GPU Requirements for LLM Fine-Tuning AI Guides, Tutorials & GPU Infrastructure Insights | Runpod Your first Claude Code project within Runpod: a complete setup guide 10 billion Serverless requests and counting Building for resilience: Runpod’s response to the AWS us-east-1 outage How to Connect Google Colab to Runpod Founder Series #1: The Runpod Origin Story How to Run the FLUX Image Generator with ComfyUI on Runpod Run Llama 3.1 405B with Ollama on Runpod: Step-by-Step Deployment How to Run FLUX Image Generator with Runpod (No Coding Needed) How to Use 65B+ Language Models on Runpod Deploy Llama 3.1 with vLLM on Runpod Serverless: Fast, Scalable Inference in Minutes Open Source Video & LLM Roundup: The Best of What’s New Run vLLM on Runpod Serverless: Deploy Open Source LLMs in Minutes Introduction to vLLM and PagedAttention New update to Github integration: release rollback! | Runpod Blog A note to the developers who built Runpod with us Deploy ComfyUI as a Serverless API Endpoint Setting up Slurm on Runpod Clusters: A Technical Guide Building an OCR System Using Runpod Serverless From No-Code to Pro: Optimizing Mistral-7B on Runpod for Power Users Lessons While Using Generative Language and Audio For Practical Use Cases Runpod RoundUp 3 – AI Music and Stock Sound Effect Creation New Navigational Changes To Runpod UI Use alpha_value To Blast Through Context Limits in LLaMa-2 Models Runpod Roundup 5 – Visual/Language Comprehension, Code-Focused LLMs, and Bias Detection Runpod is Proud to Sponsor the StockDory Chess Engine Runpod Roundup 4 – Open Source LLM Evaluators, 3D Scene Reconstruction, Vector Search Meta and Microsoft Release Llama 2 as Open Source SuperHot 8k Token Context Models Are Here For Text Generation How to Manage Funding Your Runpod Account Encrypted Volumes on Runpod: Protect Your Data at Rest How to Run a "Hello World" on Runpod Serverless Runpod AI field notes: December 2025 Faster GitHub Builds: Major Performance Improvements to Our Automated Integration Partnering with Defined AI to Bridge the Data Wealth Gap How to Run Serverless AI and ML Workloads on Runpod How to fine-tune a model using Axolotl Transcribe and translate audio files with Faster Whisper Runpod Achieves SOC 2 Type II Certification: Continuing Our Compliance Journey Orchestrating GPU workloads on Runpod with dstack Exploring Runpod Serverless: Create Workers From Templates DeepSeek V3.1: A Technical Analysis of Key Changes from V3-0324 Deep Cogito Releases Suite of LLMs Trained with Iterative Policy Improvement Wan 2.2 Releases With a Plethora Of New Features Iterative Refinement Chains with Small Language Models The New Runpod.io: Clearer, Faster, Built for What’s Next Introducing Clusters: On-Demand Multi-Node AI Compute Run DeepSeek R1 on Just 480GB of VRAM How Do I Transfer Data Into My Runpod? Spot vs. On-Demand Instances: What’s the Difference? Deploy GitHub Repos to Runpod with One Click Run GGUF Quantized Models Easily with KoboldCPP on Runpod How to Work with GGUF Quantizations in KoboldCPP Introducing Better Forge: Spin Up Stable Diffusion Pods Faster Supercharge Your LLMs with SGLang: Boost Performance and Customization Mastering Serverless Scaling on Runpod: Optimize Performance and Reduce Costs RAG vs. Fine-Tuning: Which Is Best for Your LLM? Run Larger LLMs on Runpod Serverless Than Ever Before – Llama-3 70B (and beyond!) How to Run vLLM on Runpod Serverless (Beginner-Friendly Guide) Embracing New Beginnings: Welcoming Banana.dev Community to Runpod Stable Diffusion + ComfyUI on Runpod: Easy Setup Guide Runpod RoundUp 2 – 32k Token Context LLMs and New StabilityAI Offerings Runpod Roundup: High-Context LLMs, SDXL, and Llama 2 16k Context LLM Models Now Available On Runpod Savings Plans Are Here For Secure Cloud Pods – How To Purchase a Monthly Plan And Save Big Pygmalion-7b from PygmalionAI has been released, and it's amazing Ada Architecture Pods Are Here – How Do They Stack Up Against Ampere? Spin up a Text Generation Pod with Vicuna and Experience a GPT-4 Rival Using OpenPose to Annotate Poses Within Stable Diffusion Set Up a Chatbot with Oobabooga on Runpod Connect VSCode to Your Runpod Instance (Quick SSH Guide) Deploy a Stable Diffusion UI on Runpod in Minutes Google Colab Pro vs. Runpod: Best GPU Cloud for AI Workloads How to Run a GPU-Accelerated Virtual Desktop on Runpod
AMD MI300X vs. NVIDIA H100: Mixtral 8x7B Inference Benchmark
Marut Pandya · 2026-02-18 · via Runpod Blog.

There’s no denying Nvidia's historical dominance when it comes to AI training and inference. Nearly all production AI workloads run on their graphics cards.

However, there’s been some optimism recently around AMD, seeing as the MI300X, their intended competitor to Nvidia's H100, is strictly better spec-wise.

Screen capture comparing the specs of AMD MI300X and Nvidia H100 SXM GPUs.

Screen capture comparing the specs of AMD MI300X and Nvidia H100 SXM GPUs.

Yet even with better raw specs, most developers don’t use AMD cards for real-life production workloads, since Nvidia's CUDA is miles ahead of AMD’s ROCm when it comes to writing software for machine learning applications.

To address the growing interest in AMD, we present benchmarks for both AMD’s MI300X and Nvidia's H100 SXM when running inference on MistralAI’s Mixtral 8x7B LLM.

Our benchmarks show that the MI300X performs better than the H100 SXM at small and large batch sizes (1, 2, 4, and 256, 512, 1024), but worse at medium batch sizes.

Benchmark Setup

We chose Mistral AI's Mixtral 7x8B LLM for this benchmark due to its popularity in production workflows and its large size, which doesn't fit on a single Nvidia H100 SXM (80GB VRAM). To handle this, we set tensor parallelism to 2 on the H100, while the MI300X, with its 192GB VRAM, can fit the Mixtral 7x8B model on a single GPU.

However, directly comparing two H100 SXMs against one MI300X wouldn't be fair, so we extrapolated the performance of two MI300Xs working together. For each batch size, we doubled the performance of the previous batch size to estimate the MI300X's performance, simulating the workload distribution of two GPUs. For instance, if handling 16 sequences, each MI300X would separately process a batch of 8 sequences, mirroring the parallel workload distribution of two GPUs.

We ran the same throughput benchmark for each graphics card to test batched offline inference that comes standard as part vLLM's LLM inference framework. All computations were done using FP16 precision with input and output lengths set to 128 tokens.

Understanding Batch Size

Before we dive into the benchmarks, it's important to understand batch size. In the context of LLM inference, batch size refers to the number of prompts processed simultaneously by the model.

Smaller batch sizes can be more manageable and less resource-intensive but may not fully utilize the GPU's capabilities, leading to higher costs per token. They also result in the highest throughput per request but the lowest overall throughput. Conversely, larger batch sizes can leverage the GPU's full potential, increasing overall throughput and reducing costs per token. However, they require more VRAM and computational power and result in a lower throughput per individual request.

Understanding this balance is crucial: smaller batch sizes prioritize quick, individual responses, while larger batch sizes optimize for bulk processing and cost efficiency. Choosing the right batch size depends on your specific use case and resource availability.

Throughput Comparison (tokens/sec)

Chart comparing the throughput (tokens per second) of the AMD MI300X and Nvidia H100 SXM when running inference on Mistral

Chart comparing the throughput (tokens per second) of the AMD MI300X and Nvidia H100 SXM when running inference on Mistral AI's Mixtral 7x8B model.

Raw data used to create the chart comparing throughput of the AMD MI300X and Nvidia H100 SXM. A batch size of 1 can't be

Raw data used to create the chart comparing throughput of the AMD MI300X and Nvidia H100 SXM. A batch size of 1 can't be split between 2 GPUs, so the performance between 1x vs 2x AMD MI300X is the same.

The Nvidia H100 SXM outperforms the AMD MI300X at smaller batch sizes, up to 128. However, as the batch size increases beyond 128, the MI300X starts to show its strengths. At batch sizes of 256 and above, the MI300X catches up and eventually surpasses the H100 SXM in throughput.

This shift suggests that the MI300X's larger VRAM (192GB) becomes more advantageous at higher batch sizes, allowing it to handle larger workloads more efficiently on a single GPU. Despite this, the H100 SXM’s consistent performance at smaller batch sizes highlights its suitability for applications where lower batch sizes are common.

Cost Comparison

Chart comparing the cost per 1 million tokens for running inference on Mistral AI's Mixtral 7x8B model using AMD MI300X and

Chart comparing the cost per 1 million tokens for running inference on Mistral AI's Mixtral 7x8B model using AMD MI300X and Nvidia H100 SXM GPUs.

It's important to note that GPU costs can vary significantly across different cloud providers. For this comparison, we used pricing from Runpod's Secure Cloud, where the H100 SXM is priced at $4.69 per hour and the MI300X at $4.89 per hour.

At smaller batch sizes (1 to 4), the MI300X is more cost-effective than the H100 SXM. For instance, at a batch size of 1, the MI300X costs $22.22 per 1 million tokens, compared to the H100 SXM's $28.11. This cost advantage continues at batch sizes of 2 and 4, with the MI300X maintaining lower costs.

At higher batch sizes (256, 512, and 1024), the MI300X regains its cost advantage, offering lower costs per 1 million tokens compared to the H100 SXM. This indicates that while the MI300X is more expensive at medium batch sizes, it becomes more cost-effective for both very low and very high batch sizes.

Serving Comparison

While throughput and cost per 1M tokens provide valuable insights into GPU performance and cost-efficiency, serving benchmarks offer a more comprehensive view of real-world capabilities.

Serving benchmarks evaluate end-to-end performance, including request throughput, token processing times, and inference latency, which are crucial for understanding user experience and responsiveness. These metrics reflect how GPUs handle different load conditions and batch sizes in production environments.

For this comparison, we measured 1x MI300X instead of 2x to highlight its performance under typical usage scenarios. This approach provides a deeper understanding of operational efficiency, consistency, and reliability beyond raw computational power and cost.

Serving benchmark results on 1x AMD MI300X.

Serving benchmark results on 1x AMD MI300X.

Serving benchmark results on 2x Nvidia H100 SXM.

Serving benchmark results on 2x Nvidia H100 SXM.

Both GPUs have their strengths: the Nvidia H100 SXM offers higher throughput at smaller to medium batch sizes, while the AMD MI300X provides lower latency and better consistency at larger batch sizes. The choice between the two will depend on the specific requirements of the workload, such as the desired balance between throughput and latency.

Conclusion

The MI300X excels at very low and very high batch sizes, offering better cost efficiency and leveraging its larger VRAM to handle larger workloads more effectively. Conversely, the H100 SXM demonstrates superior throughput at smaller to medium batch sizes, making it suitable for applications where these batch sizes are common.

Serving benchmarks reveal that the MI300X has lower latency and delivers consistent performance under higher loads, while the H100 SXM maintains robust throughput and cost-efficiency in mid-range batch sizes.

We’ll be working on more real world tests, specifically benchmarking other popular open source models like Mixtral 8x22b where AMD’s 192 GB of VRAM may be more impactful, and Llama-3 8b which is (arguably) the most popular open-source LLM out today.

If you want to try benchmarking or running AI workloads on the MI300X yourself, you can rent it by the minute on Runpod.

Try AMD MI300X

Replicate our Benchmarks

We encourage you to validate our benchmarks by running your own tests using the code outlined below. By replicating these benchmarks, you can gain a deeper understanding of the performance characteristics of the AMD MI300X and Nvidia H100 SXM in your specific use cases. Whether you're testing throughput, cost-efficiency, or serving benchmarks, following our setup and commands will help you see firsthand how these GPUs perform with MistralAI’s Mixtral 8x7B LLM.

AMD MI300X Benchmarks

Hardware:

  • GPU: 1x MI300x
  • Memory: 192 GB VRAM
  • Tensor Parallelism (TP): 1
  • Cloud Provider: Runpod

Software:

  • vLLM Version: 0.4.3 + ROCm614
  • Repository: ROCm/vllm

Commands to replicate throughput benchmarks:

Commands to replicate serving benchmarks:

Nvidia H100 SXM Benchmarks

Hardware:

  • GPU: H100 SXM
  • Memory: 80 GB VRAM
  • Tensor Parallelism (TP): 2
  • Cloud Provider: Runpod

Software:

Commands to replicate benchmarks:

Author profile: Marut Pandya