惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
量子位
大猫的无限游戏
大猫的无限游戏
Hugging Face - Blog
Hugging Face - Blog
S
SegmentFault 最新的问题
Blog — PlanetScale
Blog — PlanetScale
月光博客
月光博客
Google DeepMind News
Google DeepMind News
小众软件
小众软件
WordPress大学
WordPress大学
宝玉的分享
宝玉的分享
MongoDB | Blog
MongoDB | Blog
B
Blog RSS Feed
博客园 - Franky
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
B
Blog
博客园 - 聂微东
The GitHub Blog
The GitHub Blog
Recent Announcements
Recent Announcements
Y
Y Combinator Blog
Microsoft Security Blog
Microsoft Security Blog
雷峰网
雷峰网
Jina AI
Jina AI
酷 壳 – CoolShell
酷 壳 – CoolShell

Runpod Blog.

DeepSeek V4 in the wild, and how to run it on Runpod New Runpod datacenter now live: AP-IN-1 Track GPU spend across your team with Cost Centers The GPU supply supercycle is here. Here’s what AI builders need to know. Community Spotlight: One-click AI image and video generation on Runpod with SwarmUI | Runpod Blog Community Spotlight: LoRA Pilot Data Prep to Inference Introducing the Runpod Assistant: Manage Your Cloud GPU Resources with Natural Language OpenAI's Parameter Golf: Train the Best Language Model That Fits in 16MB on Runpod LLM inference optimization: techniques that actually reduce latency and cost Pruna P-Video and Vidu Q3 public endpoints now available on Runpod Runpod brand spelling guide Quickstart - Runpod Documentation The AI market looks nothing like the narrative Training StyleGAN3 with Vision-Aided GAN on Runpod KoboldAI – The Other Roleplay Front End, And Why You May Want to Use It How to Connect Cursor to LLM Pods on Runpod for Seamless AI Dev Community Spotlight: How AnonAI Scaled Its Private Chatbot Platform with Runpod Prompt Scheduling with Disco Diffusion on Runpod Runpod's Latest Innovation: Dockerless CLI for Streamlined AI Development Run Your Own AI from Your iPhone Using Runpod Introducing Flash: Run GPU workloads on Runpod Serverless: No Docker required Use Claude Code with your own model on Runpod: No Anthropic account required Avoid Errors by Selecting the Proper Resources for Your Pod What hackers built on Runpod at TreeHacks 2026 Easily Back Up and Restore Your Pod with Cloud Sync + Backblaze B2 The Complete Guide to GPU Requirements for LLM Fine-Tuning AI Guides, Tutorials & GPU Infrastructure Insights | Runpod Your first Claude Code project within Runpod: a complete setup guide 10 billion Serverless requests and counting Building for resilience: Runpod’s response to the AWS us-east-1 outage
Runpod Roundup: High-Context LLMs, SDXL, and Llama 2
Brendan McKeag · 2023-07-22 · via Runpod Blog.

Welcome to the Runpod Roundup! In this week we'll be discussing new text and image generation models, including an exciting new Stable Diffusion model.

The goal of Runpod Roundup is to keep you abreast of new developments over the week that you might have missed, with a focus on new models and offerings and  other actionable developments that you can run in a Runpod instance right this  moment.

High Context LLM Models Now Available - 8k Through 16k Tokens

After a very long time of languishing in the sufficient-but-not-optimal area of 2k tokens, LLM models are finally breaking through this barrier. Earlier in the month, we saw SuperHot 8k context models from TheBloke and this week we saw 16k models from Panchovix. Despite higher VRAM requirements, all of these models fit comfortably within Runpod instances (though if you were already on the borderline for your chosen GPU/model combo, you may need to bump it up to the next higher level of GPU.)

These larger context options could not have come at a better time, as many applications of LLMs have been butting up against this limit for quite some time– for example, a roleplay where the AI plays multiple characters can easily run into hundreds of tokens for a single round of responses which can fill up the window quickly. Give the new models a shot and let us know what you think!

Stable Diffusion XL (SDXL) Released For All

After spending some time percolating in beta, SDXL is now available for anyone to download. According to the creators, this version corrects some long-standing (and often parodied) problems with the original model, such as human anatomy and text. It does not appear that SDXL is compatible with Automatic1111 yet, but fear not - we'll have a how-to article coming out shortly on how to get it set up in a Runpod instance through alternate means. For now, though, you can grab the model from the StabilityAI HuggingFace site.

AI-generated image of a modern wooden house on a snowy mountainside above the clouds at sunrise

Meta and Microsoft Release their Llama 2 Open Source LLM Model

Llama 2 is now available for download, with 7b, 13b, and 70b parameter size available. The model boasts a 4k contest length and has been built with dialogue in mind using Reinforcement Learning from Human Feedback. According to human evaluators, the model performs comparably to ChatGPT and you can run it right in your own Runpod pod. The HF site advises that you may need an A100 just for the 13B model, so be aware of the heightened resource requirements compared to other commonly available open-source LLM models.

Note: You'll need to first request access to Llama 2 from Meta through the Llama access page and have it approved before you'll be able to download it from Huggingface, and the turnaround may be a few days.

Questions?

Feel free to reach out to Runpod directly if you have any questions about these latest developments, and we'll see what we can do for you!

Author profile: Brendan McKeag