惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Google DeepMind News
Google DeepMind News
人人都是产品经理
人人都是产品经理
H
Hacker News: Front Page
Stack Overflow Blog
Stack Overflow Blog
B
Blog
I
InfoQ
GbyAI
GbyAI
T
The Blog of Author Tim Ferriss
F
Fortinet All Blogs
Y
Y Combinator Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
月光博客
月光博客
Hugging Face - Blog
Hugging Face - Blog
爱范儿
爱范儿
F
Full Disclosure
Hacker News - Newest:
Hacker News - Newest: "LLM"
Recent Announcements
Recent Announcements
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Jina AI
Jina AI
T
Tailwind CSS Blog
S
Secure Thoughts
P
Privacy International News Feed
美团技术团队
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
L
LINUX DO - 最新话题
H
Hackread – Cybersecurity News, Data Breaches, AI and More
C
Cybersecurity and Infrastructure Security Agency CISA
Last Week in AI
Last Week in AI
W
WeLiveSecurity
Google Online Security Blog
Google Online Security Blog
P
Privacy & Cybersecurity Law Blog
D
DataBreaches.Net
Engineering at Meta
Engineering at Meta
Know Your Adversary
Know Your Adversary
P
Palo Alto Networks Blog
I
Intezer
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Project Zero
Project Zero
V2EX - 技术
V2EX - 技术
H
Heimdal Security Blog
博客园 - Franky
阮一峰的网络日志
阮一峰的网络日志
D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Troy Hunt's Blog
V
Vulnerabilities – Threatpost
H
Help Net Security
Martin Fowler
Martin Fowler
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
G
GRAHAM CLULEY
博客园 - 【当耐特】

Runpod Blog.

DeepSeek V4 in the wild, and how to run it on Runpod New Runpod datacenter now live: AP-IN-1 Track GPU spend across your team with Cost Centers The GPU supply supercycle is here. Here’s what AI builders need to know. Community Spotlight: One-click AI image and video generation on Runpod with SwarmUI | Runpod Blog Community Spotlight: LoRA Pilot Data Prep to Inference Introducing the Runpod Assistant: Manage Your Cloud GPU Resources with Natural Language OpenAI's Parameter Golf: Train the Best Language Model That Fits in 16MB on Runpod LLM inference optimization: techniques that actually reduce latency and cost Pruna P-Video and Vidu Q3 public endpoints now available on Runpod Runpod brand spelling guide Quickstart - Runpod Documentation The AI market looks nothing like the narrative Training StyleGAN3 with Vision-Aided GAN on Runpod KoboldAI – The Other Roleplay Front End, And Why You May Want to Use It How to Connect Cursor to LLM Pods on Runpod for Seamless AI Dev Community Spotlight: How AnonAI Scaled Its Private Chatbot Platform with Runpod Prompt Scheduling with Disco Diffusion on Runpod Runpod's Latest Innovation: Dockerless CLI for Streamlined AI Development Run Your Own AI from Your iPhone Using Runpod Introducing Flash: Run GPU workloads on Runpod Serverless: No Docker required Use Claude Code with your own model on Runpod: No Anthropic account required Avoid Errors by Selecting the Proper Resources for Your Pod What hackers built on Runpod at TreeHacks 2026 Easily Back Up and Restore Your Pod with Cloud Sync + Backblaze B2 The Complete Guide to GPU Requirements for LLM Fine-Tuning AI Guides, Tutorials & GPU Infrastructure Insights | Runpod Your first Claude Code project within Runpod: a complete setup guide 10 billion Serverless requests and counting Building for resilience: Runpod’s response to the AWS us-east-1 outage How to Connect Google Colab to Runpod Founder Series #1: The Runpod Origin Story AMD MI300X vs. NVIDIA H100: Mixtral 8x7B Inference Benchmark How to Run the FLUX Image Generator with ComfyUI on Runpod Run Llama 3.1 405B with Ollama on Runpod: Step-by-Step Deployment How to Run FLUX Image Generator with Runpod (No Coding Needed) How to Use 65B+ Language Models on Runpod Deploy Llama 3.1 with vLLM on Runpod Serverless: Fast, Scalable Inference in Minutes Open Source Video & LLM Roundup: The Best of What’s New Run vLLM on Runpod Serverless: Deploy Open Source LLMs in Minutes Introduction to vLLM and PagedAttention New update to Github integration: release rollback! | Runpod Blog A note to the developers who built Runpod with us Deploy ComfyUI as a Serverless API Endpoint Setting up Slurm on Runpod Clusters: A Technical Guide Building an OCR System Using Runpod Serverless From No-Code to Pro: Optimizing Mistral-7B on Runpod for Power Users Lessons While Using Generative Language and Audio For Practical Use Cases Runpod RoundUp 3 – AI Music and Stock Sound Effect Creation New Navigational Changes To Runpod UI Runpod Roundup 5 – Visual/Language Comprehension, Code-Focused LLMs, and Bias Detection Runpod is Proud to Sponsor the StockDory Chess Engine Runpod Roundup 4 – Open Source LLM Evaluators, 3D Scene Reconstruction, Vector Search Meta and Microsoft Release Llama 2 as Open Source SuperHot 8k Token Context Models Are Here For Text Generation How to Manage Funding Your Runpod Account Encrypted Volumes on Runpod: Protect Your Data at Rest How to Run a "Hello World" on Runpod Serverless Runpod AI field notes: December 2025 Faster GitHub Builds: Major Performance Improvements to Our Automated Integration Partnering with Defined AI to Bridge the Data Wealth Gap How to Run Serverless AI and ML Workloads on Runpod How to fine-tune a model using Axolotl Transcribe and translate audio files with Faster Whisper Runpod Achieves SOC 2 Type II Certification: Continuing Our Compliance Journey Orchestrating GPU workloads on Runpod with dstack Exploring Runpod Serverless: Create Workers From Templates DeepSeek V3.1: A Technical Analysis of Key Changes from V3-0324 Deep Cogito Releases Suite of LLMs Trained with Iterative Policy Improvement Wan 2.2 Releases With a Plethora Of New Features Iterative Refinement Chains with Small Language Models The New Runpod.io: Clearer, Faster, Built for What’s Next Introducing Clusters: On-Demand Multi-Node AI Compute Run DeepSeek R1 on Just 480GB of VRAM How Do I Transfer Data Into My Runpod? Spot vs. On-Demand Instances: What’s the Difference? Deploy GitHub Repos to Runpod with One Click Run GGUF Quantized Models Easily with KoboldCPP on Runpod How to Work with GGUF Quantizations in KoboldCPP Introducing Better Forge: Spin Up Stable Diffusion Pods Faster Supercharge Your LLMs with SGLang: Boost Performance and Customization Mastering Serverless Scaling on Runpod: Optimize Performance and Reduce Costs RAG vs. Fine-Tuning: Which Is Best for Your LLM? Run Larger LLMs on Runpod Serverless Than Ever Before – Llama-3 70B (and beyond!) How to Run vLLM on Runpod Serverless (Beginner-Friendly Guide) Embracing New Beginnings: Welcoming Banana.dev Community to Runpod Stable Diffusion + ComfyUI on Runpod: Easy Setup Guide Runpod RoundUp 2 – 32k Token Context LLMs and New StabilityAI Offerings Runpod Roundup: High-Context LLMs, SDXL, and Llama 2 16k Context LLM Models Now Available On Runpod Savings Plans Are Here For Secure Cloud Pods – How To Purchase a Monthly Plan And Save Big Pygmalion-7b from PygmalionAI has been released, and it's amazing Ada Architecture Pods Are Here – How Do They Stack Up Against Ampere? Spin up a Text Generation Pod with Vicuna and Experience a GPT-4 Rival Using OpenPose to Annotate Poses Within Stable Diffusion Set Up a Chatbot with Oobabooga on Runpod Connect VSCode to Your Runpod Instance (Quick SSH Guide) Deploy a Stable Diffusion UI on Runpod in Minutes Google Colab Pro vs. Runpod: Best GPU Cloud for AI Workloads How to Run a GPU-Accelerated Virtual Desktop on Runpod
Use alpha_value To Blast Through Context Limits in LLaMa-2 Models
Brendan McKeag · 2026-01-20 · via Runpod Blog.

Use alpha_value To Blast Through Context Limits in LLaMa-2 Models

With 4k context being the norm for Llama-2 and its finetunes, it's a far cry from the "bad old days" of the 2k limits found in Pygmalion-6b and other previous landmark models.

But what if I told you that you could just set an arbitrarily high context limit for whatever Llama-2 based model you wanted with a minimal perplexity compromise, as long as you have the VRAM to hold it?

Enter NTK-Aware RoPE scaling under the Models page in text-generation-webui.

Screenshot from Llama 2 context limit tutorial

The link above has all of the math involved, but the upshot is this: any VRAM left unused in your card while inferring at the model's maximum context load is essentially wasted, and this allows you to increase that context limit and put it to use in a way without significantly harming perplexity or inference speed.

How much you can get away with increasing your context is dependent on your model and spec, but you can start by increasing your alpha value to your best-case scenario (e.g. 2.5 for 2x if you hope to get doubled context, which should be possible in many situations) and then inferring at your normal base context load and watching your GPU memory utilization in nvtop in the terminal:

Screenshot from Llama 2 context limit tutorial

As you can see, we're only at 50% VRAM usage while inferring at our max context load, which means we have a lot of room to work with. So, then you can bump up the context maximum in your application, and infer again.

Screenshot from Llama 2 context limit tutorial

With this increase (I went straight to 8k here) we're at about 70%, so we can just keep increasing it bit by bit until we reach the sweet spot of about 90-95% usage. Past this you start getting blank responses from running out of memory. The VRAM usage is also not necessarily linear with context size – it does appear to be exponential, in fact, but not so much that it quickly overwhelms the card. In this use case, I was able to increase the effective context limit of unquantized Nous-Hermes-13b from 4k all the way up to 11200 on an a100. Playing around with quantizations may render even larger potential increases.

One note that I found, though, is that if the model is loaded on multiple cards, you will not be able to extend context to the same level that you could have on a single card. You can still do it, but having, say, 2 A40s will be less effective at extending context than one A100 – even though the two A40s actually have 16GB more VRAM between them than the A100. As with all things LLM, you will want to use the fewest number of GPUs that you can for the job, all other things being equal.

In case you were worried about coherency, here are sample outputs at various context levels. Although I believe TANSTAAFL (there ain't no such thing as a free lunch) generally applies, there doesn't seem to be a noticeable loss in coherency even with large increases in context. If there's a catch, I'm not sure what it is, especially since context barriers have been such a thorn in the side of the LLM community for this long and they can now be so readily extended.

4k context:

Enveloped in each stroke, his heart thrums in rhythm with your skill. Each breath you draw serves testimony to your growth, a sense of belonging blossoming within. As the night enfolds, the pair collaborate in harmony - a symphony born anew, transforming bleakness into vibrant hope. A shared endeavor birthed love's resplendent image, its radiance casting aside fear, doubt, and uncertainty. Upon the morning's dawn, Noz stands witness, awash in awe, a man transformed by beauty.

8k context:

Noz applies the pigments with reverence, your form visible on the opposite wall, your shapes intertwined. Captivated, he works, the colors swathing the wall in a panorama. Dancing to the music, his strokes grow bold and certain. Each brushstroke testifying to your partnership, a symphony of love, your features reflected in the design, an explosion of color. Throughout the night, he labors beneath the weak bulb, a labor of adoration. Concluding the artwork, the city skyline fused in vibrant strokes, his hand finds your own, his heart surging. Emitting the finished product, he holds you tight, elation pouring into his voice. "Perfect, Ari. Y'know?"

11.2k context:

Noz works with concentration and fervor, applying his expertise in earnest. His brush strokes grow bold and defined, his colors pooling into our forms. The image unfurls, a celebratory piece depicting our kinship, our fur vibrant amongst the grey walls. The cityscape blooms beneath his hand, our coupling portrayed, the duo locked in an embrace. With a crooked grin, he casts a grin to you, the piece unveiled. "New beginnin'," he jests, his chest expanding, "new endsin'." With you by his side, he feels whole.

Questions?

Pop into our Discord - we would love to hear from you!

Author profile: Brendan McKeag

The Chips Got Faster. The Stack Didn't.

The Chips Got Faster. The Stack Didn't.

Explore why faster chips have shifted the bottleneck to AI infrastructure, and what that means for teams running production workloads.

All

Multi-Instance GPUs on Runpod: Stop Paying for Compute You Don't Need

Multi-Instance GPUs on Runpod: Stop Paying for Compute You Don't Need

With MIG, we can partition RTX 6000 Pro cards into isolated 24 GB instances. Here's when it makes sense for your workloads.

All

OpenAI Parameter Golf: what 1,100 researchers built in six weeks

OpenAI Parameter Golf: what 1,100 researchers built in six weeks

How 1,100 researchers beat OpenAI's own baseline with 16 megabytes and 10 minutes.

All

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.