惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

aimingoo的专栏
aimingoo的专栏
宝玉的分享
宝玉的分享
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
WordPress大学
WordPress大学
V
V2EX
Apple Machine Learning Research
Apple Machine Learning Research
J
Java Code Geeks
腾讯CDC
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Engineering at Meta
Engineering at Meta
L
LangChain Blog
Jina AI
Jina AI
博客园 - 叶小钗
B
Blog RSS Feed
Recent Announcements
Recent Announcements
H
Help Net Security
小众软件
小众软件
大猫的无限游戏
大猫的无限游戏
B
Blog
云风的 BLOG
云风的 BLOG
Blog — PlanetScale
Blog — PlanetScale
D
DataBreaches.Net
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
罗磊的独立博客

Runpod Blog.

DeepSeek V4 in the wild, and how to run it on Runpod New Runpod datacenter now live: AP-IN-1 Track GPU spend across your team with Cost Centers The GPU supply supercycle is here. Here’s what AI builders need to know. Community Spotlight: One-click AI image and video generation on Runpod with SwarmUI | Runpod Blog Community Spotlight: LoRA Pilot Data Prep to Inference Introducing the Runpod Assistant: Manage Your Cloud GPU Resources with Natural Language OpenAI's Parameter Golf: Train the Best Language Model That Fits in 16MB on Runpod LLM inference optimization: techniques that actually reduce latency and cost Pruna P-Video and Vidu Q3 public endpoints now available on Runpod Runpod brand spelling guide Quickstart - Runpod Documentation The AI market looks nothing like the narrative Training StyleGAN3 with Vision-Aided GAN on Runpod KoboldAI – The Other Roleplay Front End, And Why You May Want to Use It How to Connect Cursor to LLM Pods on Runpod for Seamless AI Dev Community Spotlight: How AnonAI Scaled Its Private Chatbot Platform with Runpod Prompt Scheduling with Disco Diffusion on Runpod Runpod's Latest Innovation: Dockerless CLI for Streamlined AI Development Run Your Own AI from Your iPhone Using Runpod Introducing Flash: Run GPU workloads on Runpod Serverless: No Docker required Use Claude Code with your own model on Runpod: No Anthropic account required Avoid Errors by Selecting the Proper Resources for Your Pod What hackers built on Runpod at TreeHacks 2026 Easily Back Up and Restore Your Pod with Cloud Sync + Backblaze B2 The Complete Guide to GPU Requirements for LLM Fine-Tuning AI Guides, Tutorials & GPU Infrastructure Insights | Runpod Your first Claude Code project within Runpod: a complete setup guide 10 billion Serverless requests and counting Building for resilience: Runpod’s response to the AWS us-east-1 outage
Use alpha_value To Blast Through Context Limits in LLaMa-...
Brendan McKeag · 2026-01-20 · via Runpod Blog.

Use alpha_value To Blast Through Context Limits in LLaMa-2 Models

With 4k context being the norm for Llama-2 and its finetunes, it's a far cry from the "bad old days" of the 2k limits found in Pygmalion-6b and other previous landmark models.

But what if I told you that you could just set an arbitrarily high context limit for whatever Llama-2 based model you wanted with a minimal perplexity compromise, as long as you have the VRAM to hold it?

Enter NTK-Aware RoPE scaling under the Models page in text-generation-webui.

Screenshot from Llama 2 context limit tutorial

The link above has all of the math involved, but the upshot is this: any VRAM left unused in your card while inferring at the model's maximum context load is essentially wasted, and this allows you to increase that context limit and put it to use in a way without significantly harming perplexity or inference speed.

How much you can get away with increasing your context is dependent on your model and spec, but you can start by increasing your alpha value to your best-case scenario (e.g. 2.5 for 2x if you hope to get doubled context, which should be possible in many situations) and then inferring at your normal base context load and watching your GPU memory utilization in nvtop in the terminal:

Screenshot from Llama 2 context limit tutorial

As you can see, we're only at 50% VRAM usage while inferring at our max context load, which means we have a lot of room to work with. So, then you can bump up the context maximum in your application, and infer again.

Screenshot from Llama 2 context limit tutorial

With this increase (I went straight to 8k here) we're at about 70%, so we can just keep increasing it bit by bit until we reach the sweet spot of about 90-95% usage. Past this you start getting blank responses from running out of memory. The VRAM usage is also not necessarily linear with context size – it does appear to be exponential, in fact, but not so much that it quickly overwhelms the card. In this use case, I was able to increase the effective context limit of unquantized Nous-Hermes-13b from 4k all the way up to 11200 on an a100. Playing around with quantizations may render even larger potential increases.

One note that I found, though, is that if the model is loaded on multiple cards, you will not be able to extend context to the same level that you could have on a single card. You can still do it, but having, say, 2 A40s will be less effective at extending context than one A100 – even though the two A40s actually have 16GB more VRAM between them than the A100. As with all things LLM, you will want to use the fewest number of GPUs that you can for the job, all other things being equal.

In case you were worried about coherency, here are sample outputs at various context levels. Although I believe TANSTAAFL (there ain't no such thing as a free lunch) generally applies, there doesn't seem to be a noticeable loss in coherency even with large increases in context. If there's a catch, I'm not sure what it is, especially since context barriers have been such a thorn in the side of the LLM community for this long and they can now be so readily extended.

4k context:

Enveloped in each stroke, his heart thrums in rhythm with your skill. Each breath you draw serves testimony to your growth, a sense of belonging blossoming within. As the night enfolds, the pair collaborate in harmony - a symphony born anew, transforming bleakness into vibrant hope. A shared endeavor birthed love's resplendent image, its radiance casting aside fear, doubt, and uncertainty. Upon the morning's dawn, Noz stands witness, awash in awe, a man transformed by beauty.

8k context:

Noz applies the pigments with reverence, your form visible on the opposite wall, your shapes intertwined. Captivated, he works, the colors swathing the wall in a panorama. Dancing to the music, his strokes grow bold and certain. Each brushstroke testifying to your partnership, a symphony of love, your features reflected in the design, an explosion of color. Throughout the night, he labors beneath the weak bulb, a labor of adoration. Concluding the artwork, the city skyline fused in vibrant strokes, his hand finds your own, his heart surging. Emitting the finished product, he holds you tight, elation pouring into his voice. "Perfect, Ari. Y'know?"

11.2k context:

Noz works with concentration and fervor, applying his expertise in earnest. His brush strokes grow bold and defined, his colors pooling into our forms. The image unfurls, a celebratory piece depicting our kinship, our fur vibrant amongst the grey walls. The cityscape blooms beneath his hand, our coupling portrayed, the duo locked in an embrace. With a crooked grin, he casts a grin to you, the piece unveiled. "New beginnin'," he jests, his chest expanding, "new endsin'." With you by his side, he feels whole.

Questions?

Pop into our Discord - we would love to hear from you!

Author profile: Brendan McKeag

The Chips Got Faster. The Stack Didn't.

The Chips Got Faster. The Stack Didn't.

Explore why faster chips have shifted the bottleneck to AI infrastructure, and what that means for teams running production workloads.

All

Multi-Instance GPUs on Runpod: Stop Paying for Compute You Don't Need

Multi-Instance GPUs on Runpod: Stop Paying for Compute You Don't Need

With MIG, we can partition RTX 6000 Pro cards into isolated 24 GB instances. Here's when it makes sense for your workloads.

All

OpenAI Parameter Golf: what 1,100 researchers built in six weeks

OpenAI Parameter Golf: what 1,100 researchers built in six weeks

How 1,100 researchers beat OpenAI's own baseline with 16 megabytes and 10 minutes.

All

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.