惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
S
Schneier on Security
C
Cyber Attacks, Cyber Crime and Cyber Security
N
News and Events Feed by Topic
TaoSecurity Blog
TaoSecurity Blog
T
Threat Research - Cisco Blogs
博客园 - 三生石上(FineUI控件)
大猫的无限游戏
大猫的无限游戏
The Last Watchdog
The Last Watchdog
Latest news
Latest news
AI
AI
Webroot Blog
Webroot Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
The Hacker News
The Hacker News
Google DeepMind News
Google DeepMind News
S
Securelist
IT之家
IT之家
雷峰网
雷峰网
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
P
Proofpoint News Feed
Last Week in AI
Last Week in AI
博客园 - Franky
美团技术团队
Cyberwarzone
Cyberwarzone
C
CERT Recently Published Vulnerability Notes
Security Archives - TechRepublic
Security Archives - TechRepublic
Security Latest
Security Latest
T
Tailwind CSS Blog
S
Security Affairs
S
Security @ Cisco Blogs
H
Heimdal Security Blog
腾讯CDC
N
News | PayPal Newsroom
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
博客园 - 司徒正美
博客园_首页
Jina AI
Jina AI
M
MIT News - Artificial intelligence
Hacker News - Newest:
Hacker News - Newest: "LLM"
B
Blog
F
Full Disclosure
www.infosecurity-magazine.com
www.infosecurity-magazine.com
T
The Blog of Author Tim Ferriss
Schneier on Security
Schneier on Security
N
News and Events Feed by Topic
NISL@THU
NISL@THU
C
Cisco Blogs
T
Troy Hunt's Blog
O
OpenAI News
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More

Runpod Blog.

DeepSeek V4 in the wild, and how to run it on Runpod New Runpod datacenter now live: AP-IN-1 Track GPU spend across your team with Cost Centers The GPU supply supercycle is here. Here’s what AI builders need to know. Community Spotlight: One-click AI image and video generation on Runpod with SwarmUI | Runpod Blog Community Spotlight: LoRA Pilot Data Prep to Inference Introducing the Runpod Assistant: Manage Your Cloud GPU Resources with Natural Language OpenAI's Parameter Golf: Train the Best Language Model That Fits in 16MB on Runpod LLM inference optimization: techniques that actually reduce latency and cost Pruna P-Video and Vidu Q3 public endpoints now available on Runpod Runpod brand spelling guide Quickstart - Runpod Documentation The AI market looks nothing like the narrative Training StyleGAN3 with Vision-Aided GAN on Runpod KoboldAI – The Other Roleplay Front End, And Why You May Want to Use It How to Connect Cursor to LLM Pods on Runpod for Seamless AI Dev Community Spotlight: How AnonAI Scaled Its Private Chatbot Platform with Runpod Prompt Scheduling with Disco Diffusion on Runpod Runpod's Latest Innovation: Dockerless CLI for Streamlined AI Development Run Your Own AI from Your iPhone Using Runpod Introducing Flash: Run GPU workloads on Runpod Serverless: No Docker required Use Claude Code with your own model on Runpod: No Anthropic account required Avoid Errors by Selecting the Proper Resources for Your Pod What hackers built on Runpod at TreeHacks 2026 Easily Back Up and Restore Your Pod with Cloud Sync + Backblaze B2 The Complete Guide to GPU Requirements for LLM Fine-Tuning AI Guides, Tutorials & GPU Infrastructure Insights | Runpod Your first Claude Code project within Runpod: a complete setup guide 10 billion Serverless requests and counting Building for resilience: Runpod’s response to the AWS us-east-1 outage How to Connect Google Colab to Runpod Founder Series #1: The Runpod Origin Story AMD MI300X vs. NVIDIA H100: Mixtral 8x7B Inference Benchmark How to Run the FLUX Image Generator with ComfyUI on Runpod Run Llama 3.1 405B with Ollama on Runpod: Step-by-Step Deployment How to Run FLUX Image Generator with Runpod (No Coding Needed) How to Use 65B+ Language Models on Runpod Deploy Llama 3.1 with vLLM on Runpod Serverless: Fast, Scalable Inference in Minutes Open Source Video & LLM Roundup: The Best of What’s New Run vLLM on Runpod Serverless: Deploy Open Source LLMs in Minutes Introduction to vLLM and PagedAttention New update to Github integration: release rollback! | Runpod Blog A note to the developers who built Runpod with us Deploy ComfyUI as a Serverless API Endpoint Setting up Slurm on Runpod Clusters: A Technical Guide Building an OCR System Using Runpod Serverless From No-Code to Pro: Optimizing Mistral-7B on Runpod for Power Users Lessons While Using Generative Language and Audio For Practical Use Cases Runpod RoundUp 3 – AI Music and Stock Sound Effect Creation New Navigational Changes To Runpod UI Use alpha_value To Blast Through Context Limits in LLaMa-2 Models Runpod Roundup 5 – Visual/Language Comprehension, Code-Focused LLMs, and Bias Detection Runpod is Proud to Sponsor the StockDory Chess Engine Runpod Roundup 4 – Open Source LLM Evaluators, 3D Scene Reconstruction, Vector Search Meta and Microsoft Release Llama 2 as Open Source SuperHot 8k Token Context Models Are Here For Text Generation How to Manage Funding Your Runpod Account Encrypted Volumes on Runpod: Protect Your Data at Rest How to Run a "Hello World" on Runpod Serverless Runpod AI field notes: December 2025 Faster GitHub Builds: Major Performance Improvements to Our Automated Integration Partnering with Defined AI to Bridge the Data Wealth Gap How to Run Serverless AI and ML Workloads on Runpod How to fine-tune a model using Axolotl Transcribe and translate audio files with Faster Whisper Runpod Achieves SOC 2 Type II Certification: Continuing Our Compliance Journey Orchestrating GPU workloads on Runpod with dstack Exploring Runpod Serverless: Create Workers From Templates DeepSeek V3.1: A Technical Analysis of Key Changes from V3-0324 Deep Cogito Releases Suite of LLMs Trained with Iterative Policy Improvement Wan 2.2 Releases With a Plethora Of New Features Iterative Refinement Chains with Small Language Models The New Runpod.io: Clearer, Faster, Built for What’s Next Introducing Clusters: On-Demand Multi-Node AI Compute Run DeepSeek R1 on Just 480GB of VRAM How Do I Transfer Data Into My Runpod? Spot vs. On-Demand Instances: What’s the Difference? Deploy GitHub Repos to Runpod with One Click Run GGUF Quantized Models Easily with KoboldCPP on Runpod How to Work with GGUF Quantizations in KoboldCPP Introducing Better Forge: Spin Up Stable Diffusion Pods Faster Supercharge Your LLMs with SGLang: Boost Performance and Customization Mastering Serverless Scaling on Runpod: Optimize Performance and Reduce Costs RAG vs. Fine-Tuning: Which Is Best for Your LLM? Run Larger LLMs on Runpod Serverless Than Ever Before – Llama-3 70B (and beyond!) How to Run vLLM on Runpod Serverless (Beginner-Friendly Guide) Embracing New Beginnings: Welcoming Banana.dev Community to Runpod Stable Diffusion + ComfyUI on Runpod: Easy Setup Guide Runpod RoundUp 2 – 32k Token Context LLMs and New StabilityAI Offerings Runpod Roundup: High-Context LLMs, SDXL, and Llama 2 16k Context LLM Models Now Available On Runpod Savings Plans Are Here For Secure Cloud Pods – How To Purchase a Monthly Plan And Save Big Ada Architecture Pods Are Here – How Do They Stack Up Against Ampere? Spin up a Text Generation Pod with Vicuna and Experience a GPT-4 Rival Using OpenPose to Annotate Poses Within Stable Diffusion Set Up a Chatbot with Oobabooga on Runpod Connect VSCode to Your Runpod Instance (Quick SSH Guide) Deploy a Stable Diffusion UI on Runpod in Minutes Google Colab Pro vs. Runpod: Best GPU Cloud for AI Workloads How to Run a GPU-Accelerated Virtual Desktop on Runpod
Pygmalion-7b from PygmalionAI has been released, and it's amazing
Brendan McKeag · 2023-05-17 · via Runpod Blog.

Pygmalion-7b from PygmalionAI has been released, and it's amazing

Last month, the latest iteration of the Pygmalion model was released. Although it is not that much larger as it is still only a 7b model compared to the commonly used 6b version, what it does with that parameter space has also been improved by leaps and bounds, especially with writing that looks to the AI for creative input. The model is not only capable of taking much greater risks while generating its output, but will also succeed in doing so with actions that make sense in context.

Comparing Pygmalion 7b to 6b

There are two primary versions of the new model; Pygmalion 7b directly iterates on the pre-existing model, while the Metharme 7b model has been optimized for creative writing and roleplay.  

To test the models on extremely open-ended creative input, I used a character with an intro paragraph that does not have a great deal of direction, along with a reply that indicates that I want the AI partner to drive the story. With very little direction, it forces the model to get creative.

Chat roleplay transcript with an AI character named Morgan describing a detective arriving at the door

Here are three potential generations for each model from the given input:

Pygmalion 6b:

His tone is calm and collected, yet there is something behind his words. A sense of urgency and determination that could easily be mistaken for panic or fear. "I'm Detective Morgan. We need to talk."
The detective looks at you, his eyes narrowing slightly. He seems almost annoyed, but there is something more there. A certain tension in his voice. He's trying to hide it, but it's still there. "Wow... You have quite the place here. It's a shame about what happened though.
"My name is Detective Morgan. I'm investigating a murder. I believe you may have information that could help us."

Pygmalion 7b:

The detective looks at you, his eyes narrowing slightly. He seems almost annoyed, but there is something more there. A certain tension in his voice. He's trying to hide it, but it's still there.  "Well... I have reason to believe that a crime has been committed. In fact, I believe it may have happened right here in this building."
The detective stands there for a moment, looking you over. He seems somewhat surprised by how...nice you appear to be. "Well, I was hoping to ask you a few questions about a murder that happened earlier today. A young woman was found dead in her apartment, and I wanted to know if she had any visitors recently. If so, I would appreciate it if you could tell me who they were and what they were doing there."
The detective smiles politely at you, his tone calm but firm. He is here for business, and he will get it done quickly and efficiently. The first step is to establish trust, so he begins speaking to you. "My name is Detective Morgan. I am investigating a case where a young woman was found dead earlier today. Her name was Diana Vickers, she was 21 years old and lived alone. I would like to ask you a few questions if you don't mind."

Metharme 7b:

You notice the way his posture is set, a stance that seems confident yet relaxed. His gaze seems sharp and alert, as if he's ready to jump into action at a moments notice. His voice seems to carry a hint of authority and experience, which adds to the mystery of who this stranger is. It seems almost impossible that this could be anything other than a police officer, but you don't know why they would come here. You invite him inside, still unsure of what to expect. The detective steps inside, closing the door behind himself. He gives you a once-over, then looks around your house.
"Well..." He clears his throat, looking down at his feet. "This isn't really a social call, so please forgive my intrusion. But I am investigating a case involving a missing woman, and I think she may have come here last night. Her name is Samantha Baker, she's a nurse, and her husband called us because she never came home after work. She usually gets home late, but she's never missed work. We're worried sick about her, and we need your help finding her. Can you tell me anything?"
After looking around a bit, he steps inside and closes the door. A light breeze enters through the window, carrying the scent of leaves and the sound of children playing outside. He pulls his hat off, placing it carefully beside the door, then folds his hands together in front of him. His voice is calm, yet firm. "Well, my name is Detective Morgan, and I have come to ask you some questions regarding a case that occurred earlier today. Would you care to sit down?"

As you can see, while 6b's generations do attempt to drive the story forward, they are cautious and rather uncreative. Only the 7b models go so far as to create a victim character out of whole cloth to drive the story around, and in particular, the Metharme model includes further evocative writing flourishes while Pygmalion 7b tends to remain more factual. But regardless of your preference, both 7b models clearly wipe the floor with 6b in terms of how much they are willing to output, and how in-depth they will make their poses.

Setting up the Pygmalion 7b and Metharme 7b models on Runpod

According to the PygmalionAI Huggingface repository, the 7b models had to be released as XOR models due to licensing concerns, which require a fair bit of setup that you can read about latest iteration if you're interested. However, a kind soul on HF has already done this and put the complete models up for download, so we can easily grab them with a couple of commands.

First, go ahead and set up a pod under either the Secure or Community Cloud with the Runpod Text Generation UI template.

Runpod Text Generation UI template card for the Runpod/oobabooga:1.0.1 image

Once the pod is set up, you can easily download and set up the models through the Web Terminal with the following commands:

root@46cdf4a0da30:/# cd workspace
root@46cdf4a0da30:/workspace# cd text-generation-webui
root@46cdf4a0da30:/workspace/text-generation-webui# python download-model.py TehVenom/Metharme-7b-Merged-Safetensors
root@46cdf4a0da30:/workspace/text-generation-webui# python download-model.py TehVenom/Pygmalion-7b-Merged-Safetensors

Combined, the two models are about 10gb in size, so be sure you have enough space in your volume for them.

Once they're downloaded, be sure to go to the Models page in Oobabooga and switch to them when you're ready. It's usually a good idea to have multiple models on hand, as at times a model can get stuck on a particular piece of input, and you can switch to a different model for a few actions to get through it to help it break its 'writer's block.'

Text generation web UI Model tab with PygmalionAI 6B and TehVenom Pygmalion 7B models in the dropdown

Credit to TehVenom on HuggingFace for taking the time to do this for the community, and they have several other models that they have also applied weights to if you care to try them out.

Contact Runpod

Have any questions on text generation best practices? Please reach out to us on our Discord!

Author profile: Brendan McKeag

The Chips Got Faster. The Stack Didn't.

The Chips Got Faster. The Stack Didn't.

Explore why faster chips have shifted the bottleneck to AI infrastructure, and what that means for teams running production workloads.

All

Multi-Instance GPUs on Runpod: Stop Paying for Compute You Don't Need

Multi-Instance GPUs on Runpod: Stop Paying for Compute You Don't Need

With MIG, we can partition RTX 6000 Pro cards into isolated 24 GB instances. Here's when it makes sense for your workloads.

All

OpenAI Parameter Golf: what 1,100 researchers built in six weeks

OpenAI Parameter Golf: what 1,100 researchers built in six weeks

How 1,100 researchers beat OpenAI's own baseline with 16 megabytes and 10 minutes.

All

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.