惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Y
Y Combinator Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
The Cloudflare Blog
V
Visual Studio Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
月光博客
月光博客
IT之家
IT之家
大猫的无限游戏
大猫的无限游戏
宝玉的分享
宝玉的分享
博客园_首页
V
V2EX
WordPress大学
WordPress大学
博客园 - 三生石上(FineUI控件)
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
L
LangChain Blog
aimingoo的专栏
aimingoo的专栏
F
Fortinet All Blogs
爱范儿
爱范儿
阮一峰的网络日志
阮一峰的网络日志
GbyAI
GbyAI
Recorded Future
Recorded Future
J
Java Code Geeks
Martin Fowler
Martin Fowler
小众软件
小众软件
人人都是产品经理
人人都是产品经理
Help Net Security
Help Net Security
The Register - Security
The Register - Security
B
Blog RSS Feed
Forbes - Security
Forbes - Security
T
Tailwind CSS Blog
C
CERT Recently Published Vulnerability Notes
P
Privacy International News Feed
D
DataBreaches.Net
博客园 - 【当耐特】
K
Kaspersky official blog
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
T
The Exploit Database - CXSecurity.com
L
LINUX DO - 热门话题
Jina AI
Jina AI
G
GRAHAM CLULEY
H
Help Net Security
D
Docker
Microsoft Security Blog
Microsoft Security Blog
S
Securelist
O
OpenAI News
U
Unit 42
V2EX - 技术
V2EX - 技术
腾讯CDC
罗磊的独立博客

Salesforce

Scale Your MRR: Subscription Management For Small Business Streamlining Commerce Media Ad Inventory Management 12 AI Sales Strategies for Startups That Actually Work Sell Smarter: Ecommerce Metrics To Track For Your Small Business Shop Apply the Orchestration Density Framework to Your Next Automation Decision Wait, Black Friday Planning In Spring? It’s Time to Start Holiday Promotions AI-First Operations, One Process at a Time How BCU Is Transforming Banking Service with Agentforce Salesforce Headless 360: What the Agent Consumer Means for Your Integration Architecture Meet Customers Where They Are: Agentforce Contact Center Now Offers WhatsApp Voice 11 Free Lead Generation Tips for Small and Growing Businesses SFR-VibeTrain: The Agent That Trains Agents Why Technical Accuracy is the Wrong Metric for Agent Success Strengthening Salesforce Security Against AI-Driven Threats Join Us in the Community Hub at Connections 2026 The Best Way To Build AI Agents That Customers Trust 5 Ways AI is Changing the Communication Game For Startups Trust in the Era of Agents: Highlights from the 2nd Annual Trusted AI Impact Report You Can Be an Agentic Enterprise No Matter What Size Business How to Make Your Email Marketing Accessible for Everyone What is Headless? Don’t Lose Your Head, SMBs: It’s a Good Thing Architect the Future UI: Slack as Your Agentic Surface Point of Sale Innovations to Modernize the Shopper Experience Governing the Agentic Enterprise at Scale with MuleSoft Omni Gateway How to Cut Service Time with Case Routing Automation 5 Tips for Marketers to Get Started with Salesforce Flow No One is Vibe Coding Trade Promotion Management 7th Edition State of Sales Report: 3 Growth Trends for Startups and SMBs How the Architect Vista Brought Architectural Thinking to Life at TDX 2026 5 Steps to Develop an Architect Mindset With AI Why AI Isn’t Replacing Developers, It’s Empowering Them 10 Ways to Make Your AI Agent a Better Communicator AI in Design 2025: What Real Use Taught Us. How Salesforce Personalization Learns Which Offers Drive Revenue The 4-Step Guide to Salesforce Agent and Application Development Get Ready for Connections 2026: Top Sessions and New Reveals 8 Ways AI Agents Are Evolving in 2026 4 Principles to Make the Right Salesforce UI Decisions Apprentice Journey Shines a Light on Talent Pathways at Salesforce Agent Script: The Control Plane for Agentic Decisions Scaling the Agentforce Life Sciences Ecosystem to Drive the Future of Pharma and MedTech Unlocking Unstructured Data: Building AI-Powered Support Triage with Data 360 Asking For a Friend: What Are Rich Communication Services (RCS)? 5 Slack Shortcuts For Small Teams 195% ROI In Field Service? Here’s How They Did It Submit Your Architect Session: The Dreamforce 2026 Call for Participation Is Open What Is an AI Assistant for Small Business? B2C Commerce April release: Transforming the B2C developer experience with agents Meet Your 24/7 Prospecting Partner — And 5 More Stand-Out Features In Our April Release 11+ Small Business YouTube Channels You Need to Follow Today Stop Treating Disputes Management Like IT Tickets Limitless Service: A New Operating Model for Growth in the Agentic Era What is Transactional Reconciliation in Email and SMS Marketing? How to Prepare for National Small Business Week (2026) How to Design a High-Scale Multi-Cloud Incident Journey 10 Ways An AI CRM Can Amplify Your Startup Vibe Code Better Agents with Agentforce Free vs. Paid CRM: Which is Right for Your Business? Salesforce Customer Success Awards 2026: Lead Era of the Agentic Enterprise 12 Free Webinars for Small Business Owners (2026) The Agentforce Life Sciences Consultant Certification Maximize Growth: The Power of Partnerships for SMBs Salesforce AI Research at ICLR 2026 Data Protection For Small Business: How To Safeguard Yourself Beyond 100K Tokens: Evaluating AI Agents in Long-Context Software Engineering Top 32 Small Business Tools To Try Today 6 New Innovations Redefining Salesforce Development How to Win the Battle for Attention in the Agentic Email Inbox Mastering the Eisenhower Matrix: Prioritize Like a Pro 10 Signs It’s Time To Upgrade Your CRM and How To Get Started (2026) Celebrating 10 Years of Financial Services Innovation How SMBs Can Gain An Edge With Agentic AI: Key Trends From Our Marketing Report How MAN Truck & Bus is Shaping the Future of Sales with Salesfive Introducing the Future of Salesforce Data Protection: Backup & Recover Next Stop Making AI Slop – Build a Foundation for Authentic AI Content 5 Tips to Help Marketers Navigate AI Email Summaries Data Sharing: Is it Safe? Is it Secure? Everything You Need to Know Small Business Week Readiness: 4 Things to Do Before May (2026) How Salesforce Employees Make an Impact During Earth Month AI Agents Are Advancing Rapidly… Is Your Testing Strategy Keeping Up? What Is Microproductivity and Why Is It Helping So Many Teams? TDX 2026 Roundup: Agentforce Edition 3 Ways Salesforce Connections Has Boosted My Career AI Agents Don’t Just Answer‌ — ‌They Act. Do You Have a Governance Strategy? The Future of MedTech Field Execution is Agentic 5 Steps to Prepare Your Data For an AI CRM From Break/Fix to Profit Engine: Aftermarket Service for Robot OEMs Building Trusted Human-Agent Collaboration: A Practical Framework ISV Strategy for the Salesforce Summer ’26 Release 5 Email Marketing Tips for Small Business Commerce Shops Should You Give Your AI Agent a Human Name? Creating Pathways into AI for People with Disabilities The Organized Chaos of Upfronts: 3 Hurdles Impacting Your Yield Trying to Scale Beyond ‘One-Off’ AI Tasks? You’re Probably Using the Wrong Interface What is Cost Per Lead (CPL)? The Case for Unified CCaaS and CRM — And Why the Data Makes It Clear In the New Era of AI, You Need to Win Over Both Humans and Agents What Is SPIN Selling? A Way to Build Trust With Your Customers G2 Crowns Salesforce as Best Financial Services Software Hidden Insights: The Guide to Tableau For Small Business Owners
Can Language Models Remember What They Learn?
Ye Liu, Srijan Bansal, Shafiq Rayhan Joty, Semih Yavuz · 2026-05-29 · via Salesforce

Post-training methods (RLVR, On-policy distillation) are Episode-local

Language models are getting better at learning from feedback during post-training. In reinforcement learning with verifiable rewards (RLVR), a model tries a problem, a verifier checks the answer, and the policy is updated based on the scalar reward. Recent self-distillation methods go further by using feedback or successful sibling rollouts to create a stronger teacher signal. But most of these methods are still episode-local. A rollout is sampled, scored, used for one update, and then mostly discarded, which wastes signal. Across training, a model repeatedly sees related problems under a changing policy, and those attempts reveal more than just right or wrong answers. They show which strategies keep working and which reasoning patterns transfer to new problems. Procedural Memory Distillation (PMD) is built around this idea: self-improvement should be cumulative.

PMD turns model attempts into reusable training memory

Knowing whether an answer was correct is only part of what a model can learn from an attempt. The reasoning that produced the answer matters too, along with the patterns that are worth holding onto.

PMD converts the model’s own training-time attempts into procedural memory, uses that memory to condition a stronger self-teacher, and distills the resulting guidance into the model’s weights. The memory is used only during training. At inference time, the final model runs normally, without external retrieval or a memory bank.PMD uses memory as a training scaffold, not a deployment dependency.

PMD creates an online loop:

  1. The current policy attempts problems.
  2. A verifier or environment scores the attempts.
  3. PMD stores the attempts as memory.
  4. The model reflects on successes and failures.
  5. Reusable strategies, lessons, and behaviors are extracted.
  6. A memory-conditioned self-teacher trains the next policy.
  7. The updated policy produces new attempts, refreshing the memory.

This is the central design principle: policy and memory co-evolve.

Figure 1. PMD turns repeated attempts into online memory, uses that memory to condition a self-teacher, and distills the guidance into the student.

Hierarchy of procedural memory

PMD organizes memory into three levels.

Experience memory stores the raw trajectories including successful rollouts, failed rollouts, rewards, and feedback. It is faithful, but local and verbose.

Insight memory reflects on attempts for the same problem and extracts strategies and lessons. Strategies describe what led to correct solutions. Lessons describe recurring mistakes and why they failed. When both successes and failures are available, PMD compares them directly to identify what separates correct reasoning from incorrect reasoning.

Behavior memory abstracts across related problems. PMD clusters semantically similar questions and distills their insights into reusable behaviors such as general skills or mistakes to avoid.

This hierarchy creates a fidelity-transfer trade-off. Experience is concrete but hard to transfer, behavior is transferable but more abstract, and insight sits in between.  In our experiments, combining problem-level insights with cross-problem behaviors gives the strongest internalized policy.

How PMD trains the model

PMD builds on self-distillation where the current model plays two roles:

  • The student generates rollouts from the current policy.
  • The self-teacher produces a target distribution after seeing additional context.

Standard self-distillation gives the teacher episode-local context, such as feedback from the current attempt. PMD gives the teacher a broader view that includes strategies from prior correct attempts, lessons from prior failures, and behaviors retrieved from related problems.

The student then learns from this memory-conditioned teacher on its own rollout states. This keeps training on-policy while letting the teacher use knowledge accumulated across earlier attempts.

At inference time, the student receives no memory prompt. If performance improves, the useful procedural knowledge must have been internalized into the model weights.

Figure 2. During PMD training, memory-related concepts such as “strategy,” “lesson,” and “behavior” appear more often in memory-free student rollouts, suggesting that teacher-side guidance is being absorbed into the policy.

What we found

We evaluate PMD on two verifiable domains: SciKnowEval, a science reasoning benchmark covering biology, chemistry, physics, and materials science, and LiveCodeBench, a code-generation benchmark with execution-based unit-test feedback.

Across both benchmarks and two model families, PMD improves over GRPO and SDPO.

With Qwen3-8B, PMD improves SciKnowEval average accuracy from 74.4 with SDPO to 77.2, and LiveCodeBench from 47.9 to 51.7. With OLMo3-Instruct-7B, PMD improves SciKnowEval from 69.5 to 73.3, and LiveCodeBench from 45.0 to 51.1. Relative to SDPO, these are 3.8–5.5% gains on SciKnowEval and 7.9–13.6% gains on LiveCodeBench. PMD uses the same self-distillation backbone as SDPO, but conditions the teacher on online procedural memory. The model’s own training history appears to contain useful signal that episode-local updates miss.

Why the full loop matters

PMD’s gain depends on three pieces working together: reflection on past attempts, persistence of memory across training steps, and co-evolution between policy and memory.

Reflection alone provides some of the lift. A variant called PMD-Transient builds memory from the current batch and discards it after one step. It still improves over SDPO, showing that structured strategies and lessons are useful even without long-term memory.

Persistence adds more on top of that. full PMD keeps memory across training steps, allowing useful patterns to accumulate. This is especially important for code generation, where correct patterns may be rare and need many attempts to consolidate.

Memory by itself, though, isn’t enough. In the Evolving Memory + Frozen Policy condition, the memory keeps improving while the policy weights stay fixed. Performance remains much weaker than full PMD, because the model never internalizes the memory.

The reverse arrangement falls short for a different reason. In the Frozen Memory + Evolving Policy, condition, the policy updates but the memory bank is fixed, and memory written by an earlier policy can become stale as the learner changes.

The takeaway: The learner creates experience, experience becomes memory, memory strengthens the teacher, and the teacher updates the learner. Freezing either the policy or the memory breaks this loop.

Memory transfers across model scales

We also test whether learned memory is useful beyond the exact model that produced it. Memories learned from Qwen3-8B are transferred to models ranging from Qwen3-1.7B to Qwen3-32B.

Across model sizes, memory-augmented inference outperforms the no-memory baseline. PMD co-evolved memory also transfers better than memory built with a frozen policy. Retrieving more memory entries improves performance,which suggests the behaviors are adding signal rather than noise.

Figure 3. Learned memory transfers across model scales. PMD co-evolved memory outperforms frozen-policy memory, and retrieving more memory entries improves performance.

PMD improves test-time scaling

A stronger model should improve its top answer and also preserve useful candidate diversity when sampled multiple times. Candidate diversity matters because methods like majority voting, verifier reranking, and best-of-N selection all depend on good candidates being present in the sample set.

On SciKnowEval, PMD continues to improve as the rollout budget increases, while SDPO saturates earlier. PMD also leaves more verifier headroom: more cases where a correct answer exists among sampled candidates even if majority voting does not pick it.

Figure 4. PMD preserves answer-space coverage as rollout budget increases. The wider gap between majority voting and best-of-k indicates more recoverable accuracy for verifier reranking.
Figure 5. PMD solves a larger set of problems than SDPO across SciKnowEval subjects, suggesting that memory-augmented training expands useful coverage rather than merely shifting which examples are solved.

What happens to the memory bank?

PMD does not simply accumulate more and more text. Experience and insight memories grow quickly early in training, then plateau as per-problem memory reaches capacity. Behavior memory acts more like a consolidation layer: it keeps reusable skills while avoiding redundant behaviors.

Figure 6. Experience and insight memories accumulate problem-specific evidence, while behavior memory shows subject-dependent consolidation dynamics.

The Bottom Line

Language models already learn from feedback. PMD shows they can also learn from the history behind that feedback. The point isn’t to give the model an external notebook forever, but to let it use one while learning and then absorb the useful lessons into its own behavior. A self-improving model needs to track its own history and carry the useful parts of it forward.  That is the central idea behind Procedural Memory Distillation.