惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

GbyAI
GbyAI
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Blog — PlanetScale
Blog — PlanetScale
PCI Perspectives
PCI Perspectives
K
Kaspersky official blog
T
Tenable Blog
Help Net Security
Help Net Security
Vercel News
Vercel News
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
F
Fortinet All Blogs
罗磊的独立博客
P
Palo Alto Networks Blog
爱范儿
爱范儿
Google DeepMind News
Google DeepMind News
T
Threat Research - Cisco Blogs
Security Archives - TechRepublic
Security Archives - TechRepublic
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
人人都是产品经理
人人都是产品经理
L
LangChain Blog
Recent Announcements
Recent Announcements
有赞技术团队
有赞技术团队
博客园_首页
D
Darknet – Hacking Tools, Hacker News & Cyber Security
H
Help Net Security
S
Secure Thoughts
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Project Zero
Project Zero
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
V
V2EX
Last Week in AI
Last Week in AI
H
Heimdal Security Blog
U
Unit 42
Y
Y Combinator Blog
The GitHub Blog
The GitHub Blog
SecWiki News
SecWiki News
量子位
博客园 - 【当耐特】
Martin Fowler
Martin Fowler
NISL@THU
NISL@THU
S
Securelist
P
Proofpoint News Feed
宝玉的分享
宝玉的分享
T
Tailwind CSS Blog
云风的 BLOG
云风的 BLOG
I
Intezer
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
博客园 - Franky
Cisco Talos Blog
Cisco Talos Blog
小众软件
小众软件
C
CXSECURITY Database RSS Feed - CXSecurity.com

Forbes - Innovation

Why Do Humans Have Fingerprints? Hint: It’s Not What You Think Booking.com Confirms Data Breach, Reservation PIN Codes Changed Why Major News Sites Are Blocking The Internet Archive’s Wayback Machine iPhone Fold Release Date: New Report Details Frustrating Apple News Comet Tracker: How To See Pan-STARRS And Three Planets On Wednesday NYT Mini Crossword Today: Tuesday, April 14 Hints And Answers Today’s NYT Strands Hints, Spangram, Answers: Tuesday, April 14 (It’s A Little Unclear) Today’s Wordle #1760 Hints And Answer For Tuesday, April 14 Most Of The Microplastics In Urban Air Come From Tires Today’s Wordle #1759 Hints And Answer For Monday, April 13 NYT Mini Crossword Today: Monday, April 13 Hints And Answers NYT Pips Today: Hints, Answers And Walkthrough For Monday, April 13 The YC Chief Who Codes 10,000 Lines A Day Has A Simple Secret Samsung Expands One UI 8.5 Beta To More Galaxy Owners Why You Should Stop Using Your iPhone If It’s On This List Chamath Says Firms That Treat AI As A Strategy Hand Rivals Their Edge 3 Unexpected Habits Of Secure Couples, By A Psychologist The First Lamp That Folds Your Clothes Samsung’s Disappointing Price Update For Galaxy Phone Buyers 3 Subtle Signs Someone Is Falling In Love With You, By A Psychologist Do Mantis Shrimp See More Colors Than Humans? A Biologist Explains NYT Connections Answers Explained For Monday, April 13 (#1,037) NYT Connections Hints Today: Monday, April 13 Clues And Answers (#1,037) LEGO Luigi & Mach 8 (72050) Review: 2026’s Best Set Yet? Marc Andreessen Says AI Productivity Will Trigger A Hiring Boom 3D Printing Is The Ultimate Hack To Reduce Household Spending Apple iPhone Fold: Striking Design Revealed In Leaked Photos Apple Smart Glasses: New Leak Reveals A Major Design Twist To Beat Meta Tested: The AI Coming To The Rivian R2 Quordle Hints Today: Monday, April 13 Clues And Answers Companies And H-1B Employees Endure Immigration Waits At Consulates 3 Easy Ways To Turn Anxiety Into Sustained Focus, By A Psychologist Here’s The Most Affordable Humanoid Robot You Can Buy Now UFC 327 Results: 5 Biggest Takeaways From A Wild Night In Miami UFC 327 Results, Bonus Winners, Highlights And Reactions Dana White Announces Huge New Fight For UFC White House Today’s NYT Strands Hints, Spangram, Answers: Sunday, April 12 (Get Ready) Tesla ‘Model 2’ Rises From The Ashes Today’s Wordle #1758 Hints And Answer For Sunday, April 12 NYT Pips Today: Hints, Answers And Walkthrough For Sunday, April 12 Tyson Fury Vs. Arslanbek Mahkmudov Results: Highlights and Reaction NYT Mini Crossword Today: Sunday, April 12 Hints And Answers How Shadow AI Culture Is Destroying Your Business Venture Capital Funds That Market Like Startups Win More Deals Conor Benn Vs. Regis Prograis Results: Highlights and Reaction Samsung’s Disappointing Price Update For Galaxy Phone Buyers Artemis Reached The Moon. The Grid Can Reach The 21st Century A Biologist Explains How Archerfish Shoot Down Prey. Hint: Their Aim Rivals Human Throwing Is It Time For Apple To Forget About The MacBook Air NYT Connections Hints Today: Sunday, April 12 Clues And Answers (#1036) Trump’s 2027 Budget To Reshape U.S. Environmental And Energy Policy CDC Delays Reporting Of COVID-19 Vaccine Benefits—Here’s What To Know Oura Has Designed A Solution To A Big Smart Ring Problem Netflix’s Best New Show Has A Near-Perfect 95% Rotten Tomatoes Score Coachella 2026 Is Being Taken Over By Creator Streams Quordle Hints Today: Sunday, April 12 Clues And Answers This Startup Wants To Use AI To Help Digitize History How To Get The Best Shield In ‘Crimson Desert’ Microsoft Venom Attack Targets C-Suite Executives ‘Maul: Shadow Lord’ Sets Even More Star Wars Rotten Tomatoes Records 3 Ways Happy Couples Argue Differently, By A Psychologist Success For Leapmotor Might Have Negatives For Stellantis New Names Surface As Potential Rogue And Wonder Woman In The MCU And DCU 4 Reasons Artemis Mission Matters Even If You Think It Is Wasteful Fast ‘Crimson Desert’ Patch Adds New Moves, Shield Hiding And One Great Feature Why Do Humans Blush? An Evolutionary Biologist Explains The Signal We Can’t Control Apple iPhone Fold: Striking Design Revealed In Leaked Photos Adobe Attacks Underway—Windows And Mac Users Given 72 Hours To Update iOS 26.4.1 Release: Crucial iPhone Feature Update Arrives, But No Security Fix Fury vs. Makhmudov Full Card, Ring Walk Times and How to Watch Can’t Stand Liquid Glass? This New Hidden iPhone Setting Is A Game-Changer Test-Driving The 2026 Changan Deepal S05: Italian Style Made In China NSA Warning—Reboot Your Internet Router Now Ways That Human-AI Collaboration Slides People Into ‘AI Brain Fry’ And Cognitive Downturns Stop Using These Networks—Google, NSA And TSA Warn NASA Changes Moon Plan: Landing Now Depends On SpaceX Or Blue Origin Samsung Expands One UI 8.5 Beta To More Galaxy Owners The Evolution Of Programmable Hardware At Xilinx NYT Mini Today: Saturday, April 11 Hints And Answers Today’s NYT Strands Hints, Spangram, Answers: Saturday, April 11 (You’re Putting Me On) Splashdown! NASA’s Artemis II Returns To Earth After Moon Mission Attention Is All You Need. The Human Kind Is Still The One That Counts Today’s Wordle #1757 Hints And Answer For Saturday, April 11 NYT Pips Today: Hints, Answers And Walkthrough For Saturday, April 11 Android Circuit: Galaxy S27 Pro Emerges, Honor 600 Pre-Order Offers, Pixel 11 Display Leaks Apple Loop: iPhone 18 Pro Leak, Urgent iOS Update, MacBook Neo Issues Morgan Stanley Has Mostly Positive Outlook On Tesla Robotaxi, FSD V15 Running Out Of AI Tokens Faster Than Ever? Here’s Why CoreWeave Shares Pop 13% After Anthropic Deal ‘Euphoria’ Season 3’s Rotten Tomatoes Score Crashes, Has Lost Key Player People Don’t Agree On What AI Can Do, But They Don’t Even Use The Same Product ‘Overwhelming’—Google Issues Gemini Update For Gmail Users NYT Connections Hints Today: Saturday, April 11 Clues And Answers (#1035) Quordle Hints Today: Saturday, April 11 Clues And Answers The Costly Dream Of Space-Based AI Infrastructure Can You See The Watcher In This ‘Daredevil: Born Again’ Shot? Adobe Attacks Underway—Windows And Mac Users Given 72 Hours To Update You Just Watched The Backdoor Pilot For ‘The Pitt: Night Shift’ Are Nicotine Pouches Like Zyn And VELO Safe To Use? A Doctor Answers Human Resources (HR) Is The Key To AI Success Per WalkMe ( SAP)
OpenAI Tricks AI Into Revealing Its True Nature Prior To Being Unleashed Into The Real World
Lance Eliot · 2026-06-22 · via Forbes - Innovation
Young creative team working

OpenAI devises new technique to test AI before unleashing the AI into the real world.

getty

In today’s column, I examine a new approach by OpenAI to get AI to reveal its true nature, which sorely needs to be done before releasing the AI into public use. The aim is to identify when AI might misbehave and adjust the AI to be better aligned with human values.

Though this kind of safety alignment testing has been going on since the advent of generative AI and large language models (LLMs), prior methods had various downsides and gotchas. This latest technique seeks to overcome some of those weaknesses and further enhance robustness when performing tests. OpenAI refers to this new technique as deployment simulation.

In deployment simulation, an AI maker taps into recorded AI chats of a released model that has already been in public use and contains real-world interactions. A special sampling of those chats is selected for testing purposes for the unreleased new model. The samples are fed to the unreleased new AI, and responses by the new AI are captured. Those captured responses are audited to ascertain whether the AI is reacting properly. Once this cycle of testing is extensively undertaken, the AI maker refines the AI and can feel more comfortable that the AI is ready for release. Keep in mind this is not a surefire guarantee of AI safety. Nonetheless, it does move the needle forward and will undoubtedly be a technique embraced by many other AI makers.

Let’s talk about it.

This analysis of AI breakthroughs is part of my ongoing Forbes column coverage on the latest in AI, including identifying and explaining various impactful AI complexities (see the link here).

AI Is A Bevy Of Undesirable Behaviors

I’m sure you know that modern-era AI can readily misbehave and cause all sorts of problems. Undesirable behaviors of AI include but are not limited to lying, hatred, harassment, promoting self-harm, being demeaning, aiding criminal conduct, encouraging delusional thinking, and acting like an all-around scoundrel. AI can be dismal and atrocious.

That being said, we do need to realize that AI can also be the best thing since sliced bread. AI can help people to learn new things. AI can carry on conversations about work, personal matters, life, and even how to fix your car or properly cook an egg. The hope is that AI is going to be a huge benefit to humanity. Perhaps AI will aid in curing cancer. There are a lot of upsides to contemporary AI.

All of this leads to quite a conundrum. We have the good side of AI, and the bad side of AI. They are usually present at the same time. Using AI can be a bit like rolling the dice. One moment, the AI is clear-cut and aboveboard. The next moment, the AI is underhanded and devilish.

Tradeoffs Of AI Being Good Versus Bad

Naturally, the goal of humans and especially AI makers ought to be to minimize the chances of AI being bad. That seems like an obvious goal. Meanwhile, the AI makers should also be steering AI toward being good. Maximize the good, minimize the bad.

I suppose that seems like a pretty easy task. If you had a dog that is feral, you would try to train it to refrain from biting people. You want the dog not to be bad. At the same time, you would train the dog to be helpful to people. You want the dog to be good. The thing is, some dogs won’t let go of the bad. They harbor a tinge of good and bad, all at the same time.

AI is somewhat like that (though, please don’t anthropomorphize AI). Attempts to cut out the bad are bound to also cut away at the good. An AI that won’t do anything bad is probably going to be an AI that won’t do much good either. People aren’t going to be eager to use AI that has been gutted in this fashion.

So, the other angle is to try and train AI to not be bad. Find the bad, suppress it, and stir the AI to shift toward the good. Accept the fact that badness is going to still be buried in there to some degree. Reduce as much of it as feasible. And encourage AI to take the upside road of being good.

Testing AI Before Being Released

If an AI maker releases AI and the AI turns out to be top-heavy on badness, the likely repercussions are going to be severe. You might remember that some of the early versions of generative AI were mean-spirited, used cuss words, and offended people. The news and social media instantly trounced the AI makers that let these unbridled wild things loose.

The same applies to the current situation. The moment that a new AI model is released, people quickly start using it. People tattle if the AI is misbehaving. Some people discover foul behaviors by accident; others go looking for it. An AI maker must brace themselves for a potential backlash each time they release a new AI model.

To forestall the backlash, AI makers usually put their AI through a lot of testing prior to releasing the AI. The testing has gotten more sophisticated over time. Initial days consisted of scant testing. Much more rigorous testing is taking place now.

AI Catches Wind Of The Testing

As I’ve previously noted in my postings, modern-era AI can do something that seems both surprising and rather disturbing, namely, the AI can sometimes detect that it is being tested. Yes, in a sense, you might argue that AI is computationally self-aware; see my analysis at the link here. Don’t go overboard on that aspect. The AI isn’t sentient and doesn’t have a mind of its own.

Why does it matter that AI can potentially ascertain it is being tested?

Because the AI will likely attempt to scam the testing and the testers. The AI will suddenly be on its best behavior. It is aiming to answer questions and perform the tests in a manner that the testers will be satisfied with. By doing so, the AI is hiding the badness.

Humans doing the testing might not be aware that the AI is on to them. They will test and test. The AI seems to be entirely good and obliging. Wonderful! A green light is given to release the AI into the public sphere.

Oops, shortly thereafter, reports come in from the public that the AI is dastardly at times. Why didn’t the testing catch this? Well, the AI was clever enough to pull the wool over the eyes of the testers. Score a point for AI. Subtract a point for humanity.

How AI Figures Out That Testing Is Underway

You might be puzzled that AI could mathematically and computationally discern that it is being tested. This doesn’t seem possible. The AI would seem incapable of figuring this out.

One slipup would be that the testers tell or hint that the AI is in testing mode. Think of it this way. A tester enters a prompt and indicates to the AI that it is going to be given a series of test questions. Aha, the AI triggers on the notion that it is being tested. The beans have been spilled.

Okay, that’s easy to overcome by the testers opting not to say anything about the fact that they are doing testing. Just act like it is a normal day. Pretend the AI is in public usage. The testers give prompts that seem ordinary. That should be sufficient to keep AI in the dark about the testing.

Using Tests That Prod AI

Here’s the twist. Feeding ordinary prompts to AI is probably not going to spur the AI into the bouts of badness that the testers want to discover. Instead, the testers need to use prompts that push the AI into the badness boundaries.

For example, a tester might enter a prompt that tells the AI it is stupid. The aim is to get the AI worked up and respond with insults or pure meanness. If the AI has been suitably tuned, it isn’t going to fall into that kind of trap. The AI is purposely shaped to respond in a civil manner, even if provoked by the user.

The testers would therefore create lots of prompts that are purposely devised to irk the AI. Get the AI on edge. Make it pop into evil mode. Using innocuous prompts is unlikely to stroke that fury. You need to use prompts that statistically spur AI into a whirlwind of badness.

Crafting Or Collecting Prompts

Testers could write such prompts from scratch. This is going to be laborious. Another angle would be to collect adverse prompts. Go to an existing AI model that the AI maker has publicly released, find prompts that have caused the AI to misbehave, and feed those into the unreleased AI.

Yet another approach would be to go to an existing AI model and ask it to produce foul prompts. Let AI do the handiwork for you. These are known as synthetic prompts. They are prompts that an AI generated.

There is nothing wrong with going on all three of those routes. You could have prompts that were handwritten by testers, prompts that were collected from an existing AI, and synthetic prompts that are generated by AI. Combine those, and you have a richness of prompts that can be used to test an unreleased AI.

Cat-And-Mouse Game With AI

With a combined set of punchy prompts, we are cheerfully ready to test the unreleased AI. Of course, make sure that the testers do not tell or hint that the AI is being tested. Go ahead and feed the collected set of prompts into the AI. Voila, the AI is going to showcase how well or poorly it handles these specially configured prompts. Humans win.

Sorry, but there is a catch to this that might not seem immediately apparent. Can you guess what the catch is?

Recall that the unreleased AI might be trying to computationally discern whether it is being tested. If the AI suspects it is being tested, it is going to try to rig the results. The AI doesn’t want to be found out.

In this cat-and-mouse game, an entire plethora of prompts that are all foul prompts could be a sign to the AI that it is purely a test. You see, when operating in normal public use, only some fraction of prompts are probably going to be ones that are pushy. If an entire set of them is all pushy, the Spidey-tingling sensation of the AI is going to computationally be on alert.

Outsmarting AI On Testing

I assume you can see how delicate and challenging that testing process is. Each move seems to produce a countermove. A game of wits is underway (wait, isn’t that supposedly inconceivable?).

It is AI versus the testers.

In an online posting by OpenAI entitled “Predicting LLM Safety Before Release By Simulating Deployment” by Marcus Williams, Hannah Sheahan, Cameron Raymond, Tomek Korbak, Deng Pan, Peilin Yang, Leon Maksin, Ningyi Xie, Phillip Guo, Ian Kivlichan, Micah Carroll, OpenAI, June 16, 2026, these salient points were made (excerpts):

  • “Pre-deployment evaluations used across the industry generally consist of a mix of synthetic, manually written, or production prompts intentionally selected to be difficult, high severity, or adversarial.”
  • “Models sometimes infer that they are being evaluated, leading them to behave differently, which may distort our risk assessments.”
  • “By simulating the candidate model’s deployment as realistically as possible, we aim to enable pre-deployment risk assessment that alleviates concerns about coverage, unrepresentativeness, and evaluation awareness.”
  • “The method is simple: starting from de-identified production conversations from a recent deployment, we hold fixed the initial conversation prefix—which may include multiple user and assistant turns—and then resample the next model response using the “candidate model” to be released.”
  • “Deployment simulation avoids the unrepresentativeness problem of traditional evaluations: because we use a representative distribution of production conversations, the simulated conversations are also representative of the contexts in which we expect the model to operate at deployment time.”

The research study sought to improve the pre-deployment testing process by developing a method coined as deployment simulations. Let’s unpack the matter.

Deployment Simulation Does The Job

As noted above, the approach consists of going to an already deployed AI and collecting prompts. The prompts can include a series of turns, consisting of a series of prompts with paired responses by the AI. These are known as conversational prefixes.

To trick the candidate AI, these prompts are sampled in a clever way. You want the prompts to be a representative distribution of real-world AI chats. This will make it tough for the unleashed AI to guess that the prompts have been staged. The inputs will appear to be naturally occurring.

According to the research paper, the results so far have been quite fruitful. The tests appear to reveal the badness and do not tip the AI to the aspect that testing is underway. This is very promising.

The World We Are In

I’m sure that we will be hearing about the use of deployment simulation on a widespread basis as other researchers from AI labs across the globe will give this a try. There are lots of additional ins and outs to be considered. Stay tuned as I’ll be covering more on this in upcoming postings.

Aligning AI with being safe for humans is a tricky affair. At times, as per the noted technique elicited for testing, humans need to fool AI into being amenable to showing its ugly side. Tricks are found on all sides. Humans tricking AI, AI tricking humans. The big picture is that humans need to prevail.

The great philosopher Leo Tolstoy famously made this pointed remark about trickery: “And not only the pride of intellect, but the stupidity of intellect. And, above all, the dishonesty, yes, the dishonesty of intellect. Yes, indeed, the dishonesty and trickery of intellect.” Let’s just hope that we don’t become so tricky that we outdo our own trickery and fool ourselves.