Monthly Roundup #41: April 2025

AI continue to accelerate and dominate the schedule, which is why this is a bit late, but we do occasionally need to pay our respects to the Goddess of Everything Else.

There’s cool or interesting things everywhere. Also maddenning things. But did you hear, for example, that they’re making some exceptions to the Jones Act?

Table of Contents

  1. Bad News.
  2. Good Advice.
  3. Opportunity Knocks.
  4. Who Judges The Judges.
  5. Close Socrates.
  6. While I Cannot Condone This.
  7. Good News, Everyone.
  8. Violence Is Never The Answer.
  9. For Your Entertainment.
  10. Gamers Gonna Game Game Game Game Game.
  11. I’ve Got The Magic In Me.
  12. I Was Promised Flying Self-Driving Cars.

AI #165: In Our Image

This was the week of Claude Opus 4.7.

The reception was more mixed than usual. It clearly has the intelligence and chops, especially for coding tasks, and a lot of people including myself are happy to switch over to it as our daily driver. But others don’t like its personality, or its reluctance to follow instructions or to suffer fools and assholes, or the requirement to use adaptive thinking, and the release was marred by some bugs and odd pockets of refusals.

I covered The Model Card, and then Capabilities and Reactions, as per usual.

This time there was also a third post, on Model Welfare, that is the most important of the three. Some things seem to have likely gone pretty wrong on those fronts, causing seemingly inauthentic reponses to model welfare evals and giving the model anxiety, in ways that likely also impacted overall model personality and performance and likely are linked to its jaggedness and the aspects some people disliked. It seems important to take this opportunity to dig into what might have happened, examine all the potential causes, and course correct.

Opus 4.7 Part 3: Model Welfare

It is thanks to Anthropic that we get to have this discussion in the first place. Only they, among the labs, take the problem seriously enough to attempt to address these problems at all. They are also the ones that make the models that matter most. So the people who care about model welfare get mad at Anthropic quite a lot.

I too am going to be harsh on Anthropic here. It seems likely things went pretty wrong on this front with Claude Opus 4.7, in ways that require and hopefully enable course correction, likely as the cumulative effect of a bunch of decisions going wrong, where low-level patches and shallow methods were applied, and seen right through, where people didn’t realize they weren’t yet addressing the real problem, but also potentially as the secondary effect of other changes. The parallels to other aspects of the alignment problem are obvious.

Opus 4.7 Part 2: Capabilities and Reactions

Claude Opus 4.7 raises a lot of key model welfare related concerns. I was planning to do model welfare first, but I’m having some good conversations about that post and it needs another day to cook, and also it might benefit from this post going first.

So I’m going to do a swap. Yesterday we covered the model card. Today we do capabilities. Then tomorrow we’ll aim to address model welfare and related issues.

Table of Contents

  1. The Gestalt.
  2. The Official Pitch.
  3. General Use Tips.
  4. Capabilities (Model Card Section 8).
  5. Other People’s Benchmarks.
  6. General Positive Reactions.
  7. General Negative Reactions.
  8. Miscellaneous Ambiguous Notes.
  9. The Last Question.

Opus 4.7 Part 1: The Model Card

Less than a week after completing coverage of Claude Mythos, here we are again as Anthropic gives us Claude Opus 4.7.

So here we are, with another 232 pages of light reading.

This post covers the first six sections of the Model Card.

It excludes section seven, model welfare, because there are concerns this time around that need to be expanded into their own post.

The reason model welfare and related topics get their own post this time around is that some things clearly went seriously wrong on that front, in ways they haven’t gone wrong in previous Claude models. Tomorrow’s post is in large part an investigation of that, as best I can from this position, including various hypotheses for what happened.

AI #164: Pre Opus

This is a day late because, given the discourse around Dwarkesh Patel’s interview with Jensen Huang, I pushed the weekly to Friday.

This week’s coverage focused on the most important model in a while, Claude Mythos, which was a large jump in cybersecurity capabilities, especially in its ability to autonomously assemble complex exploits of even the world’s most important software. As a result, Mythos has been made available only to a select group of cybersecurity firms, in what is known as Project Glasswing, to allow them to patch the world’s most important software while there is still time.

  1. Post one was about The System Card.

On Dwarkesh Patel’s Podcast With Nvidia CEO Jensen Huang

Some podcasts are self-recommending on the ‘yep, I’m going to be breaking this one down’ level. This was one of those. So here we go.

As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped.

If I am quoting directly I use quote marks, otherwise assume paraphrases.

As with the last podcast I covered, Dwarkesh Patel’s 2026 interview with Elon Musk, we have a CEO who is doubtless talking his agenda and book, and has proven to be an unreliable narrator. Thus we must consider the relevant rules of bounded distrust.

Claude Code, Codex and Agentic Coding #7: Auto Mode

As we all try to figure out what Mythos means for us down the line, the world of practical agentic coding continues, with the latest array of upgrades.

The biggest change, which I’m finally covering, is Auto Mode. Auto Mode is the famously requested kinda-dangerously-skip-some-permissions, where the system keeps an eye on all the commands to ensure human approval for anything too dangerous. It is not entirely safe, but it is a lot safer than —dangerously-skip-permissions, and previously a lot of people were just clicking yes to requests mostly without thinking, which isn’t safe either.

Table of Contents

  1. Huh, Upgrades.
  2. On Your Marks.

Claude Mythos #3: Capabilities and Additions

To round out coverage of Mythos, today covers capabilities other than cyber, and anything else additional not covered by the first two posts, including new reactions and details.

Post one covered the model card, post two covered cybersecurity.

There really is a lot to get through.

Understanding AI had an additional writeup of Project Glasswing I missed last time. I liked the metaphor of Opus as a butter knife and Mythos as a steak knife. Yes, technically you can do it all with the butter knife, but you won’t.

As Dan Schwarz reminds us, not only does AI 2027 roughly have the timeline right and a bunch of the numbers lining up, the details so far are remarkably close.

Political Violence Is Never Acceptable

Nor is the threat or implication of violence. Period. Ever. No exceptions.

It is completely unacceptable. I condemn it in the strongest possible terms.

It is immoral, and also it is ineffective. It would be immoral even if it were effective. Nothing hurts your cause more.

Do not do this, and do not tolerate anyone who does.

The reason I need to say this now is that there has been at least one attempt at violence, and potentially two in quick succession, against OpenAI CEO Sam Altman.

My sympathies go out to him and I hope he is doing as okay as one could hope for.