Weekly Links: Agent Fleets, No to Agent Cruelty, and Software Decompilation

Anthropic bans cruelty to it's AI models, Muse security layers, and software is increasingly vulnerable to AI.

Weekly Links: Agent Fleets, No to Agent Cruelty, and Software Decompilation

This week: PewDiePie launches his own new small language model (with, allegedly, some help from distillation), mid-range model price drops continue, and Google launches its own agents.

On to the main stories:

  • Researchers are tracking a Chinese AI ‘agent fleet’. This is a scary headline for something simple, but interesting. The TechCrunch article is talking about a large amount of new agent traffic to some internet services and guessing that one of the large Chinese Internet companies may be launching a similar Agent product to Meta's Muse agent. It's interesting to see how the idea that agents will be the primary users of the Internet already shows up so readily in traffic data.
  • OpenAI Releases Findings on 377 Math Problems, Further Roiling Field. The publications follow on from the released proofs around the Navier-Stokes equations (which now have some doubts around them, with the lean proof diverging from the human-readable proof). The new sets of proofs are striking in that the average compute time isn't generally very high. I'm not sure picking off all these long, checked platforms really endeared OpenAI to anyone, and they could perhaps have fallen to mathematicians using AI anyway, but in a less apocalyptic way. Scott Aaronson has a thoughtful take on the effect of the proof publications on the mathematics research community.
  • How We Built Safety Into Muse. Meta published an interesting piece on how to architect the security for the Muse personal agent. The centerpiece of the design is to always use short-lived authentication tokens for everything and separate token management into a separate system. This makes sense, but I guess it will mean a large number of human approvals, which could frustrate users. People will probably be quickly looking for the "grant access forever" button, which is inherently dangerous.
  • Anthropic bans ‘abusive or cruel behavior’ toward Claude. Being cruel to your chatbot may soon result in it shutting down conversations. There has been a bit of ridicule about this in some places, and elsewhere there are a lot of warnings about anthropomorphizing AI. However, I think this policy makes a lot of sense. Allowing humans to be abusive and cruel to a system that talks back in this way risks accelerating that behavior (the Stanford prison experiment is a great example). In addition to the fact that AI systems will likely be much smarter than us one day, building respect into interactions from the beginning helps humans stay on the rails as well.
  • Someone Rebuilt Free, Open-Source Versions of Photoshop, Premiere, and Lightroom Using AI. This is just one of the software bombshells from this week. The tool suite was built as a "clean room" implementation, which means that none of the original source code was used, just listings and information about how features work. This makes it potentially free from legal take-downs (though the authors need to be careful they don't use the Adobe name in any claims). It's hard to say how good the project will really get, but as AI improves, packaged software certainly looks at risk. In another way, gaming fans have figured out that you can decompile and edit many existing games back to their source code with Opus 5.5 and above.

Wishing you a great weekend.