Weekly Links: AI Alignment, Pacing, and Two Models of Personal Agents

AI Alignment and pacing the frontier take center stage this week.

Weekly Links: AI Alignment, Pacing, and Two Models of Personal Agents

Oh boy. This is the week the large labs and many others turned the alignment narrative up to 11, with programs and warnings about AI advances happening too fast. There are clearly real risks, and they need to be considered. However, some of the proposed mechanisms seem to be striking at the wrong problem.

The main stories:

  • Both OpenAI (also see "An Alien Mind") and Anthropic (and pacing the frontier) this week released essays on "slowing down" frontier model development and controlling the evolution of models. I do think there are sensible provisions in the proposals, that international cooperation is key, and that we absolutely need to establish rules about how AI should be used and operated. However, I'm quite skeptical about a lot of the points made - particularly things like limiting progress based on inputs (compute used, etc.) or whether some of the goals (e.g., interpretability) have any legs. The issue, in my view, is that models are already at a level where, even if you keep compute constant, new methods will 2x, 3x, and 10x capabilities in a few short years. We're already at AGI by various definitions, and we're actually already on the inevitable path to SGI in some specialisms even if you stopped pre-training. The danger with now trying to "shut the gate" is that we end up with a small set of organizations "permitted" to be at the frontier, or with regulation that no one but the largest organizations can implement. This leaves whole economies dependent on a few actors. I'd rather have distributed progress and strong incentives to do alignment work. Amodei does suggest this, but it's buried in a lot of other items. I'm not sure how this will all be received, but governments certainly need to react to AI and prepare for it. Slowing the frontier makes sense, but I'm skeptical that direct legal prohibitions are the way to go. It seems better to focus on clear legal guardrails on the use of the technology: prohibiting certain uses and (above all) making liability for incidents caused by AI clear for those that deploy the technology. These rules create back pressure on labs to build controllable, safe systems. It also creates legal liability for technology providers that promise levels of safety and then don't deliver. Note that I'm not trivializing this; it's incredibly hard to build AI systems that are both effective and "safe". The key point is that mechanisms such as observers and/or restricting inputs during model development are unlikely to make things safer (maybe you need MORE compute to build a safe model).
  • The paradox at the heart of AI and science | Terence Tao (Youtube Video). This is a great video by a leading mathematics professor and researcher. I think he neatly captures the strength of human scientific/mathematical exploration v's the current cutting edge of AI. In his analogy, the human way is like a hike to a great destination. Along the way, you may well find many interesting things. The AI way is more like being parachuted right to the destination. It's valuable, but you don't learn a lot of those important lessons. I think this makes a lot of sense, and the adoption of AI in mathematics research is a super hot topic at the moment (see the Math and AI declaration for example: https://mathandai.org/). Tao does not argue AI shouldn't be used, just that it has different effects than human exploration. The Math and AI declaration signatories also express the risk that "speedrunning" famous Math problems might demotivate the field and collapse the human ecosystem of mathematicians. I can well imagine it would be highly demotivating to work on a problem for many years and then see it solved in an inelegant brute-force way (after a lot of token spend). Ultimately, though, these proofs (however they arise) are new building blocks, and so the frontier moves. I suspect for Math it will be a lot as it has been for code. The frontier of what can be built is moving; things are messy (and so are some systems), but it will probably mean an even greater need for mathematicians and their craft. Human+AI in math will likely be the winning play for many years.

  • Shopify extends lifeline to Tailwind as vibe coding erodes web dev platform's bottom line. Tailwind is a really widely used web framework, and the team behind it previously sustained itself with premium versions. However, with the rise of vibe coding, the open-source version is seeing wide use, but there's much less incentive to upgrade. Shopify stepping in may help Tailwind, but this scenario is likely playing out across 1000s of open-source products. There really needs to be a way to sustain key projects.
  • Introducing Muse - Your personal AI agent that gets things done. This could be one of the week's most consequential announcements. Meta is launching a personal assistant agent that can handle many daily tasks, such as booking restaurants, managing calendars, and more. A good way to think of this is as the "consumer" version of Codex (ChatGPT Work) and Claude CoWork. Each Muse agent comes with a fully fledged VM in the cloud that serves as the agent's 24/7 runtime. My guess is initial adoption will be slow (people don't know what these systems can do), but Meta will grind it out and later connect it to all their social experiences.
  • At the same time, Apple released the Duo, which is the latest device in the race to be an "AI hub" where AI control lives on your local device. Unfortunately for Apple, I think that if you have to choose between the two, having your AI 24/7 in the cloud is likely to be the more important part of the architecture.

Among the smaller stories: Americans can now also buy Swedish Robot Mowers again, Google Deepmind publishes predictions for 9 Billion possible gene mutations, and the fight for chatbot rights.

The AI alignment challenges are real, and it's good to see key researchers signing up to coordinate on AI safety and development. It will be interesting to see whether the focus stays on safety or development speed simply keeps going at the current pace.

Wishing you a good weekend!