Discussion about this post

User's avatar
Matthew Tromp's avatar

Nicholas, you are smart, and you have a lot of great posts. But this one is quite poorly thought out. For one, it's not clear to me from this post that you have read AI 2040.

> alignment research is simply capabilities research, patching particular bugs in the system, and we’re not going to find anything deep without actually being in contact with the systems.

First, the whole point of AI 2040 is that we slow down once we already have quite capable systems, but before they are too powerful to keep controlled with our current interpretability/scaleable oversight techniques. We then do work to try and align those systems. We aren't working in some highly theoretical domain here: we would be working with real systems, using them as assistants to our research, but also as the research subjects. Systems which are trained using the same techniques we would later be using to train our superintelligent systems.

Also, saying that alignment research is simply capabilities research is just wrong. The whole problem with highly capable systems is that when you train them to behave themselves (answer honestly, write secure code, etc.) it's impossible to tell whether you're adjusting their goals to align with yours, or simply teaching them how you expect them to behave. That's what the entire subject of deceptive alignment is about, and it remains a hard, unsolved problem. In testing these systems have already shown signs of deceptive alignment and evaluation awareness, the exact problems we predicted they would have before we knew what AI would look like exactly.

> Second, slowing down means that other companies will catch up, and if you are concerned about companies taking risks to outpace their competitors, this is bad. A monopoly is preferable.

The point of AI 2040 is that we coordinate globally to slow down all AI projects. This line is a non sequitur.

> Consider the way people thought about AI in the 2000s, at least on LessWrong. They identified the basic problems AI alignment faces – leaving the invented jargon aside – which is an AI doing things that we don’t want to do, and some plausible reasons for this to be the case, but they had the model of technology completely wrong. There is no apparent anticipation of machine learning, but rather a program, a discrete block of code, which is an agent in the same way you and I are, which could then take off.

Indeed. But again, we wouldn't be doing purely theoretical work in plan A. We would be working on highly capable systems, which were trained with the same techniques that will get us to superintelligence.

> What is so categorically different between “write code to solve this problem” and “don’t kill all the humans”?

Instrumental convergence https://www.lesswrong.com/w/instrumental-convergence

> I think there is much to be said for considering how AIs can monitor each other. We need not understand the inner workings, just as we do not need to read machine code in order to understand programming. I expect that the most important things to do are to use AIs to align other AIs. I don’t see why people expect this to fail – if you think that a dictatorship can last, how about a dictatorship with access to the inner motivations and desires of its citizenry?

This is called scaleable oversight, and it is a subject of active research. But there are some problems with it:

- The monitoring AIs are less capable than the AIs they're monitoring.

- You're assuming we have the tools to confidently evaluate the inner motivations and desires of the AI under monitoring. That's a problem called interpretability, and it is not a solved problem. It's something people are working on.

- You can simply not do any of this control work, save yourself some money, and get your newest model out before your competitors.

> Suppose, though, that we do delay AI via international treaty. What then? To the extent that the frontier labs are delayed, it allows other companies to catch up. We can imagine that there is a tradeoff between the mean and variance of capabilities, at least on the feasible frontier – do you want more companies to take riskier bets?

Yes. This is why Plan A proposes regulating and monitoring all newly manufactured compute, to avoid this exact problem. They even have a lovely chart, the "golden path", about how to balance the risk of making a system too intelligent for us to control, and trailing actors getting ahead with unaligned systems. They thought really hard about these questions!

> I simply believe that, if it were to destroy us all, or to lock us in slavery, it would be at the direction of humans.

Assuming alignment is solved magically, this is also something to be very concerned about! It's called concentration of power risk, and it's another thing that the authors of AI 2040 thought very carefully about and which their plan is designed to minimize. If we're going to build superintelligence, we probably don't want a single company or head of state to have complete control over it, and one of the goals of Plan A is to give competitors a chance to catch up and keep each other in check. From Plan D in AI 2040:

> Second, even if somehow that’s wrong and the resulting superintelligences are robustly aligned, there’s still the question “Who are they aligned to?” and we think that Plan D’s answer to that question is unacceptable. There will be too many tempting opportunities for a CEO or President to become AGI-enabled dictator, and more generally this plan seems like it would lead to the most insane concentration of power in history.

Even if you think alignment will be solved by default, Plan A is still a very good plan! Probably the best plan we have!

> or perhaps, to take off, so that it can give us a subservience preferable to the one that might be given to us by our human masters.

Why would you assume it would do that? Why would you assume it would care about and take care of humanity, instead of simply ignoring us like we ignore ants when we build skyscrapers? What understanding of alignment leads you to believe this would happen? Because AI alignment researchers do not think this is what would happen, and I'm not sure why you think you're more likely to be right about this than they are!

Nicholas, don't you ever encounter people who say things like "Rent is too high. We should just say that landlords can't charge more than X in rent and that will solve the problem"? Don't you get frustrated with these people? Don't you think "who is this person, coming in, making this terrible proposal, without any understanding of the subject and without any awareness that there's an entire field of study dedicated to answering questions like these, and that they have already roundly dismissed that idea as bad." You are doing this to the field of AI alignment. It is a difficult, technical problem which we have not solved and do not have any reason to expect we will be able to solve before we run out of time.

Locrian's avatar

I think a big part of the disagreement between your position and the more worried one is to what extent current AI is aligned for tasks like “write code to solve this problem”--whether it will reliably do the things it knows you would want, and avoid doing things it knows you wouldn't want, in addition to following the literal text of the instruction. Your view is: 1) current AI is aligned, 2) future AI will probably behave basically like current AI, except that it will be smarter, therefore 3) future AI will be aligned by default. The opposite view follows the same logic, but starting from the premise that current AI is not aligned.

4 more comments...

No posts

Ready for more?