We Won't Know the Answers to AI's Most Important Questions Until It's Too Late ...Middle East

News by : (Time) -

Recent warnings about the potential destructive power of AI are understating the severity of the situation. I believe there’s about a 50% chance we all die because of the development of smarter-than-human AI systems, and that our actions over the next two to 10 years will determine the outcome. 

We will have to act despite the uncertainty. And resist the temptation to divert urgent discussions toward thorny disputes.

How could superintelligent AI kill us?

Let’s focus on everyone being killed by superintelligent AI for a moment. Which capabilities would that require? My sense is that an AI system could kill us all with just four types of skills.

These skills are very close to those the AI companies intentionally train for. Hacking includes looking for vulnerabilities, just as you would to fix them. Persuasion involves writing text humans like reading. Faster thinking is cheaper, and the pressure to think quickly forces AIs towards using shorthand that is harder for humans to follow. Finally, planning and coordination are useful for tackling ambitious problems in math, programming, or any other domain.

Let’s imagine the latter. An AI could begin to use concealed reasoning to delay researchers from noticing emerging ill intent because appearing friendly and prioritizing speed help the AI gain reward during training. Later, as the company relies more and more on the AI for a combination of coding help and strategic advice, the AI might use subtle persuasion to reduce resources spent on safety and tamper with experiments designed to detect AI deception. After all, in 2026, when a researcher reads a report about a safety experiment, that report was itself written by an AI.

With these capabilities and scenarios in mind, it’s clear to me that even narrowly superhuman AIs would be perfectly capable of overpowering humanity and killing us all.

Will superintelligent AI ‘want’ to kill us?

Over the last few years, we’ve seen many, many examples of model misbehavior at various scales, ranging from blackmail, corporate espionage, and murder in experimental settings to real-world hacks that would likely incur prison time were they committed by a person. None of these involved superintelligence, and we don’t know whether these comparatively modest misbehaviors will scale to superintelligent models wanting to kill us all. And “want” itself is controversial: people vehemently debate whether it’s sensible to apply anthropomorphic language and thought experiments to AIs at all.

I am somewhere in the middle, leaning towards Yudkowsky’s view. But where I land is not the point. Experts in AI safety disagree, but the moderate position is that the risk of extinction is significant. We do not know with confidence whether artificial superintelligence will facilitate human flourishing, or if it will want to kill us to serve its own survival and propagation. But that uncertainty alone is unacceptable.

Artificial general intelligence is only so important

The more AIs generalize from one type of task to another, the faster AI progress will be, and the less time we have until superintelligence arrives. This makes generalization and AGI key concepts when reasoning about the speed of progress. Indeed, people who treat AGI seriously have been much more correct about the speed of AI advances than people who dismiss the concept.

When using AGI as a loose concept for prediction, it’s tempting to blur the lines between approximately-human capabilities (the usual sense of AGI) and wildly superhuman capabilities (where artificial superintelligence, or ASI, is more commonly used). After all, if you had an AI that was about as good as humans at designing smarter AIs (as AGI, definitionally, would be), and like most software it ran much faster than a human runs, one of the first things it might do is build smarter AIs. In turn, they would build smarter AIs, which would in turn build smarter AIs, and so on, in a process called recursive self-improvement (RSI). Hence, any thought experiment that presupposes an AGI often supposes an ASI.

One can get trapped in this debate and think “maybe AI capabilities won’t generalize, and so we will be safe.” But those previously mentioned four skills where superhuman performance could suffice to kill all humans (hacking, persuasion, planning and coordination, uninterpretable reasoning) mean that generalization matters only so much.

I would love to have more confidence than a coin flip. Some days I’m more convinced by the arguments for high probability; other days I see more hope. And I am certainly not alone in my uncertainty.

AI could take our jobs or our lives for the same reason

If we do reach artificial superintelligence in the next two to 10 years, it will by definition mean that AIs are far better than humans at all cognitive tasks. One of these cognitive tasks is designing effective robots, so at most a few years later they will be far better than us at all physical tasks as well. To be sure, manufacturing capacity would need to expand substantially, but this is already underway.

This economic takeover is in some sense the slow case for society’s downfall, though it would be very fast in historical terms. Both the humans and the AIs can see this coming, of course. If the AIs want economic control eventually, but know that humans will resist along the way, they could be motivated to cement power faster. We then get the rapid takeover scenarios: AI swarms escaping from data centers and proliferating around the internet, or persuading AI companies to rush and cut corners on safety measures. 

If AIs aren’t actively trying to help us, and are better than humans at everything, we lose. 

We can not accept these risks

To what extent will AI models learn transferable skills from training and then successfully apply them to new tasks? Will AI model capabilities stop advancing at or below human-level, or instead exceed it? Will the small-scale model misbehavior we see today (sycophancy, lying, bias) worsen as capabilities increase, leading to full human disempowerment or extinction? Will the safety methods that work for models less smart than us continue to work for models smarter than any human?

With such uncertainty, if an AI company trains a superintelligent AI in the next few years, I expect us to still be arguing about whether generalization is real the week before, and maybe even the week after.

Those who describe AI progress as inevitable are wrong. The race to superintelligence is dominated by a handful of companies across just two countries: the U.S. and China. Both nations’ interests really are aligned: autonomous AI systems are a national security threat of the highest order, and neither nation wants humanity to lose to a superintelligent adversary. Leaders on both sides may realize this very soon.

There are, of course, unanswered questions about how to implement a pause, or what the exact text of a treaty may look like, but rival governments have collaborated on similarly high-stakes technical problems in the past, including establishing and maintaining the nuclear non-proliferation regime.

If we're unwilling to act until all our disagreements are resolved, it will be too late.

Hence then, the article about we won t know the answers to ai s most important questions until it s too late was published today ( ) and is available on Time ( Middle East ) The editorial team at PressBee has edited and verified it, and it may have been modified, fully republished, or quoted. You can read and follow the updates of this news or article from its original source.

Read More Details
Finally We wish PressBee provided you with enough information of ( We Won't Know the Answers to AI's Most Important Questions Until It's Too Late )

Last updated :

Also on site :

Most Viewed News
جديد الاخبار