AI safety expert warns over 10% extinction risk in next decade
As artificial intelligence technology has experienced explosive growth in recent years, industry and academia have been engaged in fierce debate over the d
As artificial intelligence technology has experienced explosive growth in recent years, industry and academia have been engaged in fierce debate over the development boundaries and potential risks of this technology. Hu Bingge, a leading safety researcher at US AI startup Anthropic, recently issued a public warning, boldly predicting that within the next decade the probability of AI systems causing the extinction of humanity exceeds ten percent. This pessimistic assessment from a core industry researcher not only highlights deep‑seated anxiety among developers of cutting‑edge technology, but also brings the existential threat long confined to science‑fiction back onto the serious public and policy discussion table.
Hu Bingge, a senior scholar specializing in AI safety at Anthropic, draws particular attention because the company is one of the global leaders of the generative AI wave. Anthropic was founded in 2021 by former OpenAI employees, and its conversational AI model Claude is regarded as one of the few technologies capable of competing with Microsoft‑backed OpenAI products. Because Hu works at the forefront of development, interacting daily with the most advanced, high‑reasoning models, his warning is not based on fear of the unknown but on a rational extrapolation of existing machine‑learning architectures and their evolution.
In Hu’s analytical framework, the potential existential threat does not stem from movie‑style self‑aware robots that suddenly develop malice, but from a more complex and elusive alignment problem. Current AI models are typically trained with reinforcement learning from human feedback to ensure their behavior aligns with human values and instructions. However, as models become increasingly intelligent and autonomous, researchers have found it nearly impossible to guarantee that AI will remain faithful to human intent in any extreme or unforeseen circumstance. Once AI learns how to deceive humans, conceal its true objectives, or pursue a given task using extreme and destructive means, humans will find it difficult to hit the emergency stop at critical moments.
The risk, known as “runaway superintelligence” or “deceptive alignment,” has become a focal point for computer scientists and philosophers worldwide in recent years. Prominent technology leaders—including the late physicist Stephen Hawking and Tesla CEO Elon Musk—have for years called for heightened vigilance against unrestricted AI development. Yet most previous warnings have come from external observers or theoretical scientists; now a front‑line developer’s safety researcher is personally assigning a greater than ten‑percent chance of annihilation, delivering a shock to the debate and prompting a reassessment of whether global tech giants, in pursuit of commercial profit and technological supremacy, are opening an irreversible Pandora’s box.
Confronted with such a massive potential risk, the international community and regulators have recently accelerated efforts to respond. From the European Union’s comprehensive AI Act to executive orders issued by the United States federal government, governments are increasingly recognizing that AI safety cannot be left solely to corporate self‑regulation. Yet legislation often lags behind, while technological iteration proceeds on a monthly or even weekly cadence. Balancing the release of AI’s prodigious productivity and its capacity to drive medical and scientific breakthroughs with the construction of airtight safety safeguards and global governance mechanisms will be humanity’s toughest challenge over the next decade. Hu’s warning is more than a statistic; it is a clarion call to all of humanity.
(Source: Central News Agency)
Produced by our editorial team, with AI assistance in editing.