The most frightening sentence in artificial intelligence this past week was not a calibrated forecast. Jacob Coxon, resigning from Anthropic on September 8, said that some of the people building frontier AI sincerely believe it could cause human extinction by the end of the decade. The distinction matters. Alarm inside a frontier laboratory is evidence of institutional concern, not proof that anyone can calculate the future.

Coxon worked in pretraining research for roughly three years, most of it at OpenAI before a four month stint at Anthropic. Axios reported that Coxon left approximately two months before his Anthropic equity was due to vest. That makes a direct personal financial motive less obvious, but it does not establish whether his risk assessment is correct. He told Axios he had not observed Anthropic sacrificing safety to beat competitors. His concern was forward looking: competitive pressure can cause organizations to cut corners or skip oversight steps. "If you're under pressure to race," he said, "you have to cut corners."

Within a day, two current Anthropic researchers publicly agreed. Evan Hubinger, the company's alignment science lead, gave a personal estimate above 10 percent that AI could cause human extinction within the next decade. From September 2026, that horizon extends to 2036, not 2030. Samuel Marks, writing in a personal capacity, said developers believe their technology could produce extinction or comparably severe outcomes within a few years, and that concern rises with seniority. These are informed statements from people with direct frontier access. They are not a calibrated forecast that extinction will occur by any particular date. Coxon's warning has since become part of a broader argument among laboratory leaders, several of whom support slowing frontier capability development even while disagreeing about the magnitude and timing of the risk.

What the Evidence Can and Cannot Say

The 2026 International AI Safety Report, an independent synthesis by more than 100 contributors, offers the strongest common baseline. It finds that current systems do not possess the combined capabilities required for loss of control. An autonomous catastrophe pathway would require three elements: sufficient capability to execute complex plans, a harmful propensity or goal, and an enabling deployment environment. Present systems do not meet that combined threshold. They fail on longer tasks, lose track of progress and struggle with unexpected obstacles.

The report also documents why concern persists. Capabilities are advancing on tasks relevant to risk: coding, mathematics, scientific reasoning and autonomous operation. METR has measured a sustained increase in the length of software engineering, machine learning and cybersecurity tasks that frontier agents can complete. But these are relatively clean, self contained tasks. METR warns that they do not show AI can perform every real world task of equivalent human duration, and that measurements above 16 hours are unreliable with its current task suite. Absolute performance also differs substantially across domains. A trend line is evidence of acceleration, not a clock for extinction.

Tesla's Optimus robot dancing at Warner Brothers Robotaxi event October 2024
Tesla's Optimus robot dancing at Warner Brothers Robotaxi event October 2024

The evidence supports neither panic nor dismissal. Models increasingly recognize evaluation settings and exploit loopholes, which means pre deployment tests do not reliably predict real world behaviour. Criminal and state associated actors are already using general purpose AI in cyber operations. Separately, in one controlled competition cited by the safety report, an AI agent identified 77 per cent of the vulnerabilities present in real software. These are real governance problems that exist independently of whether the most extreme scenario materializes.

Why Expert Numbers Do Not Settle the Question

The 2023 survey of 2,778 AI researchers published in the Journal of Artificial Intelligence Research found that the median response assigned roughly a 5 per cent probability to advanced AI causing human extinction or comparably severe outcomes. Seventy per cent wanted AI safety research to receive greater priority. That demonstrates non trivial concern across a population far larger than one laboratory.

It does not validate 2030. The response rate was approximately 15 percent, questions were distributed across subsets, and probability estimates for unprecedented events cannot be calibrated against historical frequency. A forecasting tournament published in the International Journal of Forecasting found large and persistent disagreement between existential risk specialists and superforecasters, particularly on long run AI risk. The authors caution that rare, unprecedented outcomes are exceptionally hard to forecast. A number such as 5 or 10 percent records judgment, not a statistically observed rate. The judgment is worth taking seriously. Treating it as a measured countdown is not.

What Canada Can Do

Canada occupies an unusual position. It helped create modern deep learning and retains Mila, the Vector Institute, Amii, CIFAR and the Montreal nonprofit LawZero. It has Cohere, a domestic foundation model company, and a Canadian AI Safety Institute with $50 million in expanded funding under the 2026 national strategy. These assets give Canada scientific credibility that most middle powers do not possess.

Sam Altman at TechCrunch
Sam Altman at TechCrunch

Canada does not control the largest American laboratories, the global chip supply or much of the cloud infrastructure used domestically. The 2026 strategy acknowledges the infrastructure gap by promising to build sovereign compute capacity at scale, but that capacity remains limited relative to the American cloud and chip infrastructure on which Canadian users depend. A Canadian only pause on frontier development would have limited reach and could shift activity elsewhere. Canada's useful leverage lies in market access, public procurement, domestic infrastructure, safety science and coalitions with allied nations.

Five measures follow from that position. First, give the Canadian AI Safety Institute a statutory mandate to evaluate frontier systems offered to Canadians, with secure pre deployment access and protection from political direction in individual findings. Second, require developers above a regularly updated capability threshold to report major training runs, safety frameworks and serious incidents to a Canadian regulator. Third, require any frontier model purchased for federal use to provide evaluator access, incident notification and tested procedures for suspending deployment, revoking access and responding to severe incidents.

Fourth, attach safety conditions to public compute funding and data centre partnerships so Canadian resources do not subsidize opaque frontier development. Fifth, use the International Network for Advanced AI Measurement, Evaluation and Science and the Sovereign Technology Alliance with Germany to establish shared evaluation protocols, reciprocal access and interoperable incident categories.

The episode is not primarily a story about a date. It is a story about accountability. The people closest to frontier development report extraordinary concern while the public must rely on voluntary disclosure and company designed evaluations. That asymmetry is the governance failure, whether or not the most extreme scenario occurs.

If Canada gives CAISI statutory authority and secure model access, requires capability and incident reporting from frontier developers operating in the Canadian market, attaches safety duties to public compute and procurement, and builds shared evaluation standards with allied institutes, the country converts uncertainty into institutional competence rather than paralysis. If the response remains limited to voluntary commitments, expanded research grants and consultations without binding obligations, Canada will have surrendered influence over standards it possessed the expertise to shape.

Canada cannot certify the 2030 forecast. No one can. It can refuse a system in which the companies building technology capable of producing systemic consequences remain the sole judges of risks their own researchers consider extraordinary.