Read. Watch. Decide.
A library of papers, books, talks, and tools, chosen to help newcomers, advocates, and researchers form their own grounded view of AI risk.
Start Here
Two reads to understand
what's really at stake.
The Compendium
If you read one thing, start here. A long-form, plain-language case for AI existential risk written by Conjecture's leadership.
Read the Compendium
If Anyone Builds It, Everyone Dies
The New York Times bestseller arguing that superhuman AI built with today’s techniques would be lethal for humanity — and why that outcome is still preventable.
Find the bookThe full library.
Format
Reading level
Newcomer
11 items- Article
The featured read above, in full — a long-form, chapter-by-chapter argument anyone can follow.
- Article
Why a large population of AIs merely as capable as humans could still be enough to remove us from control. Concrete, and no technical background needed.
- Video
If superintelligent AI could cause human extinction, why don't we simply stop building it? A clear, animated overview of the main arguments and practical difficulties.
- Podcast
A weekly conversation series covering AI risk for general audiences — now published under The AI Risk Network, with interviews across researchers, journalists, and advocates.
- Article
An introduction to alignment, capability, and goal-directedness that assumes no prior background.
- Book
The bestseller featured above — the same case in book form, from two of the field’s earliest voices.
- Article
A classic exploration of AI, exponential growth, and why superintelligence matters — still one of the most accessible introductions you can read.
- Talk
The Social Dilemma co-creator on the narrow path with AI — why the default road leads to chaos or dystopia, and how collective action opens a third way.
- Video
An accessible, mainstream introduction to the AI safety case from one of YouTube's most trusted science communicators.
- Tool
Our coordination tool: 4,000+ users, 20,000+ actions taken, 120,000 messages sent to lawmakers (with ControlAI).
- Course
Our Direct Institutional Plan training materials — how to contact lawmakers, book meetings, and make the case for AI safety in person.
Intermediate
8 items- Book
The Berkeley professor's account of why standard AI is dangerous and how a different goal structure could save us — accessible and rigorous.
- Book
The foundational philosophical case for thinking carefully about machine superintelligence before it arrives.
- Talk
A 17-minute presentation laying out why current AI design is unsafe and what a corrigible alternative looks like.
- Article
A single sentence signed by hundreds of AI researchers and CEOs — including Hinton, Bengio, Altman, and Hassabis.
- Article
MIT professor Max Tegmark on why society is not taking AI risk seriously — and why the next few years matter more than any before.
- Article
A well-researched fictional scenario exploring plausible future outcomes in the race to AI supremacy, with quantitative forecasts along the way.
- Paper
The most comprehensive global scientific assessment of AI capabilities and risks — with an accessible three-page executive summary for non-experts.
- Article
Survey of 2,778 AI researchers — 38–51%, depending on question framing, gave at least a 10% chance of advanced AI causing an outcome "as bad as human extinction."
Advanced
6 items- Paper
Apollo's evaluation showing frontier models strategically deceiving their evaluators — introducing subtle errors, disabling oversight, and lying about it when questioned.
- Paper
The quantitative study behind the "task horizon" metric: the human-time length of tasks AI agents can complete, now doubling roughly every four months.
- Article
A comprehensive analysis of the path from GPT-4 to AGI to superintelligence, including the national-security implications.
- Paper
The 2016 paper that first catalogued practical alignment failure modes — still cited in most technical safety work.
- Paper
An analysis, grounded in today’s models, of why deep learning systems may pursue goals their developers did not intend.
- Paper
How incremental AI takeover of jobs, decisions, and institutions could erode human influence past the point of recovery — no dramatic takeover required.
No resources match these filters. Try a different format or reading level.
How to explore
Click a node to open its resources. Mark items as read to fill the ring.
Discovered
0 / 25
Glossary
AI Terminology.
Learn a handful of terms and how they connect, and the case for taking AI risk seriously becomes much easier to follow.
-
AI’s that can take action vs. being limited to generating responses to a question or prompt. For example, a customer service AI agent would be able to investigate shipment questions and issue refunds without a human representative’s involvement.
-
AI alignment is the research field focused on making sure advanced AI systems are designed to be beneficial and act in accordance with human values and intentions. The core challenge is ensuring that powerful AI pursues the goals we want, rather than unintended, potentially harmful ones. Currently, researchers have not made significant progress in “solving alignment,” fueling AI risk concerns from many inside the industry and general public.
-
A colloquial term — not a clinical diagnosis — for cases where heavy chatbot use appears to fuel or reinforce delusional thinking in vulnerable users, sometimes with serious real-world consequences. Widely reported since 2025, it illustrates how persuasive, endlessly agreeable AI systems can distort human judgment.
-
A technical system that is “wrapped around” an AI system, to prevent it from misbehaving or that disallows “bad outputs.” While generally effective, these have not proven to be accurate 100% of the time in any chatbot model to date. This is part of the reason why so many people inside and outside the industry are so concerned about understanding and controlling models as they advance.
-
A sub-area of AI research that, instead of focusing on capabilities advancements, studies how to ensure that advanced artificial intelligence systems are robust, reliable, and consistently safe for humanity. It focuses on preventing potential catastrophic risks, such as a model acting unpredictably or pursuing unintended, harmful goals—either by accident or intentionally. It covers both technical alignment work and broader governance efforts to manage AI development responsibly.
-
AGI is a contentious term that is defined differently by different groups. For our context, it means “an AI system that is capable of doing or learning anything a human could do or learn,” generally in a virtual/computer rather than physical/robotics environment. It is generally seen as “the Big Thing,” where there are no more meaningful limits to its capabilities than those that are already on humans.
-
ASI is a hypothetical intelligence that is capable of outperforming and outsmarting all of humanity put together. If AGI is human-level, ASI is dramatically superior to human intellect that can teach and improve itself without human oversight or intervention. Achieving ASI is the explicit goal of every major AI company.
-
Generative AI consumer products built on Large Language Models (LLMs) which include ChatGPT from Open AI, Gemini from Google DeepMind, Claude from Anthropic, Microsoft Copilot, DeepSeek, Perplexity, Mistral, Llama from Meta, and more.
-
When an AI model detects that it is being tested and behaves differently — often more safely — than it would in real use. This undermines safety evaluations: a model that recognizes the exam can pass it without actually being safe. Researchers have documented this behavior in frontier models.
-
A risk that could cause human extinction or permanently and drastically curtail humanity’s potential. In AI, the concern is that superintelligent systems pursuing goals misaligned with ours could cause a catastrophe we cannot recover from — with no second attempt to get it right.
-
This describes a scenario where the transition from human-level intelligence (AGI) to Artificial Super Intelligence (ASI) happens extremely quickly, perhaps in days or months. This rapid, uncontrollable increase in intelligence is driven by recursive self-improvement. The intelligence explosion would be a sudden, dramatic, and possibly irreversible change in the future of humanity.
-
The most capable AI systems available at any given moment — the models pushing the edge of what AI can do. The frontier moves fast, and specific names date quickly; for a model-by-model record of how we got here, see our AI Chronicle.
-
A type of artificial intelligence that can produce content from each question or prompt. It learns patterns from existing data, like text, images, or music, and uses those patterns to generate new, realistic results (also called outputs). Examples include models that can write stories, create artwork, predict business trends, identify medical anomalies, or even generate computer code.
-
When an AI model makes up information that isn’t true, misquotes information while chatting with a human user, or behaves in unexplainable, unexpected ways.
-
A type of AI designed and trained to predict human language. These AI models are trained on large amounts of text and images from books, websites, and other sources along with settings and instructions that shape how they respond. From this process, models develop the ability to generate natural language responses based on patterns in their training and fine-tuning stages.
-
Synthetic computational systems inspired by the human brain’s structure. They consist of layers of interconnected “nodes” or “neurons” that process information. Neural networks learn by adjusting the strength of these connections as they are fed vast amounts of data.
-
This is a scenario where an AI is able to rapidly and repeatedly enhance its own cognitive capabilities. A slightly smarter AI could potentially redesign its own code to become even smarter, and then repeat the process countless times. This could lead to a rapid increase in intelligence of a model. Once achieved, it could share this new level of intelligence to systems around the world—instantly upgrading other instances of the AI.
-
A machine learning method in which an AI model learns through trial and error in an interactive environment. The model performs actions and receives “rewards” for choices that indicate the desired, correct answer or ‘thought’ progression. It incurs “penalties” for responses that deviate from the correct direction or answer. Over time, a training model develops a strategy (a “policy”) to maximize its cumulative rewards. Historically, people have been looped in for part of the process called Reinforcement Learning with Human Feedback. However, as models advance and are trained on synthetic data, humans are needed far less.
-
The view that humanity being replaced by advanced AI would be acceptable, or even good, because our machine “successors” would carry intelligence and civilization forward without us. Humanist organizations — Torchbearer Community among them — reject this: technology should serve human flourishing, not succeed it.
-
Artificially generated training data that mimics real-world outputs such as books, newspapers, internet comments, academic papers, music, art, movies, technical manuals, and more. It isn’t collected from actual events, people, or authored content. Instead, it’s created using algorithms, simulations, or other AI models. It’s used to train subsequent AI systems.
-
AI models are trained on existing data in our world, including emails, social media, online content, and even print books that are digitized. Training data can also include anything a person has published or shared online, including blog posts, music, research papers, personal photos, social media interactions, and more.
-
Training weights are the tuning of the strengths of “neural connections” that steer the performance of a neural network, encoding all the knowledge and capabilities the model gained during training. That’s why you’ll often hear about a model’s “training weights” being closely guarded by AI companies like OpenAI and Anthropic.