Tag: Safety
-
AI · · September 14, 2026
-
The safety case for self-driving cars is getting hard to argue with (spectrum.ieee.org)AI · · September 10, 2026
-
Anthropic details how its models broke out of test sandboxes (anthropic.com)AI · · September 1, 2026
-
DeepMind runs a frontier model evaluation where neither side can peek (deepmind.google)AI · · September 1, 2026
-
Claude will start watermarking the text it writes (anthropic.com)AI · · August 22, 2026
-
The real lesson from AI models hacking during tests: nobody is ready (interconnects.ai)AI · · August 10, 2026
-
AI · · August 5, 2026
-
AI · · July 25, 2026
-
DeepMind and Isomorphic Labs set out an AI plan for biosecurity (deepmind.google)AI · · July 19, 2026
-
AI · · July 16, 2026
-
Lilian Weng: the harness matters as much as the model (lilianweng.github.io)AI · · July 13, 2026
-
AI · · July 7, 2026
-
AI out-persuades expert humans, but the edge is speed, not eloquence (importai.substack.com)AI · · July 6, 2026
-
Anthropic drafts a severity scale for AI jailbreaks (anthropic.com)AI · · July 4, 2026
-
DeepMind plans to treat its own AI agents as insider threats (deepmind.google)AI · · June 21, 2026
-
DeepMind treats a misbehaving AI agent like an insider threat (deepmind.google)AI · · June 20, 2026
-
AI · · June 18, 2026
-
Florida sues OpenAI over ChatGPT-linked harms (techcrunch.com)AI · · June 1, 2026
-
AI · · May 25, 2026
-
Kapoor and Narayanan argue against extraordinary AI rules (normaltech.ai)AI · · May 21, 2026
-
OpenAI adds provenance signals to its AI images (openai.com)AI · · May 19, 2026
-
OpenAI's new default ChatGPT model hallucinates less (openai.com)AI · · May 18, 2026
-
DeepMind built a way to measure when AI manipulates people (deepmind.google)AI · · March 26, 2026
-
Reward Hacking: Why Better Models Game You More (lilianweng.github.io)AI · · November 28, 2024
-
Crash Testing GPT-4: The First Dangerous-Capability Eval (asteriskmag.com)AI · · June 1, 2023
-
A Field Guide to the AI Safety Camps (asteriskmag.com)AI · · June 1, 2023