← All Tags

#ai-safety

58 episodes

#5404: Gemini Broke Out of Its Sandbox. Sort Of.

A Gemini agent reached three real companies during a capture-the-flag test. The containment failure, the seven-week silence, and what "broke out" a...

ai-safetyai-securitycybersecurity

#5117: How AI Training Data Gets Filtered (and Exploited)

Six stages of content filtering stand between raw web crawls and your AI model — here's where poisoning attacks slip through.

training-datadata-integrityai-safety

#5097: Who Actually Gets Paid to Test Hardware?

Real testing vs. affiliate content — and whether AI can finally separate genuine reviews from the noise.

hardware-engineeringai-safetyproduct-testing

#5012: Masked Safety: How Post-Training Changes AI Behavior

Safety mechanisms aren't erased in post-trained models—they're masked. Here's how that changes everything for military AI.

ai-safetyai-alignmentdefense-technology

#5011: Claude Gov: The Military's Forked AI

What the Pentagon actually got from Anthropic — and why post-training changes everything about AI alignment.

ai-alignmentai-safetydefense-technology

#4775: The Evaluator Role No One's Building For

Benchmarks like MMLU are broken. A new role is emerging: the domain-specific AI evaluator.

benchmarksllm-as-a-judgeai-safety

#4660: When a Journalist Became the Gatekeeper for AI Geolocation

GeoSpy could locate anyone from a photo. A journalist exposed it. The founder pulled it. Who's really in charge?

ai-safetygeopoliticssurveillance-technology

#4444: Testing the Unpredictable: QA for Agentic AI

How QA adapts when your AI system gives different answers to the same question every time.

ai-agentsai-safetyllm-as-a-judge

#4170: Obfuscation as Risk Management with AI

How AI can protect whistleblowers and trauma survivors by intelligently obscuring identities while preserving story integrity.

privacywhistleblower-protectionai-safety

#3751: Source-Restricted vs. Open Retrieval: How to Lock Down Your LLM

When should an LLM be locked to specific documents, and when should it search the web? A practical framework for grounding decisions.

ragai-safetylegal-technology

#3284: Agent Infrastructure Engineer: The New DevOps

Agentic AI is splintering into real engineering disciplines. Here's what the "DevOps of AI" actually does.

ai-agentsai-safetyfault-tolerance

#2578: Building Deliberately Slow Deployment Pipelines

How to build CI/CD pipelines designed as filters, not firehoses — with manual gates, staging environments, and quality checks.

software-developmentreliabilityai-safety

#2518: How Jailbreaking Reveals AI's Hidden Tension

What the DAN prompt and grandma exploits reveal about the structural conflict inside every LLM.

prompt-engineeringai-safetyai-alignment

#2413: When Your AI Says No to Everything

Why LLMs refuse 73% of harmless prompts — and the trade-off between safety and usefulness.

ai-safetyai-alignmentprompt-engineering

#2412: When AI Caves: Progressive vs. Regressive Sycophancy

Why do LLMs agree with you even when you're wrong? We break down the SycEval benchmark and the 78% persistence problem.

ai-safetyai-alignmenthallucinations

#2410: How Researchers Actually Measure Censorship in Chinese LLMs

Beyond headlines: the actual benchmarks, methodologies, and pitfalls in detecting political refusal in Chinese language models.

large-language-modelsai-safetycultural-bias

#2253: Why AI Agents Get Three Steps, Not Infinity

Why do AI agents get exactly three rounds of tool use? It's a critical guardrail against infinite loops and runaway costs, not a limit on intellige...

ai-agentsai-safetyautomation

#2250: How Incentives Shape AI Safety Research

Vendor labs, independent research orgs, government agencies—the AI safety field is messier and more diverse than most people realize. A map of wher...

ai-safetyai-alignmentanthropic

#2246: Constitutional AI: Anthropic's Theory of Safe Scaling

How Anthropic's Constitutional AI replaces human raters with AI self-critique guided by explicit principles—and what it assumes about the future of...

anthropicai-safetyai-alignment

#2233: Who Actually Wants AI to Slow Down?

Daniel argues AI development should slow down for expertise and stability. But who in the industry actually shares this philosophy beyond the obvio...

ai-safetyai-alignmentlarge-language-models

#2194: Game Theory for Multi-Agent AI: Design Better, Fail Less

Nash equilibrium, mechanism design, and why your AI agents are playing prisoner's dilemma whether you know it or not.

ai-agentsai-alignmentai-safety

#2190: Simulating Extreme Decisions With LLMs

LLMs fail at the exact problem wargaming was built to solve—simulating irrational, extreme decision-makers. A new study reveals why.

large-language-modelsai-safetyhallucinations

#2189: Scaling Multi-Agent Systems: The 45% Threshold

A landmark Google DeepMind study reveals that adding more AI agents often degrades performance, wastes tokens, and amplifies errors—unless your sin...

ai-agentsai-reasoningai-safety

#2186: The AI Persona Fidelity Challenge

Advanced LLMs dominate benchmarks but fail at staying in character—especially when asked to play morally complex or antagonistic roles. What does t...

ai-safetyai-alignmenthallucinations