#safety

Safety

201 stories tagged Safety · page 4 of 9

How Trees Teach Us About AI Safety
general

How Trees Teach Us About AI Safety

A LessWrong article argues that trees, which are mostly air, offer a counterintuitive lesson for AI safety: true safety comes from structural resilience, not merely from efficiency or sustainability. The draft oversimplifies this by focusing on resource efficiency rather than the article's core point about failure modes and robustness.

GPT-5.6: What You Need to Know About the Latest AI Model
general

GPT-5.6: What You Need to Know About the Latest AI Model

Zvi Mowshowitz has published a detailed system card analysis of GPT-5.6, exploring its capabilities, safety features, and potential implications. The piece examines improvements in reasoning, coding, and reduced refusal rates, while also noting ongoing concerns around alignment and evaluation transparency.

via Hacker News AI#AI#OpenAI#GPT 5.6
OpenAI Joins Effort to Create Global AI Safety Standards
policy

OpenAI Joins Effort to Create Global AI Safety Standards

OpenAI is collaborating with other organizations to develop shared standards for advanced AI through the Appia Foundation. This includes frameworks for evaluating AI safety, promoting best practices, and fostering global cooperation. The goal is to ensure AI is developed and used responsibly for the benefit of all.

via OpenAI Blog#Safety#OpenAI
New Study Compares AI Refusal Steering Techniques for Safer Chat Models
research

New Study Compares AI Refusal Steering Techniques for Safer Chat Models

Researchers compared two methods for steering refusal in AI chat models: Diff-in-Means (DiM) and Iterative Nullspace Projection (INLP). The study examined five open-weight models to see if INLP can match DiM effectiveness in controlling refusal behavior, using interventions like activation addition, directional ablation, nullspace projection, and counterfactual flipping. This could lead to more robust and steerable safety mechanisms in future AI assistants.

via ArXiv cs.AI#AI#Safety#Research