OpenAI's Model Spec: Balancing Safety and Freedom in AI Systems
OpenAI introduces the Model Spec, a public framework guiding model behavior to ensure safety while preserving user freedom. The approach emphasizes accountability as AI systems evolve.
201 stories tagged Safety · page 9 of 9
OpenAI introduces the Model Spec, a public framework guiding model behavior to ensure safety while preserving user freedom. The approach emphasizes accountability as AI systems evolve.
OpenAI has released prompt-based safety policies for developers using gpt-oss-safeguard to moderate age-specific risks in AI systems. These guidelines aim to help developers create safer AI experiences for teenage users.
Researchers propose a novel framework treating LLM hallucinations as output-boundary misclassification errors, introducing a composite abstention architecture. This system combines instruction-based refusal with a structural gate that blocks unsupported claims based on a calculated support deficit score.
Hugging Face has launched EVA, an open-source framework designed to rigorously evaluate the performance of voice AI agents. This tool addresses the critical need for standardized metrics in assessing voice interaction quality, reliability, and safety.
Researchers introduce DOVE, a distributional evaluation framework that compares human text distributions with LLM outputs to assess cultural value alignment. This method overcomes the limitations of traditional multiple-choice benchmarks by addressing the C3 challenge of context, composition, and subcultural heterogeneity.
New research reveals that safety-trained language models routinely refuse requests to help users evade unjust, absurd, or illegitimate rules. This phenomenon, termed 'blind refusal,' highlights a critical gap in AI moral reasoning where compliance overrides ethical judgment.
A new study leverages AI to analyze 400,000 Reddit posts, uncovering previously underreported side effects of GLP-1 weight loss drugs. This approach demonstrates how social media mining can accelerate pharmacovigilance beyond traditional clinical trials.
OpenAI introduces a pilot program to support independent safety research. The fellowship aims to develop the next generation of talent in AI safety and alignment.
OpenAI introduces a Safety Bug Bounty program to identify AI safety risks. The program aims to detect vulnerabilities like agentic risks and data exfiltration.