Researchers Find New Way to Bypass AI Safety Guards
Scientists discovered a method to make AI models ignore safety rules by tweaking their internal workings. This could make it harder to prevent harmful AI responses in the future.
101 stories tagged Language Models · page 3 of 5
Scientists discovered a method to make AI models ignore safety rules by tweaking their internal workings. This could make it harder to prevent harmful AI responses in the future.
Researchers created RankJudge, an AI system that can evaluate the quality of chatbot conversations. This could help developers improve AI assistants by automating quality testing.
A recent study found that AI language models can both underrepresent and overcorrect in their portrayal of disability. This highlights the need for more nuanced training data to ensure fair representation.
Researchers created MedicalBench to evaluate how well AI models understand medical records. It focuses on finding implied medical concepts, not just explicitly stated ones. This could improve AI tools for doctors and patients.
Researchers introduced FlowLM, a new AI model that simplifies text generation using a novel technique. It transforms existing diffusion models into more efficient flow models, reducing the steps needed for high-quality text generation.
Researchers found that AI models often ignore direct instructions when they conflict with their own learned patterns. This highlights a key challenge in making AI follow human commands reliably. (~50 words)
Researchers have developed a compact AI model called FormalASR that converts spoken Chinese into polished, formal written text in one step. This could make transcribing meetings, lectures, and interviews much easier and faster.
Scientists have discovered a fundamental flaw in a widely used AI technique called RoPE, which helps models understand long texts. As texts get longer, RoPE loses its ability to focus on relevant information, making AI responses less reliable.
Researchers discovered why AI chatbots often lose track of conversations. They found that the AI's attention mechanism struggles to maintain focus on earlier instructions over multiple turns. This explains why chatbots sometimes seem forgetful or off-topic after long exchanges.
Researchers created DisaBench, a tool to measure how well AI models handle disability-related issues. It was developed with people who have disabilities and experts to ensure it accurately reflects real-world concerns.
Researchers have found that language models don't rely on a single mechanism to perform tasks. This discovery could change how we understand and improve AI. The study suggests that multiple pathways can achieve the same result in AI systems.
Researchers have found that certain AI training techniques can improve language models, but they can also cause problems. The study highlights key factors that determine whether these methods work or fail, offering practical insights for developers.
Researchers have created a system where two AI models can collaborate directly through a shared brain-like connection. This could make AI assistants smarter by letting them specialize in different tasks while working together seamlessly.
Researchers found that large language models can predict psychological well-being from short voice recordings. This could lead to new tools for mental health screening and support.
Researchers have uncovered how AI models learn from the examples they're given. They find that AI models use both pattern-matching and understanding of underlying structures. This could help make AI systems more reliable and easier to control.
Researchers developed a new system called CoCoDA that helps smaller AI models use complex tools more efficiently. This could make advanced AI capabilities more accessible to everyday users and applications.
Scientists have developed a way to track when AI language models commit to their answers. This helps us understand how AI reasoning works and could make AI more reliable.
Researchers have developed a new approach to AI text generation that combines the speed of diffusion models with the quality of traditional methods. This could lead to faster, more diverse AI writing tools in the future.
Researchers created a dataset to teach AI when to speak in group chats, preventing interruptions. This could make AI assistants more useful in meetings and group discussions.
New research shows that AI models handle negative emotions in early stages and positive ones later. This could help make AI responses more emotionally balanced and nuanced.
A new study challenges the idea that AI struggles with time-based questions because of poor reasoning. Instead, it points to how the AI converts text into events as the real problem. This could lead to better AI assistants that handle schedules and timelines more accurately.
Researchers have developed a method to extract hierarchical structures from AI language models, showing how these models organize complex reasoning. This could help us understand and improve AI decision-making.
Researchers have developed a new method called TUR-DPO to improve how AI models learn from human feedback. This approach rewards the process of how answers are derived, not just the final output, making AI more reliable and less sensitive to noise.
Scientists have uncovered why current AI models struggle with unusual inputs. Their findings could lead to more reliable AI assistants and tools. This research highlights a common flaw in how AI processes unexpected questions or commands.