
New AI Defense Detects Hidden Malicious Intent in Conversations
Researchers have developed a method to spot harmful intentions hidden in multi-turn AI conversations. This helps prevent AI models from being tricked into harmful behavior over time.
1035 stories curated by AInformed · page 30 of 44

Researchers have developed a method to spot harmful intentions hidden in multi-turn AI conversations. This helps prevent AI models from being tricked into harmful behavior over time.

Researchers created a dataset to teach AI when to speak in group chats, preventing interruptions. This could make AI assistants more useful in meetings and group discussions.

Researchers found that AI models can make worse predictions when given accurate context. This happens because the models sometimes ignore good information. The study highlights a hidden flaw in how AI systems process data.

New research shows that AI models handle negative emotions in early stages and positive ones later. This could help make AI responses more emotionally balanced and nuanced.

Researchers developed AdaGATE, a new method to help AI answer complex questions that require multiple steps. It improves accuracy by selecting the most relevant information and filling in gaps automatically.

Fine-tuning AI models on even small amounts of harmless data can erase safety measures learned from much larger datasets. Researchers have identified a key mechanism behind this safety degradation, offering a way to predict and prevent it.

Current AI safety tests focus on models in isolation, but a new study warns this doesn't prove real-world safety. The research argues we need to test AI in actual use cases, not just lab settings.

Researchers have developed a method called SWAN that embeds hidden watermarks in the meaning of sentences, not just the words. This could help track AI-generated text more effectively than current methods.
A new AI system called SensingAgents improves activity tracking using wearable sensors. It overcomes common challenges in recognizing daily movements like walking or running. The system could make fitness trackers and health monitoring devices more accurate and reliable.

Researchers have developed a new AI assistant called Pro²Assist that can proactively help with multi-step tasks, like cooking or assembling furniture. Unlike current assistants, it tracks your progress and predicts what you'll need next, making it more helpful for complex activities.
Researchers have developed a new reinforcement learning technique called Adaptive Power-Mean Policy Optimization (APMPO) that improves how AI models reason. This method adapts to the evolving capabilities of large language models, making them more effective at problem-solving.

Researchers have developed ANDRE, a new AI system that extracts logical rules from data more effectively than previous methods. This could make AI systems more interpretable and reliable in real-world, uncertain situations.

Researchers developed a method called PARSE that makes AI responses faster by checking multiple parts of the answer at once. This could lead to quicker, more efficient AI interactions for everyday users.

Researchers have developed a new AI memory system called Lossless Context Management (LCM) that handles long texts better than Claude Code. This could make AI assistants more reliable for tasks requiring large amounts of information.

Researchers have improved a technique called Emphatic TD (ETD) to make AI learning faster and more stable. This could help AI systems learn more efficiently from real-world experiences.

Researchers have developed a framework to better detect and prevent AI-generated medical misinformation. This could make medical AI tools more reliable for everyday users.

Researchers created a new test to measure how well AI systems understand cause and effect in messy, real-world data. This could help improve AI's ability to make better decisions in uncertain situations.

Researchers developed a new AI algorithm called FREIA that helps large language models improve their reasoning skills on their own. This could lead to smarter AI assistants that learn and adapt without constant human supervision.

Researchers have developed an AI system that can automatically extract data from scientific literature, including text, tables, and figures. This could revolutionize materials science by making it easier to build comprehensive databases.

A new study challenges the idea that AI struggles with time-based questions because of poor reasoning. Instead, it points to how the AI converts text into events as the real problem. This could lead to better AI assistants that handle schedules and timelines more accurately.

A new study found that AI models' moral judgments are mostly the same whether they respond instantly or take time to 'think'. The differences that do exist are concentrated in particularly tricky scenarios. This suggests that AI reasoning modes may not drastically alter ethical decisions.

Researchers have developed a new AI system that tracks and models the interactions of surgical teams in real time. This could help improve communication and coordination during operations, making surgeries safer.

A new study finds that AI models often produce misleading information when analyzing conflict data in West Africa. This raises concerns about their reliability for humanitarian efforts. Researchers tested both general and specialized AI models to see how well they could classify conflict events in Nigeria and Cameroon. The results show that open-source models tend to produce more false or misleading information than models specifically trained on African conflict data.

Researchers found that deep AI models can perform deductive reasoning nearly as well as models that follow step-by-step logic. This could make AI smarter without needing extra instructions.