
OSGuard: New Benchmark Tests Safety of AI Desktop Agents
Researchers introduced OSGuard, a new benchmark to test if AI agents complete tasks safely. It checks for risky shortcuts that might bypass security or ethics rules.
1035 stories curated by AInformed · page 15 of 44

Researchers introduced OSGuard, a new benchmark to test if AI agents complete tasks safely. It checks for risky shortcuts that might bypass security or ethics rules.

Researchers found that AI memory and skill modules don't always justify their cost. The study suggests that simpler approaches — such as using the same token budget for additional actor steps — can be just as effective or better for certain web-based tasks.

A new study reveals that current AI models favor languages like English and French, making them less effective and more expensive for speakers of underrepresented languages. Researchers propose new methods to make these models more fair and efficient for everyone.

Researchers created a new test to see how well AI models handle messy, real-world time data. This could improve AI tools that analyze everything from medical sensors to industrial equipment.

Researchers introduced ReportQA, a new AI system that evaluates radiology reports by asking and answering clinically relevant questions. Unlike traditional metrics, ReportQA mimics how doctors use reports for diagnosis, potentially improving the quality and usefulness of AI-generated medical reports.

Researchers found that AI models often waste time overthinking, leading to errors. Their new method stops the AI when it's no longer helping, improving accuracy.

Researchers found that AI models often sound too confident in answers that aren't fully justified. They developed a new method to better align confidence with the quality of explanations. This could make AI assistants more reliable when giving complex answers.

Researchers developed Metric Match, a subset selection method that accurately estimates LLM judge reliability from limited human annotations, potentially reducing the cost of AI evaluation.

Researchers developed an AI system that translates natural language queries into requests for satellite imagery and environmental data. This could make complex geospatial data more accessible to non-experts.

Researchers developed a new AI framework called SERAF that enhances time series forecasting by combining historical patterns with semantic context to address non-stationarity. This approach could make predictions more accurate for real-world applications like stock markets and weather forecasting.

Researchers introduced Nemotron 3 Ultra, a massive AI model with 550 billion total parameters and 55 billion active parameters. It uses advanced techniques like a hybrid Mamba-Transformer architecture, Mixture-of-Experts, and Multi Token Prediction to handle long texts and complex reasoning tasks more efficiently than ever before.

Researchers have developed new AI models that respond instantly while maintaining strong reasoning. These models could make AI tools faster and more capable for everyday users. The models, Ling-2.6 and Ring-2.6, are designed to be efficient and practical to use, potentially improving AI assistants and other tools we interact with daily.

Researchers tested AI tools that turn human-written math proofs into computer code. They found these tools work well with clean, ideal proofs but fail with real-world, messy ones. This highlights a big gap in AI's ability to robustly handle informal mathematics.

Researchers developed a new type of AI model that can reason about unseen combinations of objects by combining causal and relational reasoning. This breakthrough could help AI systems generalize better to new situations.

Researchers introduced AdaMame, a new AI training approach that helps large language models reason better in multiple languages. This could make math-solving AI tools more reliable for non-English speakers.

Researchers introduced YeasierAgent, a system that lets users and AI agents collaborate to build apps. It rethinks how software is created, making it more flexible and social.

Researchers have developed TwinBI, an AI system that bridges the gap between interactive dashboards and natural language queries. By creating a digital twin of the dashboard state, TwinBI ensures filters, metrics, and chart contexts stay synchronized whether you're clicking or typing.

Researchers created Poker Arena, a Texas Hold'em platform to test AI's strategic reasoning and memory. It evaluates nine different cognitive skills, offering deeper insights than traditional benchmarks.

Researchers compared two methods for steering refusal in AI chat models: Diff-in-Means (DiM) and Iterative Nullspace Projection (INLP). The study examined five open-weight models to see if INLP can match DiM effectiveness in controlling refusal behavior, using interventions like activation addition, directional ablation, nullspace projection, and counterfactual flipping. This could lead to more robust and steerable safety mechanisms in future AI assistants.

Researchers have developed a new framework called Orchestra-o1 that coordinates multiple AI agents to work together. This system can handle complex tasks that require different types of information and actions, making it more versatile than current single-agent systems.

Researchers introduced a new AI safety framework called Risk-Aware Causal Gating (RACG) that helps AI models make safer decisions by evaluating potential risks. This approach could prevent costly errors in AI-driven systems by deciding when to act, defer, or abstain from actions.

Researchers developed a system that lets chatbots adjust their responses based on user feedback. This could make AI assistants more personalized without needing extensive pre-training.

Researchers created MA-ProofBench, the first formal theorem-proving benchmark dedicated to Mathematical Analysis. It could help develop smarter AI tutors and research assistants for complex math problems.

Researchers developed a Transformer-based AI model that solves the open shop scheduling problem (OSSP) faster and more efficiently, reducing the need for extensive tuning. Tested on Taillard benchmark instances (4x4, 5x5, 7x7, and 10x10), the approach could help businesses optimize production lines and reduce costs.