NeoMME: Hugging Face's Open-Source AI Model for Multilingual and Multimodal Tasks
Hugging Face released NeoMME, an open-source multimodal-native encoder that processes text, images, and audio across over 100 languages in a single efficient system.
33 stories tagged Multimodal
Hugging Face released NeoMME, an open-source multimodal-native encoder that processes text, images, and audio across over 100 languages in a single efficient system.
Researchers introduced Nemotron 3.5 Content Safety Moderator, a compact 4-billion-parameter vision-language model that jointly classifies and moderates text, images, documents, and screenshots across multiple languages and custom policies, offering a cost-effective alternative to existing guardrails.
Meta has released Muse Glimmer, an open-source AI agent designed to run locally on consumer hardware. It is multimodal, agentic, and customizable, enabling text generation, image creation, and complex workflows without sending data to the cloud.
Researchers have developed a method to predict which visual tokens are most important in multimodal AI models, making them more efficient. This could lead to faster and cheaper AI systems for tasks like image captioning and visual question answering.
Researchers introduced MultivationBench, a benchmark that evaluates how well multimodal AI models understand character motivations across sequential visual narratives, addressing a gap in existing static-text or single-image evaluations.
Researchers introduce CaM-Wolf, the first AI agent for social deduction games that integrates both text and visual perception and generation, enabling more human-like social interaction in games like Werewolf.
Researchers created Infinity-Parser2, an AI that can read and understand documents better than ever. It uses a special training method to handle all kinds of documents, from contracts to reports, in both English and Chinese. This could make document processing faster and more accurate for businesses and everyday users.
Researchers have developed a new unified multimodal foundation model that jointly models vision, language, world dynamics, and action generation. This advancement could lead to more capable robots and virtual assistants that can understand instructions, anticipate environmental changes, and execute precise actions over extended horizons.
Researchers have unveiled Gemma 4, a new suite of open-weight AI models that handle text, images, and audio together. These models are designed to be more efficient and better at reasoning, with sizes ranging from 2.3 billion to 31 billion parameters.
Researchers developed a way for AI assistants to highlight where they found answers in documents. This could make chatbots more trustworthy by showing their sources instantly.
A new paper highlights flaws in how we test AI models that handle text, images, and other inputs together. Current methods miss key aspects like understanding physical reality or combining different types of information. The authors propose better evaluation frameworks to address these gaps.
Researchers developed SciLens, an AI system that can verify scientific claims by analyzing both text and images. This tool could make it easier for scientists and the public to check the accuracy of research findings.
Researchers have traced the internal pathways through which AI models process and combine visual and audio inputs to reach decisions. The findings could lead to more transparent and reliable AI assistants and creative tools.
Researchers developed TIGER, a new AI system that improves accuracy in multimodal generation by tracking and correcting factual errors. This could make AI-generated images and text more reliable for everyday users.
Researchers have developed VFEAgent, an AI system that automates Finite Element Analysis (FEA) from images and text descriptions. This could make advanced engineering tools accessible to non-experts, speeding up design processes.
Hark, a stealthy AI startup, has raised $700M to create a universal AI interface that works across all your apps and devices. The company plans to launch its first multimodal models this summer, followed by custom hardware.
Google’s new Gemini Omni can create and edit videos from text, images, and audio using simple prompts. This could revolutionize how anyone makes video content without needing professional skills.
Researchers have developed a system called LatentRouter that can choose the best AI model for a specific image-based question before it even answers. This could make AI assistants much more efficient and accurate for visual tasks.
NVIDIA's new Nemotron 3 Nano Omni model supports long-context multimodal intelligence across documents, audio, and video. It is designed for developers to build advanced AI agents.
A new study explores source-modality monitoring in vision-language models, assessing their ability to track and communicate the origin of information. The research evaluates how models bind words to specific input components across 11 different models.
Researchers developed CognitiveTwin, a digital twin framework that predicts individual cognitive decline in Alzheimer's disease using multi-modal data. The model aims to provide accurate, fair, and robust predictions across diverse patient demographics.
Researchers developed AITP, an AI system that uses Multimodal Large Language Models to analyze traffic accidents and assign responsibility based on legal knowledge. This advancement could revolutionize accident investigations and insurance claims.
Researchers have introduced the first benchmark for extracting claims from multimodal social media posts, addressing a gap in automated fact-checking. The dataset includes text combined with images like memes and screenshots, challenging traditional methods.
Researchers introduce CFMS, the first fine-grained multimodal sarcasm dataset for Chinese social media. It includes 2,796 image-text pairs with triple-level annotations, advancing research in sarcasm detection.