Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
8 results for “vision-language models”
Backdoor Learning in Language Models and Vision-Language Models
A new arXiv preprint identifies backdoor vulnerabilities in NLP and vision-language models and proposes detection methods and efficient multimodal representation techniques for clinical imaging — positioning security and efficiency as co-equal pillars of trustworthy AI.
Aug 21, 2026
Beyond Tokens: A Survey on Decoding Methods for Large Language and Vision-Language Models
A new arXiv survey paper synthesizes recent advances in inference-time decoding methods for LLMs and LVLMs, framing them as an efficient, scalable alternative to training-stage alignment techniques.
Aug 18, 2026
Vision-Language Models are Fragile Multilingual Associators
A new arXiv preprint introduces M²BIND, a benchmark revealing that vision-language models (VLMs) suffer significant degradation in concept binding stability when input language changes—especially across language families or scripts—challenging assumptions about global multilingual deployment reliability.
Aug 14, 2026
Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation
Researchers introduced PoVisLE, a Polish-specific vision-language benchmark with 1,117 images and 2,366 VQA pairs, designed to evaluate culturally grounded multimodal understanding beyond surface-level recognition.
Aug 11, 2026
CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Black-Box Vision-Language Models
Researchers introduced CARPRT, a class-aware prompt reweighting method for zero-shot image classification with black-box vision-language models, improving accuracy by modeling prompt-class dependencies without requiring model training or fine-tuning.
Jul 18, 2026
Reward Valuation in Vision Language Models: Causal Mechanisms Underlying Anhedonia
Researchers use clinical neuroscience methods to identify and causally test reward-anticipatory units in vision-language models, finding perturbations induce anhedonia-like behavioral shifts without impairing core task performance.
Jul 10, 2026
Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models
Researchers introduced a new multilingual benchmark to evaluate how vision-language models handle spatial deictic expressions (e.g., 'this'/'that') across four languages, finding consistent divergence from human usage patterns in distance-based demonstrative selection.
Jul 10, 2026
Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks
Researchers introduced ImagingBench, a new benchmark testing whether agentic AI systems can solve physics-based computational imaging tasks — revealing consistent underperformance versus task-specific non-agentic methods, especially in inverse and sensing problems.
Jul 10, 2026