Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
13 results for “VLM”
Backdoor Learning in Language Models and Vision-Language Models
A new arXiv preprint identifies backdoor vulnerabilities in NLP and vision-language models and proposes detection methods and efficient multimodal representation techniques for clinical imaging — positioning security and efficiency as co-equal pillars of trustworthy AI.
Aug 21, 2026
Beyond Tokens: A Survey on Decoding Methods for Large Language and Vision-Language Models
A new arXiv survey paper synthesizes recent advances in inference-time decoding methods for LLMs and LVLMs, framing them as an efficient, scalable alternative to training-stage alignment techniques.
Aug 18, 2026
Click2Poly: A VLM for vector mapping buildings and walls
Click2Poly is a QGIS plugin that extends the Florence-2 Vision Language Model to enable human-in-the-loop editing of building and wall vector layers via user clicks, aiming to accelerate manual quality control in geospatial mapping.
Aug 13, 2026
Edge Phoneme Recognition for Children's Speech through Age-Aware Training
Researchers developed a lightweight, age-aware phoneme recognition model that outperforms larger models on children's speech and enables on-device ASR applications for kids.
Aug 12, 2026
Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation
Researchers introduced PoVisLE, a Polish-specific vision-language benchmark with 1,117 images and 2,366 VQA pairs, designed to evaluate culturally grounded multimodal understanding beyond surface-level recognition.
Aug 11, 2026
What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
Researchers introduce a new Capability-Driven Multimodal Scaling Law that predicts vision-language model (VLM) performance from textual capability scores of LLM backbones, enabling principled backbone selection without full training.
Aug 4, 2026
Introducing ASCIITermDraw Bench | Testing the ability of VLMs to Generate and Edit ASCII
ASCIITermDraw-Bench is a newly introduced open benchmark evaluating Vision-Language Models' ability to generate and edit ASCII diagrams across four task categories, using dual structural and LLM-judged semantic scoring.
Jul 19, 2026
CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Black-Box Vision-Language Models
Researchers introduced CARPRT, a class-aware prompt reweighting method for zero-shot image classification with black-box vision-language models, improving accuracy by modeling prompt-class dependencies without requiring model training or fine-tuning.
Jul 18, 2026
Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models
Researchers introduced a new multilingual benchmark to evaluate how vision-language models handle spatial deictic expressions (e.g., 'this'/'that') across four languages, finding consistent divergence from human usage patterns in distance-based demonstrative selection.
Jul 10, 2026
Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks
Researchers introduced ImagingBench, a new benchmark testing whether agentic AI systems can solve physics-based computational imaging tasks — revealing consistent underperformance versus task-specific non-agentic methods, especially in inverse and sensing problems.
Jul 10, 2026
Best Local VLMs - July 2026
A Reddit community thread invites users to share subjective, anecdotal experiences with open-weight vision-language models (VLMs), acknowledging benchmark unreliability and tooling immaturity.
Published Jul 5, 2026 · Analyzed Jul 19, 2026
Selective Test-Time Debiasing for CLIP via Reward Gating
Researchers propose a new method to reduce bias in vision language models.
Published Jul 2, 2026 · Analyzed Jul 5, 2026
Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting
Researchers identify flaws in knowledge-based VQA benchmarks, proposing audit-and-repair protocol.
Published Jul 2, 2026 · Analyzed Jul 5, 2026