Security researchers from Xidian University have developed CrossMPI, a novel image-based prompt injection attack targeting large vision-language models (LVLMs). Unlike text-based prompt injection, CrossMPI uses nearly imperceptible pixel-level image perturbations to alter how multimodal AI systems interpret both visual and textual inputs simultaneously. The attack achieved a 66.36% average success rate across tested models including MiniGPT4, BLIP-2, and Qwen2.5-VL, outperforming prior methods by ~41 percentage points, and demonstrated strong black-box transferability. Tested defenses like SmoothVLM reduced success rates below 5% in some scenarios but none fully neutralized the attack. The findings raise concerns for enterprises deploying multimodal AI in document processing, autonomous agents, and content moderation workflows.

5m read timeFrom csoonline.com
Post cover image
120 Impressions1 Comment