Security researchers from Xidian University have developed CrossMPI, a novel image-based prompt injection attack targeting large vision-language models (LVLMs). Unlike text-based prompt injection, CrossMPI uses nearly imperceptible pixel-level image perturbations to alter how multimodal AI systems interpret both visual and textual inputs simultaneously. The attack achieved a 66.36% average success rate across tested models including MiniGPT4, BLIP-2, and Qwen2.5-VL, outperforming prior methods by ~41 percentage points, and demonstrated strong black-box transferability. Tested defenses like SmoothVLM reduced success rates below 5% in some scenarios but none fully neutralized the attack. The findings raise concerns for enterprises deploying multimodal AI in document processing, autonomous agents, and content moderation workflows.
120 Impressions1 Comment