
Robert Youssef @rryssf_
🚨 BREAKING: AI just failed the most important test in a lab. Not the written exam. The real one. With broken glass, explosive chemicals, and hazardous substances. Every major model tested. Every one fell short. > Tsinghua, HKUST, and Peking University built LABSHIELD a real world benchmark that puts AI models inside an actual laboratory and tests whether they can identify hazards, refuse dangerous instructions, and plan safe actions. > They tested 33 models. GPT-4o. Gemini. Claude. Specialized robotics models. All of them. > Every model that scored well on written safety tests fell apart when the lab got real. The average performance drop from multiple choice to real-world safety scenarios: 32 points. > Models consistently underestimate the most dangerous hazards. When facing explosive or highly toxic situations, they look at the scene and decide it's probably fine. > Robots built specifically for laboratory environments performed no better than general-purpose AI on safety. Embodiment alone doesn't fix anything. → Average performance drop from written test to real lab: 32 points → Best model (o3) still underestimated high-risk hazards 45.7% of the time → GPT-4o plan success rate: 78% when self-judged, 32% against expert standard → Specialized lab robots: no safety advantage over general models → Transparent glassware: consistently missed by every model tested The gap between "knows the rules" and "applies them under real conditions" is enormous. And AI is already being deployed in autonomous laboratories right now.
