AI Safety
Tag1.7K stories
AI Safety news and updates covering research and practice aimed at making AI systems behave as intended. Readers can learn about alignment and evaluation methods, interpretability research, red teaming and misuse testing, incident reporting, and the policy discussions around deployment.
Meet Inspect: The Latest AI Safety Evaluations Platform Introduced By UK’s AI Safety InstituteU.K. agency releases tools to test AI model safetySnowflake Cortex LLM: New Features & Enhanced AI SafetyMachine Unlearning in 2024Deciphering Transformer Language Models: Advances in Interpretability ResearchAI Safety: Ensuring the Responsible Development of Artificial IntelligenceNIST launches a new platform to assess generative AIThis AI Paper from MLCommons AI Safety Working Group Introduces v0.5 of the Groundbreaking AI Safety BenchmarkImport AI 368: 500% faster local LLMs; 38X more efficient red teaming; AI21's FrankenmodelIndia, grappling with election misinfo, weighs up labels and its own AI safety coalition
Roadmaps
Comprehensive roadmap for AI Safety
By roadmap.sh