Apple ML Research presents a systematic study of preference alignment in Multimodal Large Language Models (MLLMs). The paper categorizes alignment algorithms into offline (e.g., DPO) and online (e.g., online-DPO) methods, finding that combining both can improve performance. It reviews existing multimodal preference datasets and their construction impact, then introduces Bias-Driven Hallucination Sampling (BDHS), a novel method for generating multimodal preference data without additional annotation or external models. BDHS achieves competitive results across multiple benchmarks compared to prior alignment approaches.
140 Impressions