Apple ML Research presents a systematic study of preference alignment in Multimodal Large Language Models (MLLMs). The paper categorizes alignment algorithms into offline (e.g., DPO) and online (e.g., online-DPO) methods, finding that combining both can improve performance. It reviews existing multimodal preference datasets and their construction impact, then introduces Bias-Driven Hallucination Sampling (BDHS), a novel method for generating multimodal preference data without additional annotation or external models. BDHS achieves competitive results across multiple benchmarks compared to prior alignment approaches.

2m read timeFrom machinelearning.apple.com
Post cover image
140 Impressions