UC Berkeley researchers propose a large world model (LWM) that can learn complex representations from images and videos in sequences of up to one million tokens. LWM aims to improve AI's understanding of complex environments by combining video and language.
Table of contents
Inside Large World Model: UC Berkeley Multimodal Model that can Understand 1 Hour Long VideosThe ProblemLarge World Model(LWM)Core ArchitectureThe Results1 Impression