LLMs are trained on next-token prediction, but the real value isn't the generative output — it's the internal representations learned during pre-training. The post questions whether generative modeling is even necessary, pointing to non-generative approaches like joint-embedding (Siamese networks) that learn powerful representations by enforcing similarity between semantically equivalent inputs, without reconstruction. This idea, explored since the 90s, suggests alternative training signals could yield equally powerful internal representations for building intelligent systems.

1m watch time
343 Impressions