<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/how-to-train-multimodal-llms-to-understand-and-interact-with-text-image-video-and-audio-model-and-umgbdauxs" -->

---
title: How To Train Multimodal LLMs To Understand And Interact...
description: The post discusses several multimodal Large Language Models (LLMs) and their capabilities, including KosMos-2, Shikra, GPT-4V, and Gemini.
canonical: https://daily.dev/posts/how-to-train-multimodal-llms-to-understand-and-interact-with-text-image-video-and-audio-model-and-umgbdauxs
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: How To Train Multimodal LLMs To Understand And Interact With Text, Image, Video And Audio: Model And Methods（Continued） | daily.dev
og:description: The post discusses several multimodal Large Language Models (LLMs) and their capabilities, including KosMos-2, Shikra, GPT-4V, and Gemini.
og:url: https://daily.dev/posts/how-to-train-multimodal-llms-to-understand-and-interact-with-text-image-video-and-audio-model-and-umgbdauxs
og:image: https://api.daily.dev/og/posts/umgbdAuxS.png
og:image:alt: How To Train Multimodal LLMs To Understand And Interact With Text, Image, Video And Audio: Model And Methods（Continued）
og:image:width: 1200
og:image:height: 630
og:locale: en
---

[AI in Plain English](https://daily.dev/sources/aiplainenglish)

[Read post](https://api.daily.dev/r/umgbdAuxS)

# [How To Train Multimodal LLMs To Understand And Interact With Text, Image, Video And Audio: Model And Methods（Continued）](https://api.daily.dev/r/umgbdAuxS "Go to post")

The post discusses several multimodal Large Language Models (LLMs) and their capabilities, including KosMos-2, Shikra, GPT-4V, and Gemini.

[#llm](/tags/llm "Check all #llm posts")

Apr 02, 2024•5m read time•From [ai.plainenglish.io](https://api.daily.dev/r/umgbdAuxS "ai.plainenglish.io")

[![Post cover image](https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/aa2754e809566645bd48cb70544c2d60?_a=AQAEufR)](https://api.daily.dev/r/umgbdAuxS "Go to post")

Table of contents

[Kosmos-2](https://api.daily.dev/r/umgbdAuxS?a=0d55 "Kosmos-2")[Shikra](https://api.daily.dev/r/umgbdAuxS?a=8d85 "Shikra")[GPT4V](https://api.daily.dev/r/umgbdAuxS?a=36a2 "GPT4V")[Gemini](https://api.daily.dev/r/umgbdAuxS?a=ad30 "Gemini")

40 Impressions4 Upvotes

Comment

Bookmark

Copy

![Placeholder image for anonymous user](https://media.daily.dev/image/upload/s--qsFuKGv_--/t_logo,f_auto/public/noProfile)Share your thoughtsPost

[![AI in Plain English's image](https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/8a17667bb9684f35ae94c89e2d61e7eb)](https://daily.dev/sources/aiplainenglish)

[AI in Plain English](https://daily.dev/sources/aiplainenglish "https://daily.dev/sources/aiplainenglish")

1.9K Followers

•

628 Upvotes

#### Would you recommend this post?

Copy link

WhatsApp

Facebook

X

New Squad

Copy linkShare with your friends

4

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"How To Train Multimodal LLMs To Understand And Interact With Text, Image, Video And Audio: Model And Methods（Continued）","url":"https://daily.dev/posts/how-to-train-multimodal-llms-to-understand-and-interact-with-text-image-video-and-audio-model-and-umgbdauxs","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/how-to-train-multimodal-llms-to-understand-and-interact-with-text-image-video-and-audio-model-and-umgbdauxs"},"datePublished":"2024-04-02T10:48:31.063Z","dateModified":"2024-05-09T08:43:33.910Z","description":"The post discusses several multimodal Large Language Models (LLMs) and their capabilities, including KosMos-2, Shikra, GPT-4V, and Gemini.","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/aa2754e809566645bd48cb70544c2d60?_a=AQAEufR","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/aa2754e809566645bd48cb70544c2d60?_a=AQAEufR","isAccessibleForFree":true,"articleSection":"AI in Plain English","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"AI in Plain English","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/8a17667bb9684f35ae94c89e2d61e7eb","url":"https://daily.dev/sources/aiplainenglish"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/how-to-train-multimodal-llms-to-understand-and-interact-with-text-image-video-and-audio-model-and-umgbdauxs","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":4},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm","timeRequired":"PT5M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"AI in Plain English","item":"https://daily.dev/sources/aiplainenglish"},{"@type":"ListItem","position":3,"name":"How To Train Multimodal LLMs To Understand And Interact With Text, Image, Video And Audio: Model And Methods（Continued）"}]}
```

