<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/new-open-source-model-turns-a-photo-and-audio-clip-into-a-talking-video-t2vhbhl1r" -->

---
title: New open-source model turns a photo and audio clip into...
description: LongCat-Video is a 13.6B parameter open-source foundational video generation model supporting text-to-video, image-to-video, and video-continuation tasks, with...
canonical: https://daily.dev/posts/new-open-source-model-turns-a-photo-and-audio-clip-into-a-talking-video-t2vhbhl1r
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: New open-source model turns a photo and audio clip into a talking video | daily.dev
og:description: LongCat-Video is a 13.6B parameter open-source foundational video generation model supporting text-to-video, image-to-video, and video-continuation tasks, with...
og:url: https://daily.dev/posts/new-open-source-model-turns-a-photo-and-audio-clip-into-a-talking-video-t2vhbhl1r
og:image: https://api.daily.dev/og/posts/T2VHbhL1r.png
og:image:alt: New open-source model turns a photo and audio clip into a talking video
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# New open-source model turns a photo and audio clip into a talking video

**[Build With GenAI](https://daily.dev/sources/buildwithgenai)** · [@chienvu62](https://daily.dev/chienvu62) · 4 upvotes · 2 comments

## Summary

LongCat-Video is a 13.6B parameter open-source foundational video generation model supporting text-to-video, image-to-video, and video-continuation tasks, with a focus on efficient long-video generation at 720p/30fps using coarse-to-fine generation and Block Sparse Attention. The repository also includes LongCat-Video-Avatar-1.5, an upgraded audio-driven avatar generation model that switches from Wav2Vec2 to Whisper-Large-v3 for improved lip sync, adds stylized domain generalization, multi-stream audio support, and accelerates inference to 8 steps via step distillation. Installation instructions, model downloads, inference scripts for multiple tasks, and MOS evaluation comparisons against Veo3, PixVerse-V5, and Wan 2.2 are provided.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://github.com/meituan-longcat/LongCat-Video>

## Community discussion

Top comments from developers on daily.dev.

**@chienvu62** · 0 upvotes

> should try

## Similar posts on daily.dev

- [Don’t just attend KubeCon \+ CloudNativeCon, Merge Forward your experience\!](https://daily.dev/posts/don-t-just-attend-kubecon-cloudnativecon-merge-forward-your-experience--l0rpp73x8) · CNCF · 0 upvotes · 0 comments
- [Announcing H2 2026 KCDs](https://daily.dev/posts/announcing-h2-2026-kcds-m96goajm1) · CNCF · 1 upvotes · 0 comments
- [Two months of Open Community Groups](https://daily.dev/posts/two-months-of-open-community-groups-asf52zhbs) · CNCF · 0 upvotes · 0 comments
- [CNCF Unveils Schedule for KubeCon \+ CloudNativeCon Europe 2026](https://daily.dev/posts/cncf-unveils-schedule-for-kubecon-cloudnativecon-europe-2026-ikhcoa5cb) · CNCF · 2 upvotes · 0 comments
- [CNCF Debuts KubeCon \+ CloudNativeCon Japan 2026 Schedule](https://daily.dev/posts/cncf-debuts-kubecon-cloudnativecon-japan-2026-schedule-xp5pyudub) · CNCF · 1 upvotes · 0 comments

---

Tags: [#python](https://daily.dev/tags/python), [#video-generation](https://daily.dev/tags/video-generation), [#diffusion-models](https://daily.dev/tags/diffusion-models)

[View this post on daily.dev](https://daily.dev/posts/new-open-source-model-turns-a-photo-and-audio-clip-into-a-talking-video-t2vhbhl1r)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"DiscussionForumPosting","mainEntityOfPage":"https://daily.dev/posts/new-open-source-model-turns-a-photo-and-audio-clip-into-a-talking-video-t2vhbhl1r","headline":"New open-source model turns a photo and audio clip into a talking video","text":"Shared: GitHub - meituan-longcat/LongCat-Video","url":"https://daily.dev/posts/new-open-source-model-turns-a-photo-and-audio-clip-into-a-talking-video-t2vhbhl1r","datePublished":"2026-08-22T00:07:29.330Z","dateModified":"2026-08-22T00:08:45.211Z","author":{"@type":"Person","name":"Chien Vu","url":"https://daily.dev/chienvu62","image":"https://lh3.googleusercontent.com/a/ACg8ocJnxXeduN9blBMay28NH-zcHaagBsmK1-9fY7M4ff4_eUiE-miT=s96-c","description":"Machine learning researcher | Ph.D","interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"EndorseAction"},"userInteractionCount":20875}},"image":"https://media.daily.dev/image/upload/s--ZrL_HSsR--/f_auto/v1722860399/public/Placeholder%2006","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":4},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":2}],"sharedContent":{"@type":"WebPage","url":"https://api.daily.dev/r/VWoHdOxex"},"comment":[{"@type":"Comment","text":"should try","datePublished":"2026-08-25T01:59:32.542Z","url":"https://daily.dev/posts/T2VHbhL1r#c-RQg7EZJjq","author":{"@type":"Person","name":"Chien Vu","url":"https://daily.dev/chienvu62","image":"https://lh3.googleusercontent.com/a/ACg8ocJnxXeduN9blBMay28NH-zcHaagBsmK1-9fY7M4ff4_eUiE-miT=s96-c"}}],"isPartOf":{"@type":"WebPage","url":"https://daily.dev/squads/buildwithgenai","name":"Build With GenAI"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Build With GenAI","item":"https://daily.dev/squads/buildwithgenai"},{"@type":"ListItem","position":3,"name":"New open-source model turns a photo and audio clip into a talking video"}]}
```

