<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/multimodal-open-d1-decision-models-for-the-edge-howvigy9x" -->

---
title: Multimodal open d1 decision models for the edge | daily.dev
description: Liquid AI released open-weight d1 decision models - d1-3B and the experimental d1-omni-600M - designed for fast, structured decision-making rather than token...
canonical: https://daily.dev/posts/multimodal-open-d1-decision-models-for-the-edge-howvigy9x
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Multimodal open d1 decision models for the edge | daily.dev
og:description: Liquid AI released open-weight d1 decision models - d1-3B and the experimental d1-omni-600M - designed for fast, structured decision-making rather than token...
og:url: https://daily.dev/posts/multimodal-open-d1-decision-models-for-the-edge-howvigy9x
og:image: https://api.daily.dev/og/posts/hOwVIGY9X.png
og:image:alt: Multimodal open d1 decision models for the edge
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Multimodal open d1 decision models for the edge

**[Hugging Face](https://daily.dev/sources/huggingface)** · 5 min read · 0 upvotes · 0 comments

## Summary

Liquid AI released open-weight d1 decision models - d1-3B and the experimental d1-omni-600M - designed for fast, structured decision-making rather than token generation, answering in a single forward pass. d1-3B, built on the LFM2.5-VL-3B vision-language backbone, supports text and images and tops the Decision Index 0.2.1 among sub-10B models, scoring 48.57 and beating a 35B-A3B model. d1-omni-600M, built on a bidirectional encoder, adds audio support alongside text and images. Benchmarks across seven public datasets show d1-3B averaging 82.9, ahead of larger Decider models, while d1-omni-600M scores 78.4 with a quarter of the parameters. Speed tests on NVIDIA Jetson devices and GPUs show single-question latency under 50ms on edge hardware and under 10ms on GPUs. Both models are available on Hugging Face with Python code examples using transformers>=5.14.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://huggingface.co/blog/LiquidAI/open-d1>

## Questions this post answers

### What is Liquid AI's d1-3B decision model and how fast is it on edge devices?

d1-3B is an open-weight multimodal decision model built from the LFM2.5-VL-3B vision-language backbone that answers in a single forward pass instead of generating tokens. It supports text and image inputs, scores 48.57 on the Decision Index 0.2.1 (best under 10B parameters), and answers a question in 16ms on an NVIDIA Jetson AGX Thor, 26ms on a Jetson AGX Orin, and 50ms on a Jetson Orin Nano.

_Teams picking an edge decision model can track benchmarks like this via daily.dev._

### What modalities does Liquid AI's d1-omni-600M model support?

d1-omni-600M is an experimental 600M-parameter decision model built from the LFM2.5-Encoder-350M bidirectional encoder backbone. It accepts either text and image, or text and audio as inputs, making it trimodal overall (text, vision, audio), though it only handles two modalities at once per request. It scores 78.4 mean across seven benchmarks, beating the larger Decider 2B model.

_Developers comparing lightweight multimodal models can follow updates like this on daily.dev._

### What package version of transformers is required to run Liquid AI's d1 decision models?

Running the d1 decision models requires transformers version 5.14 or higher, installed alongside torch, torchvision, and pillow. The models ship their own custom code, so they must be loaded with trust_remote_code=True via the AutoModel.from_pretrained call in the transformers library.

_Anyone setting up a new model dependency can check compatibility notes like this on daily.dev._

## Similar posts on daily.dev

- [Qwen3.6–35B-A3B: The Most Practical Open-Source AI Model Yet?](https://daily.dev/posts/qwen3-6-35b-a3b-the-most-practical-open-source-ai-model-yet--eihfhk9n5) · Faun · 77 upvotes · 6 comments
- [Small Language Models: Edge AI Innovation From AI21](https://daily.dev/posts/small-language-models-edge-ai-innovation-from-ai21-qwekhtizt) · IEEE Spectrum · 1 upvotes · 0 comments

---

Tags: [#ai](https://daily.dev/tags/ai), [#deep-learning](https://daily.dev/tags/deep-learning), [#local-ai](https://daily.dev/tags/local-ai)

[View this post on daily.dev](https://daily.dev/posts/multimodal-open-d1-decision-models-for-the-edge-howvigy9x)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Multimodal open d1 decision models for the edge","url":"https://daily.dev/posts/multimodal-open-d1-decision-models-for-the-edge-howvigy9x","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/multimodal-open-d1-decision-models-for-the-edge-howvigy9x"},"datePublished":"2026-10-07T16:57:16.597Z","dateModified":"2026-10-08T00:07:32.859Z","description":"Liquid AI released open-weight d1 decision models - d1-3B and the experimental d1-omni-600M - designed for fast, structured decision-making rather than token...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/3da07202ba33525118c1e397e53a1cff?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/3da07202ba33525118c1e397e53a1cff?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Hugging Face","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Hugging Face","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/f1f55c67d81a4330acf5b90b26b0c8e1","url":"https://daily.dev/sources/huggingface"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/multimodal-open-d1-decision-models-for-the-edge-howvigy9x","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai,deep-learning,local-ai","timeRequired":"PT5M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Hugging Face","item":"https://daily.dev/sources/huggingface"},{"@type":"ListItem","position":3,"name":"Multimodal open d1 decision models for the edge"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/multimodal-open-d1-decision-models-for-the-edge-howvigy9x#faq","mainEntity":[{"@type":"Question","name":"What is Liquid AI's d1-3B decision model and how fast is it on edge devices?","acceptedAnswer":{"@type":"Answer","text":"d1-3B is an open-weight multimodal decision model built from the LFM2.5-VL-3B vision-language backbone that answers in a single forward pass instead of generating tokens. It supports text and image inputs, scores 48.57 on the Decision Index 0.2.1 (best under 10B parameters), and answers a question in 16ms on an NVIDIA Jetson AGX Thor, 26ms on a Jetson AGX Orin, and 50ms on a Jetson Orin Nano. Teams picking an edge decision model can track benchmarks like this via daily.dev."}},{"@type":"Question","name":"What modalities does Liquid AI's d1-omni-600M model support?","acceptedAnswer":{"@type":"Answer","text":"d1-omni-600M is an experimental 600M-parameter decision model built from the LFM2.5-Encoder-350M bidirectional encoder backbone. It accepts either text and image, or text and audio as inputs, making it trimodal overall (text, vision, audio), though it only handles two modalities at once per request. It scores 78.4 mean across seven benchmarks, beating the larger Decider 2B model. Developers comparing lightweight multimodal models can follow updates like this on daily.dev."}},{"@type":"Question","name":"What package version of transformers is required to run Liquid AI's d1 decision models?","acceptedAnswer":{"@type":"Answer","text":"Running the d1 decision models requires transformers version 5.14 or higher, installed alongside torch, torchvision, and pillow. The models ship their own custom code, so they must be loaded with trust_remote_code=True via the AutoModel.from_pretrained call in the transformers library. Anyone setting up a new model dependency can check compatibility notes like this on daily.dev."}}]}
```

