<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/arcee-ai-s-trinity-large-thinking-a-398b-reasoning-model-for-agentic-workloads-kfvmclmn9" -->

---
title: Arcee AI&#x27;s Trinity Large-Thinking: a 398B reasoning...
description: Arcee AI has released Trinity Large-Thinking, a 398B sparse Mixture-of-Experts reasoning model under Apache 2.0. Despite its size, it activates only ~13B...
canonical: https://daily.dev/posts/arcee-ai-s-trinity-large-thinking-a-398b-reasoning-model-for-agentic-workloads-kfvmclmn9
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Arcee AI&#x27;s Trinity Large-Thinking: a 398B reasoning model for agentic workloads | daily.dev
og:description: Arcee AI has released Trinity Large-Thinking, a 398B sparse Mixture-of-Experts reasoning model under Apache 2.0. Despite its size, it activates only ~13B...
og:url: https://daily.dev/posts/arcee-ai-s-trinity-large-thinking-a-398b-reasoning-model-for-agentic-workloads-kfvmclmn9
og:image: https://api.daily.dev/og/posts/KFvmclmN9.png
og:image:alt: Arcee AI&#x27;s Trinity Large-Thinking: a 398B reasoning model for agentic workloads
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Arcee AI's Trinity Large-Thinking: a 398B reasoning model for agentic workloads

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 1 upvotes · 0 comments

## Summary

Arcee AI has released Trinity Large-Thinking, a 398B sparse Mixture-of-Experts reasoning model under Apache 2.0. Despite its size, it activates only ~13B parameters per token, making inference 2-3x faster than comparable dense models. It targets agentic workloads with 512k context windows and multi-turn tool calling. On PinchBench (agentic capability benchmark), it ranks #2 just behind Claude Opus 4.6, but costs roughly 96% less at ~$0.90 per million output tokens. The model is available via DigitalOcean Serverless Inference, Kilo Code (free trial), and Hugging Face for self-hosting or fine-tuning.

## Content

Arcee AI has released Trinity Large-Thinking, a 398B-parameter sparse Mixture-of-Experts reasoning model under Apache 2.0. Despite the massive parameter count, it only activates around 13B parameters per token during inference, which makes it 2-3x faster than comparably-sized dense models and keeps costs surprisingly low.

## What it actually does

Trinity Large-Thinking is built for agentic work: multi-turn tool calling, long-horizon tasks, and extended reasoning traces. It supports context windows up to 512k tokens, which matters for the kind of long-running agent loops where most models start to fall apart.

On PinchBench, a benchmark focused on agentic capability, it ranks #2 - just behind Claude Opus 4.6. Whether that translates to your specific use case is a fair question (benchmarks rarely do), but it's at least measuring something relevant rather than just math problems.

## The cost gap is hard to ignore

This is where things get interesting. Trinity Large-Thinking runs at roughly $0.90 per million output tokens. Claude Opus 4.6, which sits one spot above it on PinchBench, costs dramatically more - around 96% more by some estimates. That's not a rounding error. For teams running agentic workloads at any real volume, that difference matters.

## Where to run it

A few options depending on your setup:

- **DigitalOcean Agentic Inference Cloud** - available now in Public Preview via Serverless Inference. No infrastructure to provision; query it through the API or console immediately.
- **Kilo Code / KiloClaw** - free access for one week starting April 6th, if you want to try it without committing to anything.
- **Hugging Face** - model weights are available for self-hosting or fine-tuning under Apache 2.0.

## The open weights matter

Apache 2.0 means you can actually do something with this beyond just calling an API. Fine-tune it, self-host it, build on top of it. For a model performing at this level, that's not a given - most frontier-adjacent models are locked behind proprietary licenses.

The honest caveat: benchmarks like PinchBench tell you something, but not everything. If you're evaluating this for real use, the free Kilo access is a reasonable way to test it against your actual tasks before committing.

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents), [#mixture-of-experts](https://daily.dev/tags/mixture-of-experts)

[View this post on daily.dev](https://daily.dev/posts/arcee-ai-s-trinity-large-thinking-a-398b-reasoning-model-for-agentic-workloads-kfvmclmn9)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Arcee AI's Trinity Large-Thinking: a 398B reasoning model for agentic workloads","url":"https://daily.dev/posts/arcee-ai-s-trinity-large-thinking-a-398b-reasoning-model-for-agentic-workloads-kfvmclmn9","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/arcee-ai-s-trinity-large-thinking-a-398b-reasoning-model-for-agentic-workloads-kfvmclmn9"},"datePublished":"2026-04-06T18:18:27.279Z","dateModified":"2026-04-09T15:14:08.918Z","description":"Arcee AI has released Trinity Large-Thinking, a 398B sparse Mixture-of-Experts reasoning model under Apache 2.0. Despite its size, it activates only ~13B...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/1d7ad7cc96ef4f34e09e4605a2444f5c?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/1d7ad7cc96ef4f34e09e4605a2444f5c?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/arcee-ai-s-trinity-large-thinking-a-398b-reasoning-model-for-agentic-workloads-kfvmclmn9","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-agents,mixture-of-experts","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Arcee AI's Trinity Large-Thinking: a 398B reasoning model for agentic workloads"}]}
```

