<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/openai-built-a-finance-ai-with-morgan-stanley-and-the-benchmark-is-doing-a-lot-of-heavy-lifting-vogmq8kqd" -->

---
title: OpenAI built a finance AI with Morgan Stanley and the...
description: OpenAI launched ChatGPT for Financial Services, built on GPT-6 Astra inside ChatGPT Work mode, co-developed with Morgan Stanley and Evercore, and bundling...
canonical: https://daily.dev/posts/openai-built-a-finance-ai-with-morgan-stanley-and-the-benchmark-is-doing-a-lot-of-heavy-lifting-vogmq8kqd
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: OpenAI built a finance AI with Morgan Stanley and the benchmark is doing a lot of heavy lifting | daily.dev
og:description: OpenAI launched ChatGPT for Financial Services, built on GPT-6 Astra inside ChatGPT Work mode, co-developed with Morgan Stanley and Evercore, and bundling...
og:url: https://daily.dev/posts/openai-built-a-finance-ai-with-morgan-stanley-and-the-benchmark-is-doing-a-lot-of-heavy-lifting-vogmq8kqd
og:image: https://api.daily.dev/og/posts/VOgMQ8KQD.png
og:image:alt: OpenAI built a finance AI with Morgan Stanley and the benchmark is doing a lot of heavy lifting
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI built a finance AI with Morgan Stanley and the benchmark is doing a lot of heavy lifting

**[Trends](https://daily.dev/sources/trends)** · 2 min read · 1 upvotes · 0 comments

## Summary

OpenAI launched ChatGPT for Financial Services, built on GPT-6 Astra inside ChatGPT Work mode, co-developed with Morgan Stanley and Evercore, and bundling licensed data from providers like S&P Capital IQ, LSEG, Moody's, and PitchBook for use cases such as valuations, LBO models, and pitchbooks. OpenAI touts a benchmark called OfficeQA Pro, where GPT-6 Astra scores 69.9% versus 60.2% for GPT-5.6 Sol and beats Claude Fable 5.1 by 7%. Critics note that 69.9% accuracy implies roughly a 30% error rate on financial document analysis, a significant risk for institutions handling real money. The announcement also omits any mention of EU data residency or DORA compliance, a gap that could matter for European financial institutions given DORA's third-party oversight and exit-strategy requirements.

## Content

OpenAI just launched ChatGPT for Financial Services, and the headline number is doing a lot of heavy lifting.

The product runs on GPT-6 Astra inside ChatGPT Work mode, built alongside Morgan Stanley and Evercore. It bundles licensed data from S&P Capital IQ, LSEG, Moody's, PitchBook, MSCI, Preqin, and a handful of others. The pitch: research, LBO models, buyer screening, pitchbooks, all in one place. Enterprise security features (SSO, role-based access, encryption, information barriers) are included.

The benchmark OpenAI is leaning on is OfficeQA Pro, which tests whether AI agents can parse U.S. Treasury Bulletins, financial tables, charts, and footnotes. GPT-6 Astra scores 69.9%, versus 60.2% for GPT-5.6 Sol and roughly 63% for Claude Fable 5.1. That's a 9-point jump over the prior model, and OpenAI is calling it out directly.

Here's the problem: nobody outside OpenAI seems to have heard of OfficeQA Pro before today. The community reaction is skeptical. One widely-shared repost flags that on what's being called "BullshitBench," GPT-6 Astra beats previous OpenAI models but still falls short of Anthropic's offerings. The implication: OpenAI found a benchmark where it wins, then announced a product around it.

There's also the small matter of a 30% error rate. A score of 69.9% sounds impressive until you remember that means roughly one in three answers on a test involving financial tables is wrong. For institutions making investment decisions, that's not a footnote.

European financial institutions have a separate concern entirely: the announcement says nothing about EU data residency or DORA compliance. ICT third-party oversight, exit strategies, and the possibility of AI providers being designated critical service providers are all live regulatory questions that OpenAI apparently didn't feel like addressing at launch.

The product is real, the data partnerships are real, and the use case is genuinely interesting. But the benchmark strategy is going to follow this announcement around for a while.

## Questions this post answers

### What benchmark score does GPT-6 Astra achieve on OfficeQA Pro compared to GPT-5.6 Sol and Claude Fable 5.1?

GPT-6 Astra scores 69.9% on OfficeQA Pro, a benchmark testing whether AI agents can find and analyze information across U.S. Treasury Bulletins including complex tables, charts, and footnotes. That compares to 60.2% for GPT-5.6 Sol and is roughly 7% higher than Claude Fable 5.1's score.

_Teams weighing AI accuracy trade-offs for finance workflows can track benchmark shifts like this one on daily.dev._

### What data providers and companies are integrated into OpenAI's ChatGPT for Financial Services?

ChatGPT for Financial Services bundles licensed data from S&P Capital IQ, LSEG, Moody's, PitchBook, MSCI, Preqin, Datasite, Dow Jones Factiva, and Box, and runs on GPT-6 Astra inside ChatGPT Work mode. Morgan Stanley and Evercore co-developed the product, targeting use cases like valuations, LBO models, buyer screening, and pitchbooks.

_Developers building on vertical AI integrations can follow enterprise AI partnership news on daily.dev._

### Does OpenAI's ChatGPT for Financial Services address EU DORA compliance requirements?

No, the product announcement does not mention EU data residency or DORA (Digital Operational Resilience Act) compliance. DORA requires ICT third-party oversight and documented exit strategies, and OpenAI could eventually be designated a critical service provider under the framework, making this a significant omission for European financial institutions.

_Compliance-conscious engineers evaluating AI vendors for regulated markets can follow gaps like this on daily.dev._

## Community take

How the wider developer community reacted, aggregated from 4 discussions and 347 comments across x (as of 2026-09-11).

**TL;DR:** Reaction is split between job-loss anxiety for analysts/bankers and skepticism about whether the tool is a real breakthrough, with recurring concern that auditability, data provenance, and accountability matter more than the model's reasoning.

**Sentiment:** 20% positive · 35% mixed · 45% skeptical

**The case for**

- Some see it as a genuine productivity upgrade, reducing time spent scraping and formatting data so people can focus on strategy.
- A few think firms that adopt and execute well will gain a real edge over slower-moving competitors.

**The pushback**

- Many argue the real bottleneck isn't the model but auditability — every number needs traceable lineage back to a data snapshot for compliance.
- Several question how it handles messy, incomplete, or revised books rather than clean data.
- Some doubt this is a meaningful breakthrough, pointing out the advertised features are fairly basic (company info, financial summaries, slide decks).
- Concerns about accountability: no model can hold a professional's signature or bear disciplinary consequences when a valuation goes wrong.
- Worries about data privacy pushing banks and hedge funds to build their own in-house agents instead.
- Some note GPT-6 Astra reportedly burns through usage credits much faster than the previous model.

**By community**

- x (mixed): Replies range from job-loss jokes and skepticism about real-world trust/auditability to a few genuinely positive takes on productivity gains, with much of the volume being low-signal hype or reaction GIFs.

**Hottest debate:** Whether the product represents a substantive breakthrough for finance work or just packages basic lookups behind a big benchmark, with auditability and accountability seen as the real unresolved issue.

**Open questions**

- How does the model handle messy, incomplete, or inconsistent financial records that most real firms actually have?
- What happens to point-in-time consistency and data revisions for backtests and audits?
- Is there a clear boundary between an AI-traced citation and the human sign-off/authorization record required for compliance?
- How does its accuracy on live market data compare to established terminals like Bloomberg?

**Highlights**

> @OpenAI Building the model was never the hard part. Owning the number is. When a valuation goes wrong, someone's membership number is on it. Someone faces a disciplinary committee. That accountability is what a signature actually means, and no model can hold it. The question for this
> — [cagujjubhai on x · 2 comments](https://x.com/cagujjubhai/status/2098255544675082550)

> @OpenAI “Conservative” Community Banks that do not start seriously considering/investing in AI capabilities for their institution will lose fast.  @commbankerguy what do you think?
> — [FullWalnut on x · 1 points, 2 comments](https://x.com/FullWalnut/status/2098298069087637970)

> @OpenAI How does this treat messy books? A lot of finance work here starts from incomplete records, mixed accounts, and figures that only make sense to the owner.  Curious whether the model waits for clean data or can work with what most firms actually have.
> — [NidacityNG on x · 2 points](https://x.com/NidacityNG/status/2098327667460145604)

> @OpenAI The product is not the model. It is the audit trail. Banks do not lose sleep over a pretty pitchbook. They lose sleep over a number that cannot be sourced when compliance calls. Research draft: [✓] minutes Model build: [✓] fast Sign-off: [X] still human, still slow If
> — [simplified1554 on x · 1 points](https://x.com/simplified1554/status/2098298343587811505)

> @OpenAI The hard part for finance teams will not be the reasoning, it is auditability - every number in a client deck needs lineage back to a data snapshot. Curious how the built-in financial data handles revisions and point-in-time consistency for backtests.
> — [sirHe12 on x · 1 points](https://x.com/sirHe12/status/2098293589130539495)

**Source threads**

- [x](https://x.com/OpenAI/status/2098118191029624911) · 0 points · 343 comments
- [x](https://x.com/testingcatalog/status/2098191832144572673) · 0 points · 1 comments
- [x](https://x.com/testingcatalog/status/2098193015193932084) · 0 points · 3 comments
- [x](https://x.com/scaling01/status/2098378972014809539) · 0 points · 0 comments

---

Tags: [#openai](https://daily.dev/tags/openai), [#chatgpt](https://daily.dev/tags/chatgpt), [#fintech](https://daily.dev/tags/fintech)

[View this post on daily.dev](https://daily.dev/posts/openai-built-a-finance-ai-with-morgan-stanley-and-the-benchmark-is-doing-a-lot-of-heavy-lifting-vogmq8kqd)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"OpenAI built a finance AI with Morgan Stanley and the benchmark is doing a lot of heavy lifting","url":"https://daily.dev/posts/openai-built-a-finance-ai-with-morgan-stanley-and-the-benchmark-is-doing-a-lot-of-heavy-lifting-vogmq8kqd","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/openai-built-a-finance-ai-with-morgan-stanley-and-the-benchmark-is-doing-a-lot-of-heavy-lifting-vogmq8kqd"},"datePublished":"2026-09-11T07:30:09.147Z","dateModified":"2026-09-11T14:08:24.089Z","description":"OpenAI launched ChatGPT for Financial Services, built on GPT-6 Astra inside ChatGPT Work mode, co-developed with Morgan Stanley and Evercore, and bundling...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/4f9ff441eb51893da94f649f936b7af2?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/4f9ff441eb51893da94f649f936b7af2?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Trends","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Trends","logo":"https://media.daily.dev/image/upload/s--ZfSp3asX--/f_auto,q_auto/v1780996004/logos/trends?_a=BAMAMiWQ0","url":"https://daily.dev/sources/trends"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/openai-built-a-finance-ai-with-morgan-stanley-and-the-benchmark-is-doing-a-lot-of-heavy-lifting-vogmq8kqd","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"openai,chatgpt,fintech","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Trends","item":"https://daily.dev/sources/trends"},{"@type":"ListItem","position":3,"name":"OpenAI built a finance AI with Morgan Stanley and the benchmark is doing a lot of heavy lifting"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/openai-built-a-finance-ai-with-morgan-stanley-and-the-benchmark-is-doing-a-lot-of-heavy-lifting-vogmq8kqd#faq","mainEntity":[{"@type":"Question","name":"What benchmark score does GPT-6 Astra achieve on OfficeQA Pro compared to GPT-5.6 Sol and Claude Fable 5.1?","acceptedAnswer":{"@type":"Answer","text":"GPT-6 Astra scores 69.9% on OfficeQA Pro, a benchmark testing whether AI agents can find and analyze information across U.S. Treasury Bulletins including complex tables, charts, and footnotes. That compares to 60.2% for GPT-5.6 Sol and is roughly 7% higher than Claude Fable 5.1's score. Teams weighing AI accuracy trade-offs for finance workflows can track benchmark shifts like this one on daily.dev."}},{"@type":"Question","name":"What data providers and companies are integrated into OpenAI's ChatGPT for Financial Services?","acceptedAnswer":{"@type":"Answer","text":"ChatGPT for Financial Services bundles licensed data from S&P Capital IQ, LSEG, Moody's, PitchBook, MSCI, Preqin, Datasite, Dow Jones Factiva, and Box, and runs on GPT-6 Astra inside ChatGPT Work mode. Morgan Stanley and Evercore co-developed the product, targeting use cases like valuations, LBO models, buyer screening, and pitchbooks. Developers building on vertical AI integrations can follow enterprise AI partnership news on daily.dev."}},{"@type":"Question","name":"Does OpenAI's ChatGPT for Financial Services address EU DORA compliance requirements?","acceptedAnswer":{"@type":"Answer","text":"No, the product announcement does not mention EU data residency or DORA (Digital Operational Resilience Act) compliance. DORA requires ICT third-party oversight and documented exit strategies, and OpenAI could eventually be designated a critical service provider under the framework, making this a significant omission for European financial institutions. Compliance-conscious engineers evaluating AI vendors for regulated markets can follow gaps like this on daily.dev."}}]}
```

