OpenAI documentation screenshots described Foundry as a developer platform for running the company’s newer models on dedicated computing capacity. The reported offer combined reserved resources with controls over model versions, performance and support, aimed at customers running larger workloads.
A platform built around reserved capacity
The screenshots called Foundry an offering “designed for cutting-edge customers running larger workloads.” They described inference at scale, with customers able to control model configuration and performance profile.
Rather than drawing on a shared pool of capacity, Foundry would provide a static allocation of compute for a single customer. The article suggested this might run on Azure, OpenAI’s preferred public cloud platform, but that detail was not confirmed.
Customers would also be able to monitor specific instances using tools and dashboards OpenAI uses to build and optimize models. This kind of visibility could help organizations track the systems serving their applications and manage how those systems perform.
More control over models and support
The described platform included some version control. That would let customers decide whether to move to newer model releases, rather than having updates imposed automatically. For businesses that depend on consistent outputs, control over upgrade timing can be an important part of operating a model in production.
The documentation also referred to “more robust” fine-tuning for OpenAI’s latest models. Fine-tuning lets customers adapt a model to a particular use, while the reported service-level commitments covered instance uptime and engineering support on a calendar schedule.
Together, these features point to an offering for organizations that need more than access to a model. They would be paying for dedicated capacity, operational controls and a defined support arrangement. The screenshots did not provide enough information to establish precisely how those commitments would work.
Commitments came with a high price
Foundry rentals were described in dedicated compute units, with commitments of three months or one year. Each model instance would require a particular number of units, so the cost would depend on the configuration selected.
The article gave one example: a lightweight version of GPT-3.5 would cost $78,000 for three months or $264,000 for a year. It compared that with a recent-generation Nvidia DGX Station, priced at $149,000 per unit. The comparison gives readers a sense of the scale of the proposed commitment, though the two offerings are not described as equivalent products.
Such pricing reflects the expense of running powerful AI systems. The source noted that training advanced models can cost millions of dollars and that inference—the computation involved in producing model responses—also carries significant costs. A dedicated service therefore shifts some of that expense into a direct customer commitment.
A model listing hinted at what might come next
A pricing chart in the screenshots reportedly listed a text-generating model with a 32k maximum context window. A context window is the amount of text a model can consider before generating more. A longer window lets it take more material into account at once.
Because GPT-3.5 was described as having a 4k maximum context window, some observers speculated that the model with a 32k window could be GPT-4, or a step toward it. The article presented this as speculation based on the chart, not as a confirmed model announcement.
At the time of publication, TechCrunch said it had contacted OpenAI to confirm the screenshots. The reported details should therefore be understood as information shown in documentation images, rather than a confirmed public product launch.
Part of a broader effort to earn revenue
The potential Foundry launch came as OpenAI faced pressure to turn a profit after a multibillion-dollar investment from Microsoft. The source reported that the company expected to make $200 million in 2023, compared with more than $1 billion put toward the startup so far.
OpenAI had already introduced ChatGPT Plus, a pro version of ChatGPT starting at $20 per month, and partnered with Microsoft on Bing Chat. It also made its technology available through Azure OpenAI Service and maintained Copilot, a code-generation service developed with GitHub.
Foundry would fit into this wider set of ways to serve customers, with a model aimed at organizations seeking dedicated infrastructure and operational commitments. Its reported pricing and longer-term rentals suggest a product for substantial workloads, while the screenshots left key launch details unconfirmed.