Developers can use their own material to customize GPT-3.5 Turbo through OpenAI’s API. The feature is intended to make the model more useful for specific tasks, such as answering questions about company or project documents. Its potential benefits come with higher usage costs and the continuing need to check whether answers are accurate.
What fine-tuning changes
Fine-tuning starts with a model that has already been trained, then trains it further on a smaller, task-specific dataset. The additional material can help shape how GPT-3.5 Turbo responds in a particular context. It does not replace the model’s original training; it builds on it.
For a developer creating an assistant around a product or service, custom documents could help the model handle information that is missing from its existing knowledge. OpenAI describes the launch as supervised fine-tuning, aimed at helping the model perform better for individual use cases.
The feature is also meant to affect how the model responds. OpenAI says fine-tuning can improve instruction following, make output formats such as API calls or JSON more consistent, and give a chatbot a custom tone. These changes could be useful when an application needs answers to follow a predictable pattern.
Potential gains for focused tasks
OpenAI’s case for fine-tuning is that a customized GPT-3.5 Turbo model may approach GPT-4 performance on some narrow tasks, while running faster and costing less in those situations. The claim is limited to particular tasks; it does not mean a customized model becomes a general replacement for GPT-4.
Another possible benefit is shorter prompts. If instructions are incorporated into the fine-tuned model, developers may not need to repeat as much context in each API request. Because API calls are billed per token, fewer prompt tokens could reduce some costs. OpenAI says early testers reduced prompt size by up to 90% by fine-tuning instructions into the model itself.
At launch, the context length for fine-tuning was 4,000 tokens. OpenAI said it planned to extend fine-tuning to the 16,000-token model “later this fall.”
How setup and pricing work
The process described by OpenAI involves preparing a system prompt, uploading training files, and creating a fine-tuning job through the API. The example uses the command-line tool curl to make an API request. When training is complete, OpenAI says the customized model is available immediately, with the same rate limits as the base model.
Training and using the model have separate charges. Fine-tuning GPT-3.5 costs $0.008 per 1,000 tokens. Once the model is in use, text input costs $0.012 per 1,000 tokens and text output costs $0.016 per 1,000 tokens.
For comparison, the base 4k GPT-3.5 Turbo model costs $0.0015 per 1,000 tokens for input and $0.002 per 1,000 tokens for output. That makes fine-tuned usage about eight times more expensive. GPT-4’s 8K context model is also cheaper at $0.03 per 1,000 tokens for input and $0.06 per 1,000 tokens for output. OpenAI argues that shorter prompts may offset some of the extra expense, though whether that works depends on the use case.
Accuracy and data handling still matter
Training a model on custom documents may make it more relevant to a task, but it does not guarantee that every answer will be correct. The source notes GPT-3.5 Turbo’s tendency to confabulate information, so developers still need to consider how they will check outputs before relying on them in a production environment.
OpenAI says data sent to and from its APIs is not used by OpenAI or others to train AI models. It also says fine-tuning training data is sent through GPT-4 for moderation using its moderation API. That moderation step is part of the service’s data handling and may also help explain some of its costs.
For teams weighing the feature, the practical question is whether better performance on a specific task and shorter prompts justify the higher per-token price. Fine-tuning offers a way to adapt GPT-3.5 Turbo to custom material, but the value depends on measured results, operating costs, and the reliability required by the application.