What Closed-Model Vendors Actually Let You Customize (And What Stays Locked)
When technical teams evaluate AI models, the open versus closed debate usually focuses on philosophy. But if you're trying to ship a product, the real question is simpler: what can you actually change?
Closed-source vendors like OpenAI, Anthropic, and Google don't give you model weights. You can't inspect the training data, modify the architecture, or run the model on your own hardware. But that doesn't mean you're stuck with pure vanilla API calls. These companies have built surprisingly elaborate customization layers on top of their locked foundations.
The spectrum runs from simple prompt engineering to vendor-provided fine-tuning to things that look like customization but aren't. Understanding where the real boundaries sit helps you avoid architectural decisions based on capabilities that don't exist.
Table of Contents
- The Control Surface: What You're Actually Touching
- Prompt Engineering: The Free Layer Everyone Gets
- Fine-Tuning APIs: The Official Customization Path
- What Stays Permanently Locked
- The Cost-Benefit Reality Check
- When Closed Models Make Sense Anyway
The Control Surface: What You're Actually Touching
Analogy: Closed models are like leasing a car with driver-assist features. You can adjust the seat, set cruise control preferences, and tune the stereo, but you can't rebuild the engine or modify the transmission software.
Closed-source providers operate on a control panel model. You interact with pre-built interfaces designed by the vendor. These interfaces let you modify behavior within carefully bounded limits.
Here's what the customization spectrum actually looks like:
| Customization Level | What You Control | What Stays Locked | Who Provides It |
|---|---|---|---|
| Base API calls | Nothing | Everything | All vendors |
| System prompts | Instructions, context | Model weights, training | All vendors |
| Few-shot examples | Input-output patterns | Base capabilities | All vendors |
| Fine-tuning API | Task-specific adjustments | Core architecture | OpenAI, Google, some others |
| Model weights access | Everything | Nothing | Never on closed models |
The jump from row three to row four is not as big as marketing suggests. Fine-tuning through vendor APIs is more like permanent prompt tuning than actual model modification.
Prompt Engineering: The Free Layer Everyone Gets
Every closed API lets you send a system prompt. This is your primary customization tool and it's more powerful than people expect.
System prompts let you:
- Define role and expertise level
- Set tone, style, and formatting rules
- Establish domain-specific context
- Provide examples inline
- Create guardrails for output structure
A well-engineered prompt can make GPT-4 behave like a completely different model. You can turn it into a terse code reviewer, a patient tutor, or a ruthless editor. The model's underlying knowledge doesn't change, but the interface you're interacting with transforms completely.
The practical limit is prompt length. Most vendors cap context windows at 128k tokens or less. If your customization needs more context than that, you're looking at retrieval systems, not prompt engineering.
Fine-Tuning APIs: The Official Customization Path
Fine-tuning sounds impressive. You're training the model on your data. But on closed APIs, you're not touching the actual weights the way you would with an open model.
Here's what vendor fine-tuning actually does:
It takes your examples (usually hundreds to thousands of input-output pairs) and adjusts how the model responds to similar inputs. Think of it as teaching the model your preferred patterns, not injecting new knowledge.
What fine-tuning handles well:
- Consistent formatting (always output JSON in this exact schema)
- Domain-specific terminology (use our company's product names correctly)
- Tone calibration (be more formal, less verbose, etc.)
- Task-specific improvements (better at extracting entities from our document types)
What fine-tuning cannot do:
- Add knowledge the base model doesn't have
- Change fundamental reasoning capabilities
- Fix bias or safety issues in the base model
- Give you any visibility into what changed under the hood
OpenAI's fine-tuning API charges you per training token and per inference token on your custom model. Google offers similar capabilities. The process is entirely black-box. You upload data, wait for training to complete, get back a model ID, and hope it works better.
You never see gradients, loss curves, or intermediate checkpoints. You can't inspect what changed. You definitely can't export the fine-tuned version and run it elsewhere.
What Stays Permanently Locked
No amount of money or enterprise contracts will get you access to certain things on closed platforms.
You will never get:
- Model architecture specifications
- Training data composition or sources
- Safety filter implementation details
- Actual weight values or checkpoints
- The ability to run the model offline
- Migration paths to other providers
- Guarantees about future API compatibility
The vendor can change the base model at any time. OpenAI has updated GPT-4 multiple times, sometimes changing behavior in ways that broke existing applications. Your fine-tuned model sits on top of whatever base they're currently serving.
If the vendor raises prices, changes terms, or goes out of business, you have zero portability. Your fine-tuning work is locked to that specific API forever.
The Cost-Benefit Reality Check
Closed model customization costs more than it appears.
Fine-tuning charges:
- Training cost per token (usually 8x base rate)
- Inference cost on custom models (usually 2-3x base rate)
- Storage fees for keeping the model alive
- Re-training costs when the base model updates
For many use cases, better prompts plus retrieval (RAG) deliver equivalent results at a fraction of the cost. You're paying for convenience and the vendor's brand name, not fundamentally better capabilities.
One large enterprise compared fine-tuned GPT-3.5 against well-prompted GPT-4 for a document classification task. The fine-tuned model cost 40% more per API call and performed worse. The prompt-engineered approach won on both accuracy and cost.
When Closed Models Make Sense Anyway
Despite all these limitations, closed models are often the right choice.
Choose closed when:
- You need cutting-edge reasoning (frontier models are closed)
- Compliance allows external APIs
- You value vendor-managed updates over control
- Your team lacks ML infrastructure expertise
- Time to production matters more than customization depth
The best closed models are meaningfully better than open alternatives at complex reasoning tasks. If you need that capability and can work within API constraints, the locked nature is a reasonable tradeoff.
Just be honest about what you're getting. You're renting access to a constantly shifting foundation with a limited set of adjustment knobs. That's fine for many products. It's terrible for others.
Conclusion
Closed-model vendors give you more customization than pure API access, less than actual model ownership. The spectrum runs from prompt engineering (free, powerful, limited by context) to fine-tuning APIs (expensive, black-box, narrowly useful).
The weights stay locked. The architecture stays hidden. The vendor keeps all the power.
For technical teams, this means building with an exit strategy or accepting permanent vendor lock-in. Prompt engineering plus retrieval can replicate most fine-tuning benefits at lower cost. When you do fine-tune, treat it as API configuration, not model training.
The real customization decision isn't open versus closed. It's whether you need control over the foundation, or just the interface.