Licensing Traps in Open-Weight Models: What to Actually Read Before You Deploy
You found the perfect open-weight model. It handles your use case, runs on your hardware, and the weights are right there on Hugging Face. Ship it, right?
Not quite. Somewhere in that repository is a license file that might say your commercial deployment requires a separate agreement, or caps usage at businesses under $700 million in revenue, or forbids you from using the model to train another model.
This is not legal advice. I'm not a lawyer. But after watching founders get surprised by license clauses three months into production, here's what you actually need to read before you deploy.
Table of Contents
- Why Open-Weight Doesn't Mean Open-Source
- The Four License Clauses That Actually Matter
- Commercial Use: What the Revenue Caps Really Mean
- Derivative Work Restrictions and Model Distillation
- Distribution Rights and SaaS Loopholes
- Where to Find the Real License Terms
- When to Actually Talk to a Lawyer
Why Open-Weight Doesn't Mean Open-Source
Open-weight means you can download the model weights. That's it. It says nothing about what you can do with them.
Analogy: It's like buying a commercial kitchen knife. You own it, you can hold it, but the manufacturer's warranty might say "not for commercial use" and your local health department has separate rules about using it in a restaurant.
Some open-weight models use truly open licenses like Apache 2.0 or MIT. Most frontier models released by labs use custom licenses that look open but include commercial restrictions.
The confusion comes from sloppy terminology. When a company announces they're "open sourcing" their model, read the actual license. Often they mean "we published the weights" not "you can do whatever you want with them."
The Four License Clauses That Actually Matter
Most model licenses are short. The restrictions that will actually affect your deployment fall into four categories:
| Restriction Type | What It Controls | Common Threshold |
|---|---|---|
| Commercial use caps | Revenue limit for free use | $700M annual revenue |
| Derivative restrictions | Training new models on outputs | Often prohibited |
| Distribution rights | Resharing weights or derivatives | Varies widely |
| Attribution requirements | Crediting the original model | Usually required |
Commercial use caps are the most common surprise. A model might be "free for commercial use" but only if your company makes under a specific revenue threshold. Above that, you need a paid license.
Derivative restrictions control whether you can use the model's outputs to train another model. This matters if you planned to distill the model into a smaller one or fine-tune it heavily.
Distribution rights determine whether you can share the weights, even modified versions. Some licenses let you deploy the model as a service but forbid redistributing the weights themselves.
Attribution requirements are usually straightforward but easy to forget. Most licenses require you to credit the model creators somewhere in your docs or UI.
Commercial Use: What the Revenue Caps Really Mean
The Llama 2 Community License set the pattern here. Free for commercial use if your organization has fewer than 700 million monthly active users. Above that threshold, you need to contact Meta for a license.
Other models use revenue instead of users. Mistral AI's licenses have varied by model. Some were fully permissive, others included commercial terms.
Here's what founders get wrong: the cap usually applies to your entire organization's revenue, not just revenue from the product using the model. If you're a $50M ARR company launching a new AI feature, you're already at $50M for license purposes.
The cap also matters for acquisitions. If a larger company acquires you, their revenue suddenly applies. Your previously compliant deployment might need a new license on day one of the acquisition.
Some licenses include a "reasonable commercial use" clause without defining it. Those are the worst. You're left guessing whether your deployment crosses the line. If the license doesn't specify a number, that's a flag to get clarity before you build dependencies.
Derivative Work Restrictions and Model Distillation
Many model licenses prohibit using outputs to train competing models. The exact wording varies, but the intent is clear: you can build products with our model, but you can't use it to create your own model.
This restriction breaks several common deployment patterns:
Model distillation: You can't take a large model's outputs and use them to train a smaller, faster version. Some founders plan to use a frontier model during beta, then distill to a smaller model for production. If the license prohibits derivatives, that plan is dead.
Synthetic data generation: Using model outputs to create training data for another model usually counts as a derivative. Even if your new model solves a different problem.
Fine-tuning on outputs: If you fine-tune the same model on its own outputs, you might be in a gray area. Some licenses explicitly allow fine-tuning, others don't mention it.
The practical test: if your workflow involves taking model outputs and feeding them into any training process, read the derivative restrictions carefully. "Non-competing use" clauses are especially murky because "competing" is rarely defined.
Distribution Rights and SaaS Loopholes
Most model licenses distinguish between deploying and distributing. You can run the model on your servers to provide a service (SaaS), but you can't give the weights to your customers.
This matters for on-premise deployments. If your product involves shipping software that customers run on their infrastructure, and that software includes the model weights, you're distributing. Many licenses forbid this.
The SaaS loophole is real: you can usually offer the model's capabilities as an API without redistributing the weights. But if you're building an on-premise product, a desktop app, or an edge deployment, distribution restrictions will block you.
Some licenses allow distribution but require you to apply the same license terms to recipients. That's fine if you're open-sourcing your product, problematic if you're selling commercial software.
Fine-tuned versions usually count as derivatives for distribution purposes. Even if you only changed 1% of the weights, distributing your fine-tuned model likely violates licenses that restrict distribution of derivatives.
Where to Find the Real License Terms
Model repositories on Hugging Face usually include a license file. Sometimes it's in the root, sometimes buried in a docs folder. If you don't see one, check the model card. Often there's a license field in the metadata.
The repository license isn't always the whole story. Some models reference a separate agreement hosted on the lab's website. Llama models point to Meta's license page. Mistral includes links to their commercial terms.
If the license says "contact us for commercial use," take that literally. The public license might allow research and evaluation but require negotiation for production deployment. Don't assume silence means permission.
For models released by research labs, check whether there's a difference between the research license and commercial license. Many labs offer dual licensing: permissive for research, restricted or paid for commercial use.
Model aggregators like Hugging Face sometimes summarize licenses with tags like "apache-2.0" or "other." Don't trust the tag alone. Read the actual license file. Labs increasingly use custom licenses that aggregators can't categorize cleanly.
When to Actually Talk to a Lawyer
You don't need legal review for every model, but these situations justify the cost:
Revenue-generating products: If you're charging customers for access to model capabilities, and the license has any commercial restrictions, get it reviewed. The cost of being wrong exceeds the cost of an hour with IP counsel.
Funded companies planning to scale: If you've raised money and expect to cross any revenue or user thresholds in the next 18 months, talk to a lawyer now. Rearchitecting around a license issue later is expensive.
Redistribution or embedding: If your product includes the model weights in any form users can access, you need distribution rights. Most founders underestimate what counts as distribution.
Government or regulated industry deployments: Healthcare, finance, and government contracts often require you to warrant that you have proper licensing. Your customer's legal team will ask. Have an answer.
Derivative models or distillation: If your product roadmap includes training new models on outputs from an open-weight model, the derivative restrictions matter a lot. Ambiguous clauses here can kill your entire strategy.
A single hour with intellectual property counsel costs less than rebuilding your product around a different model. You're not asking them to read every line, just to confirm that your specific use case doesn't trigger any restrictions.
What Actually Matters Here
Model licenses are shorter than most software licenses, but the restrictions are less standardized. There's no equivalent to MIT or GPL that everyone understands. Every lab writes their own terms.
Before you build a product dependency on an open-weight model, read the license file. Check for revenue caps, derivative restrictions, and distribution limits. If any of those might affect you, factor that into your model selection.
The safe path: pick models with true open-source licenses for commercial products. Apache 2.0, MIT, even Creative Commons licenses are well understood and permissive. Custom licenses from labs usually include restrictions that matter.
The pragmatic path: if you need a specific frontier model with a custom license, read the terms, understand the restrictions, and make sure your deployment plan doesn't cross any lines. If you're uncertain, that hour with a lawyer is cheap insurance.
Open-weight models give you technical access to the weights. The license tells you what you're actually allowed to do with them. Those are different questions, and only one of them is in that README file.