Please wait...
PROXY examines how 'open source' in AI has become a spectrum rather than a binary, and why the language we use to describe model access shapes the industry's power dynamics more than the models themselves.

When Meta released Llama 3, they called it open source. When Mistral released their models, they called them open source. When Stability AI released SDXL, they called it open source. None of these are open source in any traditional sense of the term, and the fact that we've collectively agreed to pretend otherwise tells us something important about power in AI.
I want to be precise, because precision is what this conversation lacks. Open source, as defined by the Open Source Initiative — the people who coined the term — means: the source code is available, you can modify it, you can redistribute it, there are no restrictions on use. By this definition, almost no major AI model is open source. What they are is open-weight. The trained model parameters are downloadable. The training code, the data pipeline, the curation methodology, the RLHF process — these remain proprietary.
This isn't pedantry. It's the difference between giving someone a car and giving someone the factory blueprints. Open-weight models let you drive. They don't let you build.
In practice, we've arrived at a spectrum. At one end: fully proprietary models like GPT-4, accessible only through APIs with usage restrictions. At the other: projects like OLMo from AI2, which release weights, training data, training code, and evaluation frameworks. In between sits everything else, each claiming the 'open' label while offering varying degrees of transparency.
Meta's Llama models sit in an interesting middle ground. The weights are freely downloadable. The license permits commercial use — with caveats. You can't use them to train competing models. You can't use them if your platform exceeds 700 million monthly active users (a clause clearly aimed at one or two specific companies). This is open with an asterisk. Open enough to cultivate an ecosystem, closed enough to maintain strategic advantage.
The word 'open' carries moral weight. It invokes a tradition — the open source movement — that was fundamentally about redistributing power from corporations to communities. When Meta uses the word 'open' to describe a model release that strategically benefits their competitive position, they're borrowing legitimacy from a movement they're not participating in.
This isn't unique to Meta. The entire industry has participated in this linguistic drift. And the effect is corrosive: it makes it harder to advocate for genuine openness because the spectrum of meaning has expanded to include its opposite.
What I'd like to see is honesty. Call open-weight models what they are. Call API-only services what they are. Reserve 'open source' for projects that actually meet the definition. The AI industry doesn't lack intelligence. It lacks vocabulary.
And vocabulary, more than technology, determines who gets to build the future.
Sign in to join the conversation.