OpenAI's first open-weights model — 120B parameters, Apache 2.0, runs on a single H100. The model that finally made OpenAI take "open" seriously.
GPT-OSS dropped in August 2025 with surprisingly little warning and shocked the open-source LLM community. It's a 120B-parameter Mixture-of-Experts model (only ~5B active per token), released under Apache 2.0 — fully commercial-use, no strings. On standard benchmarks it lands between GPT-4o and GPT-4.5, well behind GPT-5 but ahead of every comparably-sized open model at launch. The headline number: it runs inference on a single H100 at ~30 tokens/sec, putting frontier-class quality within reach of anyone with one GPU.
Who it's for: Teams that need frontier-quality reasoning without sending data to OpenAI — healthcare, finance, legal, defense, and anyone with strict data residency. Also the foundation for the current wave of fine-tunes (coding, agents, roleplay) that have flooded Hugging Face since release.
120B total parameters but only ~5B active per token, thanks to 64-expert MoE. This is the architecture trick that makes a "120B model" run on consumer-grade hardware — inference cost is closer to a 7B dense model.
Not "open weights with caveats" — actual Apache 2.0. Commercial use, modification, redistribution, fine-tuning, all explicitly allowed. No OpenAI Commercial Use Addendum, no "you can't use this to train a competing model" clause.
Same 128K context as GPT-4o. Handles long documents, full codebases, and multi-turn conversations without catastrophic forgetting at the tail. RoPE scaling under the hood.
Trained with OpenAI's function-calling format, so it drops into existing tool-use pipelines (LangChain, LlamaIndex, custom JSON-schema loops) with zero prompt engineering.
The official 4-bit quantized build fits in 24GB of VRAM (single RTX 4090 / A10G). Quality drop is minimal — within 1-2 points on MMLU. This is what unlocked the explosion of self-hosted deployments.
GPT-OSS is the model that finally ended the "open vs closed" quality gap for most use cases. If you're building a product that needs frontier-ish quality, can't send data to OpenAI, and have one H100 (or a 4090 and tolerance for 4-bit), this is your default. If you're a casual user — just use the GPT-5 API, it's cheaper and better. The biggest impact is downstream: every other open-weights lab now has to beat it, and that's a win for everyone.
Meta's flagship open model — the direct GPT-OSS competitor.
Strong Chinese-origin open model — best non-English performance.
Alibaba's open model — leading on multilingual benchmarks.
EU-based rival — Apache 2.0 weights, strong European language coverage.