By Lily Caruso · July 24, 2026 There is a ritual that repeats every time an organization decides — by conviction or, increasingly, by contract — to bring its AI in-house. Someone opens a leaderboard. Someone else asks which open model is best.

And the room spends an hour on a question that has no answer, because it is the wrong question — roughly like asking which vehicle is best while declining to mention whether you're hauling gravel or children. The right question is best for what. And by that light, the open-model world sorts into six families with personalities so distinct you could cast a sitcom with them.

Consider this a field guide, written the weekend before the field changes again — which, in this genre, is every weekend. The giant. Monday, Moonshot's Kimi K3 releases its weights: 2.8 trillion parameters, a million-token context window, the largest open model ever built, benchmark wins on the coding boards, and a frontier-or-not debate the coverage hasn't settled (we log the dispute; we decline to referee it).

Here is the thing to understand about K3, and it is the first hidden truth of this market: almost nobody will run it, and that was never the point. A model this size is less a download than a lifestyle — the hardware bill alone reads like a small defense budget. Its actual job is dynastic: move the ceiling, then let a thousand distilled descendants fall from it.

The flagship is not the product. The flagship is the parent. (The ambitious teenager downloads the full weights anyway, on principle, onto a laptop that will never forgive him. There is one in every era, and civilization quietly depends on them.) The insurgent.

DeepSeek deserves its own business-school case, if only for the origin story: a Chinese quantitative hedge fund's side project that, one January morning, cost Nvidia $589 billion in a single session — history's most expensive proof-of-concept, briefly mistaken for an obituary. The buildout it supposedly doomed then tripled. But DeepSeek's lasting contribution was quieter and more subversive: it made efficiency the axis of competition — frontier-adjacent reasoning at a fraction of the compute — and every lab on earth promptly studied its homework.

You choose this family when the binding constraint is the electric bill. In the self-hosted world, the binding constraint is always the electric bill. The corporate standard.

Meta's Llama line is what the phrase "enterprise-grade" was invented to describe — rarely the chart-topper, always the safe conversation. Deepest tooling, widest talent pool, provenance that makes conservative counsel exhale. In the adoption surveys it is simply there, the way Excel is there.

The old industry proverb held that nobody ever got fired for buying IBM; its successor is alive and well, and it runs on Llama. Organizations don't pick the standard to win benchmarks. They pick it so the choice never requires a meeting.

The workhorse ecosystem. Alibaba's Qwen family is — by the numbers leaderboard tourists never check — arguably the most used open lineage on earth, and here is the second hidden truth: not through its flagships, but through its descendants. The open-model commons is carpeted in Qwen fine-tunes, distillations, and a size variant for every budget between "research cluster" and "old ThinkPad." Which suggests the metric nobody publishes but everyone should watch: derivative fertility.

Count a family's grandchildren, not its trophies. If Llama is the standard, Qwen is the substrate — the flour of the ecosystem, in everything, credited rarely. The European hedge.

Mistral occupies a position no benchmark measures: the serious open lab that is neither American nor Chinese. In a world where model provenance is becoming a procurement question and export politics a roadmap risk, a passport is now a feature. Efficient engineering, genuine culture — and a certain kind of buyer chooses it for the same reason certain governments buy Airbus.

The weights are good. The neutrality closes the deal. The small-and-local class.

Google's Gemma, Microsoft's Phi, and their kin — models sized for one machine, a laptop, a kiosk. Leaderboard tourists scroll past them; deployment engineers marry them. Because the honest secret of enterprise AI — hidden truth the third — is that most workloads are narrow: classify this, summarize that, extract the other.

Yesterday's pickup-truck server mostly runs these. The future's most common AI will be its least glamorous, which is how infrastructure has always worked: nobody writes poems about the water main. Now, the silences — what the leaderboard conversation steps around.

First, the geography, stated at full awkwardness: the free world's best free models are, at this writing, substantially Chinese — K3, DeepSeek, Qwen — while the strongest closed models remain American. The own-your-AI movement and the export-control regime are two trains sharing one track with great mutual politeness, and almost nobody planning a five-year deployment is pricing the intersection. Second, the flagships are marketing; the ecosystems are the market — the model you'll actually run in 2027 is a shrunken descendant of something above, which makes fertility, not scores, the predictive statistic.

Third, the 2 a.m. gap: charts rank what models do in examinations; enterprises live with what models do under load, on last year's hardware, against this morning's contract amendments. Between those two rankings lies most of the industry's disappointment, and nearly all of its arbitrage. The honest hedges, because this field punishes confident prose.

Version numbers in this genre age like produce — treat any ranking, ours included, as dated the week it ships. Benchmark supremacy claims, K3's very much among them, remain contested. And "open" itself spans a licensing spectrum from genuinely-free to free-the-way-a-timeshare-is-free; the distance between those deserves a lawyer's full afternoon before it meets your data, not after.

The question to carry into Monday. The weights drop, the counters spin, and the coverage will ask what K3 scores. The better question — the one this taxonomy exists to sharpen — is which family the deploying world actually reaches for when the ceiling moves: the giant, the efficient, the standard, the small?

Watch where the downloads settle, not where they spike — spikes measure curiosity; settlement measures architecture. Because the future of on-site AI will not be decided by the best model on the chart. It will be decided the way every infrastructure era gets decided — quietly, one purchase order at a time, by the model that best fits the building it has to live in.

The chart is a beauty contest. The building always wins. Markets.

Tech. The Edge.