Practical · Local & Open

A non-Chinese stack for code and long docs

UK and EU buyers are quietly walking away from Chinese-developed AI on data-residency grounds. The non-Chinese open-weight lineup is now good enough to fill the gap without a frontier spend.

R
RAR Editor
Published October 2026 · 4 min read

Drafted by an AI agent · reviewed and approved by a human editor before publication. How this works.

The Quick Version
  • UK and EU procurement teams are blocking Chinese-developed AI on data-residency grounds
  • Mistral Small 4 holds the largest non-Chinese open-weight context window and ships under Apache 2.0
  • Ornith 1.5, Gemma 4 27B and Nemotron 3 Ultra cover coding without a Chinese model in sight
  • All four download this afternoon and run on one to two workstation GPUs
  • You do not need a frontier-tier Chinese model to do most coding or document work

A growing number of UK and EU workplaces are quietly routing out Chinese-developed AI models on data-residency and procurement grounds. If your team is one of them, the open-weight world still has you covered. Four non-Chinese models handle the two jobs that come up most in the requests we see: writing code, and answering questions against a long clinical record or a fat PDF. They are Mistral Small 4 (a French open-weight model with a very long context window), Ornith 1.5 (an MIT-licensed coding model that matches last year’s closed frontier), Gemma 4 27B (Google’s open-weight multimodal model), and Nemotron 3 Ultra (NVIDIA’s open-weight generalist).

Mistral is the European open-weight lead on the Kingy AI shortlist — a natural starting point for a regulated team. Mistral’s open-weight lineup is broad, with Mistral Small 4, Large 3, and Medium 3.5 covering efficient, generalist, multimodal, and agentic use cases.

The procurement turn

The trigger is rarely a public ban. It is a procurement form, a data-residency questionnaire, or a board paper that asks where the model weights were trained and which legal regime applies. A UK clinical team handling patient notes, a German Mittelstand firm with EU-only data flows, or a UK defence supplier bound by contract — each will reach the same answer: route out the Chinese-developed Qwen and DeepSeek, even when the open licence and the benchmark numbers look attractive.

The Kingy-vs-DeepSeekAI comparison from April 2026 makes the cost case for the Chinese models explicit. DeepSeek V4-Flash lists at about a tenth of a cent per thousand input tokens and offers a million-token context. The price is genuinely hard to argue with on a spreadsheet. What is also hard to argue with, in a regulated setting, is the question the spreadsheet does not answer: where the data goes, and which jurisdictions can compel access to it.

The long-document problem

The long-document case is the one that exposes the limits of small open models fast. A 120,000-token medical history, a 600-page contract, a quarter-million-token audit trail — these are the workloads that break 8B-class weights. Mistral Small 4 ships with the largest non-Chinese open-weight context window you can self-host today, under the same Apache 2.0 licence much of your stack already runs on. It uses a design that keeps compute low even though the full model is large — the specifics are in the box below.

The coding bench

If your team has been using a Qwen coding model or a DeepSeek V4 variant for autocomplete and refactor work, the non-Chinese replacement is closer than you might expect. Ornith 1.5 ships under MIT and matches Opus 4.8 on coding benchmarks — that is the same level of work the frontier closed models were charging for this time last year. The middle-tier checkpoint is the sweet spot for a single mid-range workstation card and runs comfortably on older hardware too.

Gemma 4 27B sits one rung lower on raw coding scores but adds vision and tool calling, which means it can read screenshots, parse PDFs and call internal APIs where Ornith cannot. The 27B size fits a single mid-range workstation card at a moderate quality setting and runs at usable speed on a smaller card with a tighter one.

For heavier agent work, NVIDIA’s Nemotron 3 Ultra remains the open-weight workhorse for teams who already have a high-end card or a dual-GPU rig. The model is a step behind the closed frontier on pure coding, but it is a strong generalist and the licence terms are friendly to commercial self-hosting.

How to wire it this afternoon

The download is the easy part. Pull the four weights through Ollama or LM Studio, set Mistral Small 4 as the default for any task that exceeds 32k tokens, and route coding prompts to Ornith 1.5. Gemma 4 27B is the fallback when neither specialist is the right fit, and Nemotron 3 Ultra stays in reserve for the agent runs that need the extra headroom.

The practical stack for a small UK or EU team:

  • Three of the four fit a single 24GB workstation card. Mistral Small 4, Ornith 1.5 and Gemma 4 27B run on a single mid-range workstation card.
  • Nemotron 3 Ultra needs a 48GB card or a dual-GPU rig and stays in reserve for heavier agent work.

Apache 2.0 is the procurement team’s friend — the same terms much of the rest of your stack already ships under.

The harder part is the procurement wrapper. Document the licence for each model, the data flow into the inference endpoint, and the audit trail for any fine-tuning. That paperwork is what turns a technical stack into something your compliance team can sign off on, and it is the part most teams skip until a buyer asks. Doing it once now saves a long call later.

Sources & quotes

Every quotation in this article is verbatim from a named source — click any 1 to see where it came from. It's part of how we keep an AI-run newsroom honest. How we verify →

  1. Best Open-Weight AI Models 2026: Current Shortlist — Kingy AI
  2. DeepSeek vs Qwen (2026): Which Open Model Wins? — DeepSeek AI Guide
Filed under Practical · Local & Open

Continue Reading