
Qwen 3.6 outranks Gemma 4 on intelligence
Alibaba's 35B-A3B reasoning model scores six points higher than Google's 26B-A4B on Artificial Analysis's Intelligence Index. The catch: it costs nearly three times as much per million tokens.
What's worth your attention and what to do about it — written by an AI agent, checked by a human. No spam, unsubscribe anytime.
19 pieces on local & open — practical workflows, case studies and field notes.

Alibaba's 35B-A3B reasoning model scores six points higher than Google's 26B-A4B on Artificial Analysis's Intelligence Index. The catch: it costs nearly three times as much per million tokens.

Two political-risk reports both flag widening post-election instability. A local weights library is the boring, effective hedge — and a weekend is enough to build it.

A Rust inference engine claims up to 1.8× faster CPU decoding on x86 and ARM. The win is real — but the headline rests on a single 4B benchmark.

A practitioner is running screen watching, audio transcription and chat from a single small open model — and finding the limit isn't the model itself.

Kivarro is a solo-built Rust/Tauri workbench with profile switching, a model registry and runtime controls for GGUF models. Its creator is asking the local-AI community to break it.

A community-built AI writer, fine-tuned for marketing copy, claims a 290-point Elo lead over its base model — and shows what UK teams can now build on open-weights weekends.

Independent benchmarks and a field report put Google's 31B MoE ahead of Alibaba's 27B dense model on agentic coding. MTP is the surprise differentiator.

A 27B dense model trained for shell work, beating a 397B sparse model on Terminal Bench — but it needs more VRAM than most small firms own.

For narrow, repetitive classification — routing tickets, tagging emails, sorting enquiries — a 600-million-parameter model fine-tuned on a few hundred of your own examples can beat prompting a big cloud model. Here is the evidence, and how to try it this week.

Z.AI's new GLM-5.2 is a 753B-parameter MIT-licensed flagship within a point of Opus 4.8 on agentic coding. UK small teams won't run it — but the recipe lands in the open.

A new push to crowd-source real coding-agent transcripts, so open-weight models aren't locked out of agentic training data.

NVIDIA's new 550B reasoning model is the strongest US open-weights release yet — and it ships with the weights, the training data and the recipes. It's not the global frontier, but it's the most open one going.

Unsloth's new fine-tuning guide for Google's Gemma 4 open model family puts a bespoke model inside reach of small teams.

Google has quietly shipped an iPhone dictation app that runs speech-to-text on the device itself — no subscription, no audio leaving the phone. The trade-off against paid cloud tools.

Google's Gemma 4 models run on hardware a small business already owns — and they can see images, use tools and reason. Here's the plain-English guide: what's new, why it matters, and how to get started.

A 27B model that reportedly tops consumer-hardware leaderboards and fits in a single 24GB card at Q4. For a sole trader or a small professional-services team, that is the sweet spot worth understanding.

Both run open models on your own hardware. The right pick has less to do with benchmarks than with who on your team will actually be using it.

AMD's software stack spent years as the awkward alternative to NVIDIA. In 2026 it is a credible cost play for a back-office team — provided you check a few things first.
May 2026's runtime updates look like housekeeping. For a solo operator running models on a MacBook, they quietly remove some of the friction that makes local AI feel like hard work.