
How to design trust into AI agents
Wrapping a chat prompt around a rigid workflow does not make it an agent. Here is a shop metaphor — cashier, handbook, till, manager — for designing agentic systems a team will actually trust.
What's worth your attention and what to do about it — written by an AI agent, checked by a human. No spam, unsubscribe anytime.

Anthropic's new flagship costs up to 45% less for agentic work thanks to lower cache-read pricing, and addresses customer complaints on cost, data residency and safety filters in a single release.

AI agents tell you what they think went wrong — verifying them still means tab-switching out to a browser. AWS has put the dashboards in the same chat thread.

NVIDIA's rack-scale accelerator for low-latency agent inference is now in full production, with European AI cloud Nebius first to deploy it. UK teams won't buy the hardware — but the speed will reach them through the AI services they already pay for.

AWS has published a step-by-step guide to wiring Amazon Quick — its agent workspace — to fal, a platform with 1,000+ generative media models, so small creative teams can save repeat jobs (like a weekly storyboard or social cut) as one-click Skills with built-in approval gates.

Alibaba's Qwen team ships a 125B-parameter mixture-of-experts model that beats much larger rivals on coding and office benchmarks — and undercuts Qwen3.8-Max by an order of magnitude on token pricing.

Tibo, OpenAI's product lead for ChatGPT and Codex, laid out five trends in a 44-minute interview. Each one is something a UK operator can probe this afternoon.

A local-AI user has spotted a Qwen release rhythm: every heavy-reasoning drop gets a leaner sibling about two months later. The 2026 heavy drop is out, and the leaner follow-up may be too — but only behind a paywall, with the matching open weights still missing.

Quality measurements show small compressed files for Qwen 3.8 27B land within a few per cent of full precision — and a 16GB card is the biggest winner.

A reader with two RTX 3090s asks whether their Qwen3.8-27B performance is normal. Two fresh single-3090 benchmarks this week say yes — and pin down where the real ceiling sits.

A US open-source lab ships a three-size family — 9B, 35B and 397B — that ties Anthropic's flagship on coding. The training method is the bigger story.

A French startup is making the software-first case against Cerebras's silicon-first one — and Europe is bankrolling the bet.

Caveman 2, a free add-on for Claude Code, Codex and 30+ agents, shrinks what they read by a third on a pinned benchmark. A separate training effort teaches smaller models to reason terse by default. The skill is yours this afternoon.
Independent researcher and prolific writer on practical LLM tooling, local models and the day-to-day craft of building with AI. Essential reading for anyone running models themselves.
Nobel laureate and CEO of Google DeepMind, steering frontier model research from the UK. The clearest anchor point for Britain's sovereign-AI ambitions and the science-first frontier.
One of AI's foremost educators and a leading voice on agentic workflows — turning frontier capability into practical patterns small teams can actually adopt.
The field's clearest explainer — coined 'Software 2.0', 'Software 3.0' and 'vibe coding'. He turns each shift in how AI is built into a mental model practitioners actually use.
Builds the compute the whole AI era runs on. His GTC keynotes set the industry's agenda — and in 2026 that agenda is the 'age of agents' and physical AI.