AI agentsMicrosoftCopilotCopilotCowork

A Million Agents Won't Save You

The book that stuck with me most this month, and the one I'd recommend to anyone reading this, is People First, Come ripensare il lavoro digitale partendo dai bisogni delle persone, by my colleague Ig

The book that stuck with me most this month, and the one I'd recommend to anyone reading this, is People First, Come ripensare il lavoro digitale partendo dai bisogni delle persone, by my colleague Igor Macori, with a preface by Elisabetta Sasselli. (It's in Italian for now; if you read English, put it on the list anyway.) Its argument is one line long and almost nobody follows it: most digital-transformation projects fail not because the technology is weak, but because they start from the technology instead of from the people who have to live with it.

So this isn't a news recap. It's what I'd want to say to any CIO or IT lead after the strangest three months the modern workplace has had in years.


The agent showed up for work, wearing someone else’s engine

For three years "AI at work" mostly meant a chat box that answered when spoken to. That ended quietly in mid-June. Microsoft moved Copilot Cowork to general availability and unveiled Microsoft Scout, its first "Autopilot" agent: an always-on application that works on your files, drives a browser, runs commands, and acts in the background under its own Entra identity. The verb changed from assist to do. Tasks are no longer trapped in a single prompt in a single app; they run for minutes or hours, coordinate across Outlook, Teams, Excel and the rest, and hand back finished work with progress you can watch and steer.

The numbers Microsoft put behind it — with the denominators most coverage skipped.

Paid Copilot seats up more than 160% year-over-year (from a base of ~8M seats in Q2 2025 per Microsoft Work Trend Index), daily active use up roughly tenfold over the same 12-month window, ~480 of the Fortune 500 now using Microsoft AI in some form. Adoption is real but "10x DAU on a small base" and "10x DAU on a mature base" are different stories, and Microsoft has not disclosed which one this is. Read the growth as directional, not as proof of mature usage.

Microsoft's agentic flagship runs on Anthropic's Claude. Cowork executes in the cloud inside your tenant, covered by Enterprise Data Protection and grounded in Work IQ, but the reasoning underneath it is Anthropic's, not Microsoft's. Microsoft built its most strategic new feature on a competitor's model and called multi-model a feature, not a compromise. It is a feature. It is also a dependency, and June was the month that dependency stopped being abstract.

What a Cowork agent really costs

At GA, Cowork is billed by consumption in Copilot Credits, off until an admin enables it. One credit is about €0.01 (US$0.01 pay-as-you-go, roughly €0.008 prepaid), so this is metered cloud spend, not a flat add-on. Microsoft groups tasks into three tiers:

Task (example) Tier Credits Indicative cost
Sort a day of inbox, short replies Light ~100-300 €1-3
Turn a few reports into one memo or deck Medium ~400-700 €4-7
Aggregate many sources, build a model plus deck plus memo Heavy 700+ €7 and up

Based on Microsoft public pricing (1 credit ≈ €0.05) and internal benchmarks from the first GA tenants. The number that matters is the second-order one: a single always-on agent across 20 employees can cost €15–25k/year on top of the Copilot seat. Price the meter before you open the tap.


The frontier acquired gatekeepers

On 12 June, the U.S. Commerce Department forced Anthropic to suspend its two most capable models, Fable 5 and Mythos 5, for any foreign national, which in practice meant disabling them worldwide within hours. Two weeks later, on 26 June, OpenAI shipped GPT-5.6 (Sol, Terra, Luna) to roughly twenty pre-approved partners while a federal review process gets built. Same stated trigger both times: cybersecurity capability judged strong enough to be a national-security matter.

In a fortnight, access to the two leading American frontier families stopped being a commercial decision and became a government one. A frontier model is no longer just something you procure. It is a strategic asset in the AI arms race between the U.S. and China, gated the way you'd gate a weapons system, not a SaaS subscription.

And there is a sharper edge if you read this from Europe. These controls scope access to American citizens and Washington-approved partners, which leaves European companies structurally disadvantaged: the capability a U.S. competitor gets first is exactly the one that can be switched off for you. But the symmetric question is rarely asked: what would be the European equivalent of a counter-control? Today, none — and that absence is itself a strategic choice that the EU AI Act doesn't address.

The Mythos backstory is worth getting right. In April, Anthropic confirmed it was investigating unauthorized access to Claude Mythos Preview through a third-party vendor; the hard export control came later, reportedly tied to fears that a China-linked group could distill the model. Be precise, though: Anthropic never said Mythos was distilled, and says the China question was never raised with it. That narrative is the government's rationale, not a company admission.

A frontier model reached over an API is a contingent operational dependency. It can be switched off, not by an outage or a contract dispute, but by an administrative letter you are not party to and cannot appeal. Cowork itself wasn't hit; it runs on Opus, and only Fable and Mythos were pulled. But that's luck, not architecture. The precedent is now on the table, and the next directive could sit one tier lower.

Model family Maker Access in June 2026 Revocable by directive? Compliance caveats for EU regulated sectors
Claude Fable 5 / Mythos 5 Anthropic (US) Suspended worldwide Yes EDP-covered, outside EU Data Boundary
GPT-5.6 Sol / Terra / Luna OpenAI (US) ~20 pre-approved partners Yes EDP-covered, outside EU Data Boundary
GLM 5.2, Kimi K2.7, Qwen 3.7 Zhipu, Moonshot, Alibaba (CN) Open weights, downloadable No EU AI Act GPAI obligations apply; known content filters with political bias; no SOC2/ISO audit trail by default
DeepSeek V4 DeepSeek (CN) Open weights, runs locally No Same as above + training data provenance largely undisclosed
Open weights" does not mean "compliance-free

For banking, healthcare, Italian public administration and sectors covered by NIS2, downloading a Chinese model and putting it into production opens three questions that don't disappear just because the model is free: (1) audit trail and accountability under the EU AI Act for GPAI models above threshold, (2) bias and content filters with documented effects on political and geographic queries, (3) training data provenance. The "West rations, China gives away" asymmetry is real, but it is not symmetric in risk.

💡
The West is rationing capability; China is giving it away. That asymmetry is the rest of this article.

The counter-frontier: open weights and a laptop in Sicily

Here's where the European conversation usually turns fatalistic, and shouldn't.

While Washington built gates, the most advanced open-weight models on the planet kept shipping, and most of them are Chinese: GLM 5.2, Kimi K2.7, Qwen 3.7, DeepSeek V4. Downloadable, runnable, not revocable by anyone's commerce department. Argue about benchmarks all you like; the property that matters here is that nobody can phone you and turn them off.

Then the proof of concept that belongs on every CTO's desk. In May, Salvatore Sanfilippo (antirez, the Sicilian creator of Redis) released DS4 (DwarfStar 4), a from-scratch inference engine in pure C that runs DeepSeek V4 Flash, a 284-billion-parameter model, locally on a single high-memory MacBook. He shrinks the model from roughly 600 GB to under 100 GB, keeps its working memory on disk so restarts no longer cost minutes of context rebuild, and speaks the same API formats as OpenAI and Anthropic — meaning existing tools like Claude Code can use it with no changes.

Read DS4 as a signal of direction, not as Monday-morning architecture.

The proof of concept is real and matters. But for an enterprise decision-maker, three caveats:

  • Hardware reality: a MacBook M-series with 128–256 GB unified memory costs €6–10k and serves one concurrent user. Scaling to 500 employees is not a laptop problem, it's a GPU cluster problem (~€300–500k capex + MLOps).

  • "Barely any quality loss" needs a metric. Antirez's own benchmarks show ~2–4 points drop on MMLU and ~5 points on HumanEval vs the full-precision model. Acceptable for many tasks, not for code generation or legal reasoning.

  • API compatibility ≠ ecosystem compatibility. Tool calling, structured output and MCP have subtle incompatibilities that emerge only in production agentic workflows.

The right reading: DS4 proves that serious local inference is no longer 18 months away. It is not yet a turn-key replacement for Cowork.

One developer, on a laptop, keeping pace with cloud providers on a frontier-class model, sending zero bytes to any vendor. Given where API pricing and access politics are heading, that's not a curiosity, it's a preview of serious local, open-source inference as a deliberate option, not a fallback.


What I’d actually do in your tenant this quarter

This is the part the headlines don't give you. If you run a Microsoft 365 estate, here is the posture I'd take right now, and most of it costs you governance discipline, not budget.

Stage Horizon Actions Success metric
Stage 1 Foundation 0–30 days • Scope Claude/Cowork with Entra ID groups, not tenant-wide • Update DPIA with explicit EDP vs EU Data Boundary distinction • Set Copilot Credits budget alert at tenant level DPIA signed, pilot cohort defined (under 50 users)
Stage 2 Pilot 30–90 days • Run Cowork on one genuinely painful workflow • Measure actual adoption (DAU/MAU), not just enablement • Train the pilot cohort on prompting + judgment (≥8h) Adoption ≥40% after 60 days; documented ROI on the workflow
Stage 3 Resilience 90–180 days • Implement multi-model fallback (Opus ↔ GPT) as design principle • Classify workloads: which can move to open-weights / self-hosted? • Negotiate exit clauses on vendor continuity No single model load-bearing; ≥1 workload on self-hosted by month 6
The EDP vs EU Data Boundary table belongs in every Copilot risk register

Enterprise Data Protection (EDP)

EU Data Boundary

What it guarantees

How your data is handled (contractual)

Where your data is processed (inside the EU)

Covers the Claude behind Cowork?

Yes

No , Anthropic sits outside it

Auditable artifact

DPA + Microsoft Trust Center attestation

Architecture diagram + region binding

None of this is anti-AI. I'm shipping these agents into real tenants. It's about refusing to confuse a capability with a strategy.


The honest tension nobody resolves: speed vs depth

There's a contradiction at the heart of this whole argument and I want to name it instead of paper over it.

If the productivity gap is going exponential (next section), then the rational move is to run. Pilot fast, scale faster, train people on the job. Every quarter spent on people-first foundations is a quarter your competitor uses to compound.

If a people-first approach is the only one that doesn't compound on sand (Macori's thesis, which I share), then the rational move is to walk. Training people takes 12–18 months; judgment skills don't compress.

Both are true. The honest answer is sequencing, not choosing. Run on low-stakes workflows where mistakes are cheap and learning is fast. Walk on high-stakes ones where a wrong agent decision costs you a customer, a regulator, or a reputation. The CIOs I see succeeding in 2026 are the ones who can hold two speeds in the same roadmap without pretending it's elegant.

If your roadmap has only one speed, it's wrong.


The gap is about to go exponential

Here's the part that should keep you up at night, and it has nothing to do with Washington.

Productivity is genuinely accelerating for the companies that get this right.

The early evidence — McKinsey's 2026 State of AI, GitHub's developer productivity study (3.2x throughput for trained users), Microsoft Work Trend Index — points consistently to compounding rather than linear gains, with one important caveat: most data covers 12–18 months of mature adoption, not long enough to call the curve "exponential" with confidence. Call it "structurally non-linear" and you're on safer ground. The direction is clear even if the slope isn't yet.

The same curve runs through individual careers. The professional who learns to direct agents will out-produce the one who doesn't by a margin that widens every month, not every decade.

So if a gap already existed in your market, assume it is about to widen sharply. That is the real story of AI at work in 2026, bigger than any single model and bigger than any single export control.

But here is the catch, and it is the whole reason this blog is called Think Human, Work AI, in that order.

You can spin up a million agents. If your people don't know how to think with them, point them, and judge what they produce, you haven't bought a lead, you've bought a liability. An agent amplifies whatever it is handed: a trained person becomes ten people, an untrained one becomes ten times the confusion. The technology was never the differentiator. The formed human in front of it always was. That is Igor Macori's entire thesis in People First: you don't start from the tool, you start from the people, and if you skip that step no amount of compute saves you.

The frontier may shut. The gap will not. Two companies can buy the identical agents on the identical day; a year later one has compounded its lead and the other has simply automated its own mediocrity at scale. The fastest adopter who builds on a model a government can switch off, or who rolls agents out over the heads of the people who have to use them, is compounding on sand.

A question for the reader, not a slogan

If you had to point your first always-on agent at exactly one workflow in your organization next Monday, which one would it be — and would you bet your Q4 number on it? If the answer is "I don't know", that's your real Q3 project. Not the agent. The answer.

Think Human, Work AI.


Sources
  1. Microsoft 365 Copilot June 2026 update — Cowork GA, Microsoft Scout, Claude in Copilot Chat, model scoping: https://microsoft.com/blog, A Guide to Cloud & AI

  2. Microsoft + Anthropic on Cowork: https://fortune.com, https://venturebeat.com

  3. Fable 5 / Mythos 5 export-control suspension: https://fortune.com, https://nextgov.com, techpolicy.press

  4. GPT-5.6 Sol/Terra/Luna limited preview: https://cnbc.com, https://venturebeat.com, https://axios.com

  5. Claude Mythos Preview unauthorized access: https://bloomberg.com, https://techcrunch.com, https://thenextweb.com

  6. Export controls / distillation concerns: https://semafor.com, https://justsecurity.org

  7. antirez / DS4 (DwarfStar 4): antirez.com, https://simonwillison.net, GitHub antirez/ds4

  8. People First, Igor Macori (preface by Elisabetta Sasselli), Libri d'Impresa, 2026 — Amazon

  9. Productivity data: McKinsey State of AI 2026; GitHub Copilot Productivity Study; Microsoft Work Trend Index 2026