Skip to content
BlackOak AgencyBlackOak.agency

// AI RADAR

AI Radar

What is shipping in AI across models, agents, video and capital — with what it is, why it matters and what to do with it on Monday. Below the brief sits the landscape: prices, tools and numbers that hold up longer.

The week in artificial intelligence

Seven models in seven days. The question is no longer which model wins.

Between 7 and 24 July five frontier models shipped, Sora left the market, Claude Cowork moved from desktop to server, and a German image lab put its model on an Audi production line. Here is what holds up once the dust settles — and what to do with it on Monday.

July 2026 — the release calendar

← drag to scroll

  1. AnthropicCowork on web and mobile07 JUL
  2. xAIGrok 4.508 JUL
  3. OpenAIGPT-5.6 — Sol, Terra, Luna09 JUL
  4. Helsing$1.8B13 JUL
  5. MoonshotKimi K3 — 2.8T params16 JUL
  6. AlibabaAgent Native Cloud18 JUL
  7. GoogleThree Gemini Flash models21 JUL
  8. Black Forest LabsFLUX 323 JUL
  9. AnthropicClaude Opus 524 JUL
  10. MoonshotKimi K3 open weights27 JUL

The brief

Six developments that change your stack

By date, with what happened, why it matters and what to do next.

07.07.2026Agents

Cowork leaves the desktop — and keeps working after you shut down

Anthropic moved Claude Cowork — desktop-only since January — to the web and to iOS and Android, in beta, starting with Max subscribers. The new surfaces run a materially different execution model: sessions live on Anthropic’s servers, belong to your account, and keep running with no device online. Scheduled tasks included.

Why it matters

This is the first time delegation is genuinely asynchronous. Not "the agent works while I watch", but "the job is done by morning". Anthropic’s own usage data shows where it lands: 33.4% business process and operations, 16.4% content and copywriting. That is exactly the work an agency still does by hand.

What next

Give an agent a folder, never your drive. Anthropic itself warns about destructive actions on ambiguous instructions and about indirect prompt injection through files and websites. Human review on anything outbound — mail, publishing, payment — is not caution, it is design.

09.07.2026Models

GPT-5.6 arrives in three tiers: Sol, Terra and Luna

On 9 July OpenAI made all three tiers generally available through the API and Codex, after a preview from 26 June. Sol is the flagship ($5/$30 per million tokens), Terra the everyday workhorse ($2.50/$15), Luna the volume tier ($1/$6). All three: 1M token context, 128K output, knowledge cutoff 16 February 2026. The Responses API gained programmatic tool calling.

Why it matters

The mini and nano suffixes are gone. Sol, Terra and Luna are durable capability tiers that can each advance at their own pace. In practice: you route per task on price rather than per project on model name, and your prompts survive the next release.

What next

Put Luna on classification, extraction and bulk summarisation; Terra on editorial work; Sol only where reasoning is genuinely worth the money. The gap between Luna and Sol is five times on input and five times on output.

16.07.2026Open

Kimi K3 is open — and fits on no machine you own

Moonshot AI shipped Kimi K3 through its API on 16 July and publishes the full weights on 27 July under a modified MIT licence. It is a mixture-of-experts with 2.8 trillion total parameters, roughly 50 billion active per token, a 1M token context, and text, image and video input. In the Frontend Code Arena K3 took first place with 1,679 points, ahead of Fable 5 (1,631) and GPT-5.6 Sol (1,618).

Why it matters

"Open" is not the same as "self-hostable". Even at four-bit precision K3 needs about 1.4 terabytes of fast memory — that means Blackwell or MI400 server nodes, not a workstation under your desk. The bill shifts from API subscription to hardware and operations, and for most agencies that is the worse trade.

What next

Treat open weights as leverage, not as a migration plan. They push hosted prices down and give you an exit under vendor risk. You run them at a provider that already owns the nodes.

23.07.2026Video

FLUX 3 learns image, video, sound and action in one architecture

On 23 July Black Forest Labs announced FLUX 3: a single multimodal foundation model that learns image, video and audio jointly, plus action prediction. The headline is video up to twenty seconds with native, in-sync audio in one generation — including multilingual dialogue, across styles and aspect ratios. Video opened in gated early access on day one, Image follows in the coming weeks, the open-weight Dev release comes last. A robotics variant is being tested on Audi production lines.

Why it matters

For an agency, picture-plus-sound in one generation removes an entire edit step. No separate voice-over laid over the footage, no lip-sync correction afterwards. That an image lab is simultaneously moving into robotics says something about where generative models are heading: from picture to action.

What next

The evaluations are early — the lab says so itself. Use FLUX 3 for pitch material and social where twenty seconds is enough; keep Veo or Kling alongside while you wait on gated access.

24.07.2026Models

Claude Opus 5 wins the benchmark that actually matters for agents

Anthropic shipped Opus 5 on 24 July at the same price as Opus 4.8: $5 in, $25 out per million tokens, with a Fast mode at double the rate running roughly 2.5× faster. One million tokens of context, 128K output, knowledge cutoff May 2026 — the most current of any Claude model. On Frontier-Bench v0.1 it more than doubles Opus 4.8; on ARC-AGI 3 it scores three times the runner-up; on OSWorld 2.0 it beats Fable 5’s best result at just over a third of the cost.

Why it matters

OSWorld measures computer use: clicking, forms, real screens. That is the benchmark that predicts whether an agent can take over your work, far more than a knowledge quiz. A jump there at a third of the cost changes which tasks you are willing to delegate at all.

What next

Opus 5 is the new default on Claude Max and the strongest model on Pro. If your own jobs still run on a 2025 model you are paying the same for less. Re-benchmark your two heaviest workflows this week.

H1 2026Market

AI search is no longer an experiment, it is a channel — and almost nobody measures it

ChatGPT passed 900 million weekly active users in late February 2026. HubSpot’s State of Marketing 2026 found half of consumers now use AI search, and that 44% of them call it their primary source for product discovery — ahead of traditional search at 31%. Meanwhile only 14% of marketers track how they perform in AI search.

Why it matters

The gap between "half your market searches here" and "14% measure it" is the whole story. Whoever can answer today, with numbers, what share of demand runs through AI answers is selling a measurement, not a trend.

What next

Start with a baseline: for which questions are you cited in ChatGPT, Perplexity, Gemini and Google’s AI overviews, and from which source? That is a report you can produce today and almost nobody has.

Signals

Short, but not small

21 JUL · GOOGLE

Three Flash models, no flagship

Gemini 3.6 Flash (up to 17% fewer tokens), 3.5 Flash-Lite and a security-tuned 3.5 Flash Cyber for governments and partners. The 3.5 Pro promised "next month" back in May is still with test partners. Pre-training for Gemini 4 has started.

26 APR / 24 SEP · OPENAI

Sora has left the market

Web and app shut down on 26 April 2026, the API goes dark on 24 September. The category OpenAI opened is now split between Veo, Kling, Seedance and Runway. Anyone still building on the Sora API has two months.

2026 · MCP

From protocol to infrastructure

41% of surveyed software organisations run MCP servers in limited or broad production. The registry counted 9,652 servers in late May; the Python and TypeScript SDKs together see roughly 97 million downloads a month. Governed since December 2025 by the Agentic AI Foundation under the Linux Foundation.

18–22 JUL · ENTERPRISE

Agents move into the existing stack

HubSpot opened Agent Hub and Agent Builder in public beta, Alibaba Cloud announced an Agent Native Cloud at WAIC with multi-agent orchestration and a sandbox, and Ushur launched a platform that runs customer journeys end to end. The shift: from demo to one real process with human review.

H1 2026 · KAPITAAL

$510 billion, and two names take 43%

Crunchbase counted a record $510B in global startup investment in the first half of 2026. OpenAI and Anthropic together raised $217B — 43% of every venture dollar deployed worldwide. Databricks signed a term sheet on 16 July at a $188B valuation.

Radar

The landscape, not the news

What is standing right now, with prices and characteristics. This part ages slower than the brief above — but always verify rates with the vendor.

Frontier models

price per million tokens · checked July 27, 2026
ModelLabDateContextIn $Out $Where it stands out
Claude Opus 5Anthropic24 jul1M525ARC-AGI 3 at 3× the runner-up; OSWorld 2.0 above Fable 5. Fast mode $10/$50, ~2.5× faster.
GPT-5.6 SolOpenAI9 jul1M530Flagship tier; programmatic tool calling in the Responses API.
GPT-5.6 TerraOpenAI9 jul1M2,5015The workhorse: GPT-5.5-level at about half the cost.
GPT-5.6 LunaOpenAI9 jul1M16Volume work: classification, extraction, bulk summarisation.
Grok 4.5xAI8 jul26Undercuts Opus 4.8 pricing by over 60%.
Gemini 3.6 FlashGoogle21 julUp to 17% fewer tokens than 3.5 Flash, and cheaper. 3.5 Pro still with partners.
Kimi K3Moonshot16 jul1M2.8T param MoE, ~50B active. Open weights 27 Jul (modified MIT). #1 Frontend Code Arena, 1,679 points.

Moving image

after Sora’s exit · checked July 27, 2026
ModelStrengthAudioPrice per secondStatus
Veo 3.1The strongest cinematic option across 2026 comparisonsNative 48 kHz synchronised dialogueLite $0,05 · Fast $0,15 · Std $0,40Available
Kling 3.0 TurboPhoneme-level lip-sync, including multiple characters≈ $0,11 – $0,14Available
Seedance 2.530 seconds in a single pass, up to 50 multimodal referencesAvailable
Runway Gen-4.5The most control over shot and directionAvailable
FLUX 3 Video20 seconds of picture + sound in one generation, multilingual dialogueNative, in syncGated early access
SoraOpened the categoryApp ended 26 Apr · API ends 24 Sep

The agent stack

what sits where · checked July 27, 2026

Protocol

  • MCPde facto standard, with the Agentic AI Foundation (Linux Foundation) since Dec 2025
  • 9,652 servers in the registry · ~97M SDK downloads a month · 41% of organisations in production
  • Natively supported by Anthropic, OpenAI, Google and Microsoft

Workspace

  • Claude Coworkdesktop, web and mobile; remote sessions and scheduled tasks
  • Agentic ComputerAlibaba’s sandbox for secure execution
  • CodexOpenAI, with GPT-5.6 across three tiers

Orchestration

  • LangGraph · CrewAI · Difycode-first multi-agent
  • AgentTeamsAlibaba’s multi-agent layer in Agent Native Cloud
  • n8nvisual, for teams without engineering

Inside the existing stack

  • HubSpot Agent Hubpublic beta, agents share customer context
  • Salesforce Agentforce · SAP Joule · ServiceNow
  • UiPath · Copilot Studio · Automation AnywhereRPA that went agentic

Where attention is going

numbers for the pitch · checked July 27, 2026
900M+

weekly ChatGPT users, late February 2026

OpenAI / ChatGPT
50%

of consumers now use AI search

HubSpot, State of Marketing 2026
44%

of those call AI search their primary source for product discovery, vs 31% traditional search

HubSpot, State of Marketing 2026
14%

of marketers track their AI search performance

HubSpot, State of Marketing 2026
$510B

global startup investment in H1 2026 — a record

Crunchbase
43%

of every venture dollar went to OpenAI and Anthropic combined

Crunchbase

Sources

Where this comes from

About this page. The brief above carries the date of its edition; each block in the radar below states when it was last checked. Everything comes from public sources, with the primary source where one is available — the list sits above. Prices, benchmark scores and release dates partly come from trade press and trackers; always verify rates with the vendor before budgeting on them. Benchmarks are snapshots and say nothing about your specific work: re-test on your own tasks.