← Agile Factor

GRADIENT PULSE

AI News, Explained

Release radar

Agentic commerce · Meta / Amazon / Shopify

Amazon blocks Muse; Shopify opens the door

Sunday night Amazon began blocking Meta's Muse from Amazon checkouts, calling it unauthorized AI-agent access under its Conditions of Use (Amazon says Muse doesn't identify itself and appears to capture credentials; Meta cites VM isolation + Sentinel monitoring). Monday, Shopify partnered with Meta so Muse can complete purchases via Shop Pay across its stores. One door slams, another opens.

Sources: Barron's

Alibaba · Chip + plan

Zhenwu V900 + 5–10T parameter plan

CEO Eddie Wu announced plans for a 5–10 trillion parameter model and unveiled the Zhenwu V900 AI chip — claimed "most powerful in China," 3× its predecessor (vendor claim). Cloud data-center target: above 20 GW by 2032. Hype Check verdict: unproven — a plan announcement, not a launch; the chip is a testable hardware claim, the model is a press promise.

Sources: Reuters

Google DeepMind · Voice

Gemini 3.8 Live voice models

Native voice models via Gemini API, AI Studio, Workspace and the Gemini app. Google claims the Extended Thinking variant tops Artificial Analysis' speech-to-speech leaderboard at 82.6% vs 81.5% for OpenAI's GPT-Live-1 (independent replication pending). Pricing $0.005/$0.018 per min in/out — far under OpenAI's voice pricing; no full duplex.

Sources: The Decoder

OpenAI · Vertical

Astra for Law

Legal-vertical platform on GPT-6 Astra: 230M-URL US case-law index, Free Law Project/CourtListener partnership, legal-writing instructions. Harvey and Legora can build on it; integrations with Relativity, Clio, Intapp, Thomson Reuters; early program with Sullivan & Cromwell, Ropes & Gray, Cooley, Latham & Watkins, Wachtell Lipton. OpenAI measured 54.0% on Vals' Legal Research Bench vs 38.7% for Astra with plain web search — independent verification pending.

Sources: Reuters

Anthropic · Frontier model

Claude Opus 5.5

Launched Sept 22 — first model in the Claude 5.5 family. Anthropic claims new state-of-the-art in coding and knowledge work, outscoring the larger Fable 5.1 on many tests (Terminal-Bench 4.0 66.4%, FrontierCode 54.4%, CursorBench 57.8%, GDPval-AA 1846, HLE 67.7% with tools) — all self-reported, no independent rerun yet. Priced at $4/$20 per million tokens (20% below Opus 5), cache reads cut 60% to $0.20, ~40% lower operating cost on typical workloads, 30%+ faster generation. Fable-level safeguards; 85% less likely to bypass boundaries in an internal eval. Sonnet 5.5 + Haiku 5.5 in the coming weeks. Hype Check verdict: unproven — the price list is confirmed, the benchmarks await independent testing.

Sources: TechCrunch, Reuters, The Decoder

xAI · Frontier model

Grok 4.7

Launched Sept 21 — 2.1T parameters (company claim), 500K context, $2/$6 per million input/output tokens, unchanged pricing. xAI's self-run table claims 71% DeepSWE and 38% Terminal-Bench, but independent runs show a wide gap (Artificial Analysis: 46 vs 53 for GPT-6 and Fable 5.1; Terminal-Bench 26% vs claimed 38%). Genuine bright spots: Harvey Legal 19.6% leads the field, AA-Briefcase 1657 Elo is top-tier, and Cursor measured $4.69 average task cost — roughly half the ~$9 rivals. Hype Check verdict: overblown.

Sources: AI Release Tracker, The Decoder

Meta · Agent

Muse personal AI agent

Launched Sept 8–9 in the US (iOS, Android, muse.ai, WhatsApp). Runs on Muse Spark; books travel, sends email, negotiates bills, purchases via Stripe Link one-time cards. Free tier + paid plans; glasses support planned.

Sources: TechCrunch, Reuters

Salesforce + NVIDIA · Reasoning model

Koa

First CRM reasoning model for Agentforce, built on NVIDIA Nemotron 3 Super and 27 years of CRM data. Matches or beats leading models on CRM actions with 3× fewer errors; open-weight; fewer tokens per task.

Sources: Salesforce, TechCrunch

OpenAI · Image generation

GPT Image 2.5 — Flare & Sunburst

Dual models (Flare for speed, Sunburst for precision), 4K output, multi-turn editing. High-quality tier uses ~¼ the tokens of the old high setting.

Source: Reuters

Meta · Speech

Muse Voice Transcribe

Real-time multilingual transcription in 80ms chunks, 70+ languages, 20+ speaker labeling. Meta claims #1 on Artificial Analysis' streaming speech-to-text leaderboard (vendor claim, third-party board).

Source: TechRepublic

Anthropic · Product

Unified Claude + Docs/Slides beta

Chat and Cowork folded into one interface; new Claude Docs and Claude Slides (beta) with export to Google Docs, Word, PowerPoint, PDF. Pro/Max first.

Source: Reuters

OpenAI · Advertising

Sponsored agents + Ads Manager

Testing business-sponsored agents inside ChatGPT (clearly labeled, separate from ChatGPT's answers). Ads Manager builds campaigns from plain-English prompts; HubSpot and Shopify first partners.

Source: Reuters

Performance per dollar

Click a column header to sort. Figures are vendor-stated unless noted.

Model ⇅Task ⇅Cost signal ⇅Caveat ⇅
Claude Opus 5.5 (Anthropic)Frontier coding + knowledge work$4/$20 per 1M tokens; ~40% lower operating cost vs Opus 5; cache reads $0.20 (60% cheaper)Price list confirmed; benchmark scores vendor-reported, no independent rerun yet
Grok 4.7 (xAI)Coding agent tasks~$4.69 / task on Cursor — roughly half the ~$9 rivalsCursor's own Sept evaluation; xAI launch scores self-reported
Gemini 3.8 Live (Google DeepMind)Native voice conversation$0.005/$0.018 per min in/out — far under OpenAI voice pricingLeaderboard top spot from Google's own chart; independent replication pending
Koa (Salesforce/NVIDIA)CRM reasoning3× fewer errors than leading models; fewer tokens per taskSalesforce's own benchmark
GPT Image 2.5 FlareFast image generation≈ $0.006 / image (low quality, 1024²)Token-math estimate
GPT Image 2.5 SunburstPrecise image gen + editing≈ $0.211 / image (max quality); high tier ~4× cheaper than old highToken-math estimate
Muse Voice TranscribeReal-time multilingual STT80ms chunks; claims #1 streaming STTVendor claim, third-party board
GPT-6 Astra (OpenAI)Frontier assistantPowers ChatGPT for Financial ServicesSpeed/accuracy claims vendor-stated

Through-line: frontier intelligence keeps getting cheaper by the token — voice at fractions of a cent per minute, coding at half the cost — while the agents spending those tokens get more autonomous.