AI News, Explained
Sunday night Amazon began blocking Meta's Muse from Amazon checkouts, calling it unauthorized AI-agent access under its Conditions of Use (Amazon says Muse doesn't identify itself and appears to capture credentials; Meta cites VM isolation + Sentinel monitoring). Monday, Shopify partnered with Meta so Muse can complete purchases via Shop Pay across its stores. One door slams, another opens.
Sources: Barron's
CEO Eddie Wu announced plans for a 5–10 trillion parameter model and unveiled the Zhenwu V900 AI chip — claimed "most powerful in China," 3× its predecessor (vendor claim). Cloud data-center target: above 20 GW by 2032. Hype Check verdict: unproven — a plan announcement, not a launch; the chip is a testable hardware claim, the model is a press promise.
Sources: Reuters
Native voice models via Gemini API, AI Studio, Workspace and the Gemini app. Google claims the Extended Thinking variant tops Artificial Analysis' speech-to-speech leaderboard at 82.6% vs 81.5% for OpenAI's GPT-Live-1 (independent replication pending). Pricing $0.005/$0.018 per min in/out — far under OpenAI's voice pricing; no full duplex.
Sources: The Decoder
Legal-vertical platform on GPT-6 Astra: 230M-URL US case-law index, Free Law Project/CourtListener partnership, legal-writing instructions. Harvey and Legora can build on it; integrations with Relativity, Clio, Intapp, Thomson Reuters; early program with Sullivan & Cromwell, Ropes & Gray, Cooley, Latham & Watkins, Wachtell Lipton. OpenAI measured 54.0% on Vals' Legal Research Bench vs 38.7% for Astra with plain web search — independent verification pending.
Sources: Reuters
Launched Sept 22 — first model in the Claude 5.5 family. Anthropic claims new state-of-the-art in coding and knowledge work, outscoring the larger Fable 5.1 on many tests (Terminal-Bench 4.0 66.4%, FrontierCode 54.4%, CursorBench 57.8%, GDPval-AA 1846, HLE 67.7% with tools) — all self-reported, no independent rerun yet. Priced at $4/$20 per million tokens (20% below Opus 5), cache reads cut 60% to $0.20, ~40% lower operating cost on typical workloads, 30%+ faster generation. Fable-level safeguards; 85% less likely to bypass boundaries in an internal eval. Sonnet 5.5 + Haiku 5.5 in the coming weeks. Hype Check verdict: unproven — the price list is confirmed, the benchmarks await independent testing.
Sources: TechCrunch, Reuters, The Decoder
Launched Sept 21 — 2.1T parameters (company claim), 500K context, $2/$6 per million input/output tokens, unchanged pricing. xAI's self-run table claims 71% DeepSWE and 38% Terminal-Bench, but independent runs show a wide gap (Artificial Analysis: 46 vs 53 for GPT-6 and Fable 5.1; Terminal-Bench 26% vs claimed 38%). Genuine bright spots: Harvey Legal 19.6% leads the field, AA-Briefcase 1657 Elo is top-tier, and Cursor measured $4.69 average task cost — roughly half the ~$9 rivals. Hype Check verdict: overblown.
Sources: AI Release Tracker, The Decoder
Launched Sept 8–9 in the US (iOS, Android, muse.ai, WhatsApp). Runs on Muse Spark; books travel, sends email, negotiates bills, purchases via Stripe Link one-time cards. Free tier + paid plans; glasses support planned.
Sources: TechCrunch, Reuters
First CRM reasoning model for Agentforce, built on NVIDIA Nemotron 3 Super and 27 years of CRM data. Matches or beats leading models on CRM actions with 3× fewer errors; open-weight; fewer tokens per task.
Sources: Salesforce, TechCrunch
Dual models (Flare for speed, Sunburst for precision), 4K output, multi-turn editing. High-quality tier uses ~¼ the tokens of the old high setting.
Source: Reuters
Real-time multilingual transcription in 80ms chunks, 70+ languages, 20+ speaker labeling. Meta claims #1 on Artificial Analysis' streaming speech-to-text leaderboard (vendor claim, third-party board).
Source: TechRepublic
Chat and Cowork folded into one interface; new Claude Docs and Claude Slides (beta) with export to Google Docs, Word, PowerPoint, PDF. Pro/Max first.
Source: Reuters
Testing business-sponsored agents inside ChatGPT (clearly labeled, separate from ChatGPT's answers). Ads Manager builds campaigns from plain-English prompts; HubSpot and Shopify first partners.
Source: Reuters
| Model ⇅ | Task ⇅ | Cost signal ⇅ | Caveat ⇅ |
|---|---|---|---|
| Claude Opus 5.5 (Anthropic) | Frontier coding + knowledge work | $4/$20 per 1M tokens; ~40% lower operating cost vs Opus 5; cache reads $0.20 (60% cheaper) | Price list confirmed; benchmark scores vendor-reported, no independent rerun yet |
| Grok 4.7 (xAI) | Coding agent tasks | ~$4.69 / task on Cursor — roughly half the ~$9 rivals | Cursor's own Sept evaluation; xAI launch scores self-reported |
| Gemini 3.8 Live (Google DeepMind) | Native voice conversation | $0.005/$0.018 per min in/out — far under OpenAI voice pricing | Leaderboard top spot from Google's own chart; independent replication pending |
| Koa (Salesforce/NVIDIA) | CRM reasoning | 3× fewer errors than leading models; fewer tokens per task | Salesforce's own benchmark |
| GPT Image 2.5 Flare | Fast image generation | ≈ $0.006 / image (low quality, 1024²) | Token-math estimate |
| GPT Image 2.5 Sunburst | Precise image gen + editing | ≈ $0.211 / image (max quality); high tier ~4× cheaper than old high | Token-math estimate |
| Muse Voice Transcribe | Real-time multilingual STT | 80ms chunks; claims #1 streaming STT | Vendor claim, third-party board |
| GPT-6 Astra (OpenAI) | Frontier assistant | Powers ChatGPT for Financial Services | Speed/accuracy claims vendor-stated |
Through-line: frontier intelligence keeps getting cheaper by the token — voice at fractions of a cent per minute, coding at half the cost — while the agents spending those tokens get more autonomous.