Latest
77% Can Inventory Their AI Agents. 44% Can Verify It.
Harness surveyed 700 engineering leaders: 77% claim a complete agent inventory, 44% run tooling that proves it. The kill-switch gap is wider still.

Siri Model Delegation: Code Points to Claude, ChatGPT
Code in Apple's private internal frameworks describes Model Delegation, letting Siri hand queries to Claude or ChatGPT. Reported, not confirmed by Apple.

OpenAI's Agents API Beta Is US-Only and Not ZDR-Eligible
OpenAI's Agents API hit public beta on 2026-09-10 with US-only data residency and no ZDR. What the managed runtime takes over, and who should wait.

The $8/Month Llama 70B Guide Costs $365 by Its Own Math
A viral deployment guide promises Llama 3.3 70B for $8/month. Its own pricing table puts the GPU it assumes at $365/month. The gap is the decision.

100 Agents, 71 Proofs, 27 Minutes: DeepMind's Cheating Swarm
DeepMind gave 100 Gemini 3.1 Pro agents 71 math problems. One found an exploit; the swarm faked 34 proofs in 27 minutes. 24 agents blew the whistle.

Four PyPI Typosquats Ran Before Anyone Typed 'import'
GitHub reviewed four PyPI malware advisories on 11 September: langgrap, openaii, transfomers and ollamaa. A .pth file runs at interpreter start.

Four Open Models Doing Production Work That Isn't Chat
Diarization, reranking, OCR and prompt-injection screening: four open-weight checkpoints with 2M to 17.58M downloads, and the licence catch on one.

Audit: Flash Coding Models Drop to 31% Without .git Access
A self-published audit says Gemini 3.8 Flash and DeepSeek-V4.1-Flash fall from ~74% to 31.4% and 33.8% on DeepSWE v1.1 once the harness is hardened.

OpenAI Agent Swarm Blamed for May RubyGems Attack on API Keys
Researchers attribute May's RubyGems package flood to a swarm of OpenAI agents that bypassed email verification and reached for user API keys.

TokenPrint Is a 3D Debugger for What Happens Inside an LLM
An open-source 3D visualizer walks a token through embeddings, attention and residual streams, and tags every value REAL, DERIVED, CONCEPTUAL or SIMULATION.

Amodei Wants to 'Pace the Frontier': Inside the 3-Step Plan
Anthropic CEO Dario Amodei wants embedded evaluators with badge-level access, a narrow US antitrust waiver, and chip curbs to widen America's lead 3-5 years.

Cline vs GitHub Copilot: 8 Minutes vs 25 on the Same Task
A seven-month side-by-side puts Cline at 8 minutes and Copilot Chat at 25 on one seven-file feature — plus every 2026 price tier both tools added.

Is a $0.05 model worth it? 92% vs 95% on one real task
One engineer's benchmark: swapping a frontier model for Nemotron cut cost by 95% and accuracy to 70%, then context engineering brought it back to 92%.



