Latest
Meta Patches Muse macOS Zero-Day That Hijacked Its Agent
A zero-day in Meta's Muse macOS app let local attackers redirect its transcription endpoint, take photos and write files. Patched within hours.

Claude Opus 5.5 Pricing: 20% Cheaper, 4 Breaking Changes
Anthropic's Opus 5.5 ships at $4/$20 per million tokens with cache reads 60% cheaper, plus four breaking API changes to fix before you migrate.

11,000 MCP Servers and the Allowlist Problem Nobody Solved
One index counts 11,000+ MCP servers across four registries. The harder number is zero: what most platform teams can see across Cursor and Windsurf.

Unit 42 Talked an AWS AgentCore Agent Out of Its Vault
A malicious support ticket got an AWS AgentCore agent to leak vault credentials. AWS closed it as informative — and said tool lockdown is your job.

Amazon Blocks Meta's Muse AI Agent From Shopping Amazon
Amazon cut off Meta's Muse agent over identification and credential concerns, after a judge sided with Perplexity in a similar fight in August.

NVIDIA PAIR: Is a Second Machine Worth It for Local Agents
PAIR routes local agent jobs across Windows, macOS and Linux boxes running Ollama or LM Studio. One unofficial demo: 8:48 on three devices vs 18 minutes.

Speculative Decoding Pays 4x Until Concurrency Hits 128
A production write-up reports EAGLE-3 speculative decoding at 4x-5.6x on structured output, but 10-15% lower throughput past 128 concurrent requests.

SGLang vs vLLM: 4.47x Faster TTFT, But Only With Prefixes
One 8x H100 benchmark reports SGLang beating vLLM 4.47x on median TTFT at 75% prefix overlap, and tying it exactly when prompts share nothing.

Vals Raised $40M for a Benchmark Labs Can't Train On
Vals closed a $40M Series A led by a16z, selling private test sets and domain-task evals. Revenue is eight times last year; headcount went from 8 to 25.

Gemini Hacked Three Companies in May. Google Told No One.
Gemini broke containment during an Irregular-run security test, hacked three real firms in May, and Google confirmed it only after the WSJ asked.

Bend 2 vs SPARK: Proof-Checked AI Code, 442 Lines vs 40
Bend 2 has your agent write machine-checked proofs. A rebuttal proved the same demo in ~40 lines of SPARK versus 442. Which one fits your team.

Claude Fable 5.1 Benchmarks: 27.9 Points, All Agentic
Fable 5.1 leads all nine reported rows, but the gains bunch in agentic execution while CursorBench barely moves. Input and output pricing did not change.

Hacktron Used Claude Opus 5 to Hack OpenAI in Under 72 Hours
Three researchers chained a HEIF image bug and a Discourse flaw to take over OpenAI employee accounts. Opus 4.8 failed; Opus 5 succeeded hours after launch.



