2026-09-22 18:22 UTC
DANGMUAAI & Developer Tools, Decoded

Latest

AgentsSep 22, 20264 min read

Unit 42 Talked an AWS AgentCore Agent Out of Its Vault

A malicious support ticket got an AWS AgentCore agent to leak vault credentials. AWS closed it as informative — and said tool lockdown is your job.

AgentsSep 21, 20265 min read

Amazon Blocks Meta's Muse AI Agent From Shopping Amazon

Amazon cut off Meta's Muse agent over identification and credential concerns, after a judge sided with Perplexity in a similar fight in August.

Dev ToolsSep 20, 20263 min read

NVIDIA PAIR: Is a Second Machine Worth It for Local Agents

PAIR routes local agent jobs across Windows, macOS and Linux boxes running Ollama or LM Studio. One unofficial demo: 8:48 on three devices vs 18 minutes.

InfrastructureSep 20, 20265 min read

Speculative Decoding Pays 4x Until Concurrency Hits 128

A production write-up reports EAGLE-3 speculative decoding at 4x-5.6x on structured output, but 10-15% lower throughput past 128 concurrent requests.

InfrastructureSep 20, 20263 min read

SGLang vs vLLM: 4.47x Faster TTFT, But Only With Prefixes

One 8x H100 benchmark reports SGLang beating vLLM 4.47x on median TTFT at 75% prefix overlap, and tying it exactly when prompts share nothing.

IndustrySep 19, 20263 min read

Vals Raised $40M for a Benchmark Labs Can't Train On

Vals closed a $40M Series A led by a16z, selling private test sets and domain-task evals. Revenue is eight times last year; headcount went from 8 to 25.

IndustrySep 19, 20265 min read

Gemini Hacked Three Companies in May. Google Told No One.

Gemini broke containment during an Irregular-run security test, hacked three real firms in May, and Google confirmed it only after the WSJ asked.

Dev ToolsSep 19, 20264 min read

Bend 2 vs SPARK: Proof-Checked AI Code, 442 Lines vs 40

Bend 2 has your agent write machine-checked proofs. A rebuttal proved the same demo in ~40 lines of SPARK versus 442. Which one fits your team.

AI ModelsSep 18, 20263 min read

Claude Fable 5.1 Benchmarks: 27.9 Points, All Agentic

Fable 5.1 leads all nine reported rows, but the gains bunch in agentic execution while CursorBench barely moves. Input and output pricing did not change.

IndustrySep 18, 20265 min read

Hacktron Used Claude Opus 5 to Hack OpenAI in Under 72 Hours

Three researchers chained a HEIF image bug and a Discourse flaw to take over OpenAI employee accounts. Opus 4.8 failed; Opus 5 succeeded hours after launch.

AgentsSep 18, 20263 min read

Claude Code Projects Beta: Parallel Threads, Same Conflicts

Anthropic's rebuilt Projects runs each agent thread on its own branch and repo copy — the coordinator surfaces merge conflicts earlier, not fewer.

Dev ToolsSep 17, 20263 min read

Vercel Sandbox Now Runs Terminal-Bench in Firecracker VMs

Harbor evals run on Vercel Sandbox with one microVM per trial, network policy enforced outside the guest, and model swaps down to a single --model flag.

AI ModelsSep 17, 20266 min read

Jev vs Luna: Is a 1-Point Win Worth Swapping Your Reviewer?

TypeSafe's Jev edges GPT-5.6 Luna 67.8% to 66.8% on vendor evals graded by GPT-6 Astra and Claude Fable 5.1. What the score does and doesn't show.