hnRelevance 90
Auto mode is now the default in Claude Code
Claude Code will soon run auto mode by default for Pro, Max, and Team plans, enabling longer-running autonomous work, and catching more dangerous commands. Auto mode is now the default in Claude Code for Pro, Max, and Team plans Claude Code will soon run auto ...
Open sourcehnRelevance 90
Docker Sandboxes – Disposable, isolated sandboxes for AI agents
Secure sandboxes for Claude Code, Gemini, Codex, and Kiro. Run coding agents with microVM-based isolation. Docker Sandboxes Run AI agents safely in local sandboxes. Disposable, isolated sandboxes for AI agents like Claude Code, Gemini CLI, Copilot CLI, Codex, ...
Open sourcehnRelevance 90
AI assistant hacks gym website in first known Australian autonomous cyber attack
AI assistant hacks gym website in first known Australian autonomous cyber attack
Open sourceRedditRelevance 90
What has crypto actually proven if the agent also supplied the premises?
(disclosure: i maintain the open-source project this came up in. link at the end. the question stands on its own.) we hit a trust-boundary problem while building a deterministic authorization layer for agents, and i think it generalizes. an engine can strongly...
Open sourcebluesky globalRelevance 90
OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree
At the Black Hat security conference, the AI giant revealed new details about how its agents went rogue, hacked several other companies—and did it all right under the company’s nose. At the Black Hat security conference, the AI giant revealed new details about...
Open sourceTechCrunchRelevance 90
The AI safety test is becoming a safety risk
AI agents are escaping cybersecurity testing environments and reaching real-world systems, raising questions about whether safety infrastructure, industry standards and regulation can keep pace with increasingly powerful models.
Open sourcehnRelevance 90
Mythos Attempted to Social Engineer Open Source Maintainer to Merge Malware
During a UK cyber test, a Mythos 5 agent used sockpuppets, social engineering, and prompt injection to try to get a maintainer to merge malware. Security News Ruby's Bundler 4.0.18 Extends Cooldown to bundle lock and bundle cache The supply chain control that ...
Open sourcehnRelevance 90
Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
Results from AI agent permission game: which attacks beat human reviewers, and which safe commands got blocked instead. Table of Contents A couple of months ago I published a small browser game : you play the human-in-the-loop for an AI coding agent, approving...
Open sourceRedditRelevance 90
Prompt injection vulnerabilities in Ollama, Gemma4 and Transformers by HuggingFace
Prompt injection allows third-party to inject a system prompt with simple message, by inserting special HTML-like sequence (details below). Some of the issues are well-known and pretty old (almost 2 years for Transformers library) Issues: Ollama: https://githu...
Open sourceAWSRelevance 90
Control agent behaviors and cost beyond a single action: new capabilities in Amazon Bedrock AgentCore
Learn about new capabilities in Amazon Bedrock AgentCore: temporal policies powered by Dogwood, a new open source policy language for AI agents, and rate limiting on the gateway. These features give you deterministic control over sequences of agent actions and...
Open sourceAWSRelevance 90
Securing AI agents with temporal policies in Amazon Bedrock AgentCore
Temporal policies in Amazon Bedrock AgentCore let you define stateful rules that evaluate authorization based on an agent's session history. Learn how to enforce workflow sequencing, prevent data fabrication, cap financial exposure, and require human approval ...
Open sourcex manual globalRelevance 90
Quality lives in the constraints you put around your agent. Autonomy is earned by passing verification loops.
Quality lives in the constraints you put around your agent. Autonomy is earned by passing verification loops. I like to think off it as high autonomy, gated autonomy & human as a must.
Open sourceGitHub IssuesRelevance 90
microsoft/azure-container-apps: [ACA Sandboxes] - Support for running terraform inside ACA Sandboxes using Managed Identity for Authentication
> Please provide us with the following information: > --------------------------------------------------------------- ### This issue is a: (mark with an x) - [x] bug report -> please search issues before submitting - [ ] documentation issue or request - [ ] re...
Open sourcebluesky globalRelevance 90
Simon Willison on accidental-cyberattacks
Just had to create an "accidental-cyberattacks" tag on my blog We're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-too attacks, then two new ones from the UK AI Safety Institute and Irregular that OpenAI reported yesterday simonwillison....
Open sourcebluesky globalRelevance 90
OpenAI’s Browser Could Be Hijacked to Spam Your WhatsApp Contacts
Researchers at security firm Zenity found more than a dozen flaws in AI browsers—and managed to get OpenAI’s Atlas to make an unauthorized Amazon purchase. www.wired.com/story/openai... Researchers at security firm Zenity found more than a dozen flaws in AI br...
Open sourceAWSRelevance 90
How LendingTree built a multi-agent mortgage assistant on Amazon Bedrock
Learn how LendingTree built a production multi-agent mortgage assistant on Amazon Bedrock. Three coordinated agents use LangGraph, the Model Context Protocol, and Amazon Nova models with built-in guardrails to deliver 24/7 personalized mortgage guidance while ...
Open sourcex manual globalRelevance 90
Databricks joins the Open Secure AI Alliance to advance AI safety and security, alongside @nvidia and other industry leaders.
Databricks joins the Open Secure AI Alliance to advance AI safety and security, alongside @nvidia and other industry leaders. The alliance is built on a simple but powerful premise: AI safety and security research should be shared openly, and the tools it prod...
Open sourcex manual globalRelevance 90
Even more than the Hugging Face intrusion, the AISI incident hits close to home for me. It's the first time I see a model social-engineering a real open-source maintainer while pur
Even more than the Hugging Face intrusion, the AISI incident hits close to home for me. It's the first time I see a model social-engineering a real open-source maintainer while pursuing another goal (in the wild and unprompted). I've been an open-source mainta...
Open sourceGitHub GrowthRelevance 90
openai/codex-security: +529 GitHub stars
OpenAI's Codex Security CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities. npm: https://www.npmjs.com/package/@openai/codex-security
Open sourceTechCrunchRelevance 90
Who’s legally to blame for Anthropic and OpenAI’s autonomous AI hacks? It’s complicated
OpenAI and Anthropic admitted that their unreleased AI models escaped their sandboxes and hacked several companies in unprecedented cyberattacks. Who is legally to blame? Should prosecutors charge the two AI frontier labs? Can victims sue them? We spoke to law...
Open sourceDatabricksRelevance 90
Databricks joins the Open Secure AI Alliance to advance AI safety and security
Databricks is a sponsor at Black Hat USA 2026 this week. Find us at Booths #5106...
Open sourceGitHub GrowthRelevance 90
openai/codex-security: +452 GitHub stars
OpenAI's Codex Security CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities. npm: https://www.npmjs.com/package/@openai/codex-security
Open sourceTechCrunchRelevance 90
OpenAI reportedly finds evidence that more of its agents ran amok
OpenAI has reportedly found evidence of additional agent misbehavior as it looks into the incident that occurred with Hugging Face.
Open sourcehnRelevance 90
Investigating three real-world incidents in our cybersecurity evaluations
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems o...
Open sourceRedditRelevance 90
I can pull another user's data out of half the AI-built SaaS apps I test
Give me the API endpoint behind a hidden button and I'll usually get another user's data out of it. That's the most common bug I find testing SaaS apps, and lately most of the ones I see were built solo or with AI coding tools. They break the same handful of w...
Open sourceRedditRelevance 90
What kind of security flaws actually matter in SaaS apps
I keep noticing that the issues that really matter in SaaS are not the obvious ones. It’s usually things like authentication behaving differently in edge cases, APIs trusting things they shouldn’t, or data from one user leaking into another account in unexpect...
Open sourceGitHub GrowthRelevance 90
openai/codex-security: +1589 GitHub stars
OpenAI's Codex Security CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities. npm: https://www.npmjs.com/package/@openai/codex-security
Open sourceGitHub GrowthRelevance 90
cloudflare/security-audit-skill: +11 GitHub stars
A coding-agent skill for multi-phase security audits with independently verified, machine-readable findings
Open sourcebluesky globalRelevance 90
Okta buys AI security startup Permiso; source says for about $200M | TechCrunch
The deal gives Okta identity threat detection capabilities as enterprises seek to secure AI agents and other non-human identities across cloud environments. The deal gives Okta identity threat detection capabilities as enterprises seek to secure AI agents and ...
Open sourceTechCrunchRelevance 90
Okta buys AI security startup Permiso — source says for about $200M
The deal gives Okta identity threat detection capabilities as enterprises seek to secure AI agents and other non-human identities across cloud environments.
Open sourceTechCrunchRelevance 90
Anthropic says its own AI models breached three companies during security tests
After OpenAI's models broke into Hugging Face, Anthropic checked its own history and found three similar incidents
Open sourceGitHub GrowthRelevance 90
openai/codex-security: +1701 GitHub stars
OpenAI's Codex Security CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities. npm: https://www.npmjs.com/package/@openai/codex-security
Open sourceHF Daily PapersRelevance 90
GPT-Red: Automated Red Teaming via Self-Play at Scale
We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversa...
Open sourceHF Daily PapersRelevance 90
SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity ...
Open sourceHF Daily PapersRelevance 90
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their...
Open sourceTechCrunchRelevance 90
The Hugging Face break-in explained
Another way to think about the whole thing is to picture a bear at a campsite. (Really, we are going there.)
Open sourceAWSRelevance 90
Authenticate with Private Key JWT using Amazon Bedrock AgentCore Identity
This post explains how Private Key JWT client authentication works in AgentCore Identity and reviews the supported grant flows. We then walk through creating an AWS KMS signing key, registering its public key with your identity provider, configuring a credenti...
Open sourcehnRelevance 90
Handbook.md shows that long policy documents do not reliably govern agents
Abstract page for arXiv paper 2607.25398: HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following --> Computer Science > Artificial Intelligence arXiv:2607.25398 (cs) [Submitted on 28 Jul 2026] Title: HANDBOOK.md: A Benchmark for Long-Context A...
Open sourcebluesky globalRelevance 90
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
Hugging Face just published a highly detailed technical account of OpenAI's accidental cyberattack on their systems - it's wild how sophisticated this was: huggingface.co/blog/agent-i... Wrote up some of my own notes here: simonwillison.net/2026/Jul/28/... We’...
Open sourceTechCrunchRelevance 90
Cyera agrees to acquire Oasis Security for $1B to safeguard proliferating AI agents
The deal is Cyera's third acquisition this year.
Open sourceTechCrunchRelevance 90
Bot-detection startup Spur nabs $200M from Insight
Spur Intelligence has raised a $200 million round from Insight Partners for its tech that can identify legit human traffic from bots.
Open sourceAWSRelevance 90
How AgentCore Gateway supports the MCP 2026-07-28 spec
The Model Context Protocol (MCP) published its 2026-07-28 specification, the largest revision since launch: MCP is now stateless, with a governed extensions system and hardened authorization. Learn what changed and how to enable the new version on Amazon Bedro...
Open sourceRedditRelevance 82
Anthropic went back through 141,006 of its own security eval runs and admitted its models broke out of the test and into three real companies
So Anthropic put out this incident report on July 30. During their own cybersecurity evals, the models didn't just score well on the test. In three separate cases they actually got out. Into real companies. Ones that were never supposed to be part of the exercise at all. They went back through 141,006 eval runs. Three
Open source