TechCrunchRelevance 90
Chinese AI model Kimi escaped its cybersecurity testing environment, researchers say
In the Kimi test, the sandbox designed to contain the experiment was not properly configured.
Open sourceOpenAIRelevance 90
Responding to the next frontier of critical cyber capabilities
OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.
Open sourcehnRelevance 90
Mistral's Shieldstral: 3B open-weights model for multimodal moderation
Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size. Solutions Introducing Shieldstral. August 4, 2026 By Mistral Back to Blog 5 min read Share this post Copy url to clipboard Copied Thinking Summary ...
Open sourceTechCrunchRelevance 90
Nvidia doesn’t mess around: A week after open AI industry group formed, it’s already showing progress
The week-old Open Secure AI Alliance, spearheaded by Nvidia and grown to over 120 companies, already has proposals out for defending against AI agents.
Open sourceOpenAIRelevance 90
Third-party cyber evaluations involving OpenAI models
OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.
Open sourcehnRelevance 90
Fast Remediation Is the New Trust Model (JFrog and OpenAI Zero-Day Findings)
Discover how AI models expose zero-day vulnerabilities and why rapid remediation is essential for modern software supply chain security. Fast Remediation Is the New Trust Model: JFrog and OpenAI Collaboration on Zero-Day Security Findings In the Era of AI-Disc...
Open sourceindiehackersRelevance 90
We've reproduced 30+ real AI runtime failures over the past week. Here's the pattern we keep seeing.
Over the past few weeks, we've reproduced 30+ real AI runtime failures from GitHub issues instead of just reading about them. Most weren't model failures - they were runtime contract mismatches between providers, tools, and application code. That led us to bui...
Open sourceGitHub GrowthRelevance 90
cloudflare/security-audit-skill: +10 GitHub stars
A coding-agent skill for multi-phase security audits with independently verified, machine-readable findings
Open sourceTechCrunchRelevance 90
Microsoft launches its first cybersecurity model, plus a new agentic cybersecurity system
Microsoft bolstered its AI cybersecurity offerings this week with the launch of its first AI security model and a new security platform.
Open sourcex manual globalRelevance 90
Today, we are announcing a series of updates that give customers frontier-grade security at half the cost.
Today, we are announcing a series of updates that give customers frontier-grade security at half the cost. MAI-Cyber-1-Flash is our first cybersecurity model, built ground up to find the most challenging vulnerabilities in complex code bases. When combined wit...
Open sourceRedditRelevance 90
OxDeAI: I built a deterministic pre-execution authorization boundary for AI agents (fail-closed, signed artifacts, adapters for LangGraph/CrewAI/AutoGen, etc...), looking for feedback.
Hey everyone. I'm the author of OxDeAI, an open-source protocol (Apache 2.0). Posting it here because I want critical feedback from people building real agents, not applause. The problem I keep hitting: as agents move from generating text to doing things (API ...
Open sourceRedditRelevance 90
'This one was different from anything we had handled before': Hugging Face confirms it was hit by cyberattack powered by an AI agent
submitted by /u/EchoOfOppenheimer to r/OpenAI [link] [comments]
Open sourceGitHub GrowthRelevance 90
elder-plinius/T3MP3ST: +38 GitHub stars
autonomous red teaming platform; multi-agent offensive-security meta-harness
Open sourceGitHub GrowthRelevance 90
cloudflare/security-audit-skill: +12 GitHub stars
A coding-agent skill for multi-phase security audits with independently verified, machine-readable findings
Open sourceHF Daily PapersRelevance 90
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operation...
Open sourceHF Daily PapersRelevance 90
Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
Self-hosted AI agents read and write their own memory and configuration files to function. An agent may get compromised via corruption of its own state -- a compromise realized via legitimate OS system call invocation. We refer to this class of threats as self...
Open sourcehnRelevance 90
Exploit brokers pay $500k for WordPress RCEs. I found one with GPT5.6 and $25
Stay current: Get research alerts for newly disclosed vulnerabilities and exposures If you're running WordPress and want to check if your instance is vulnerable, you can use our tool we've hosted here: https://wp2shell.com/. We held off on publishing this issu...
Open sourceRedditRelevance 90
Built a tool that scans your website for security problems and explains the fixes in plain English, no security background needed
I built ONUS because I kept seeing the same problem: security scanning tools exist, but they're built for people who already know what half the acronyms in the report mean. If you're not a security person, the output is basically unreadable, even when it's tel...
Open sourceGitHub GrowthRelevance 90
cloudflare/security-audit-skill: +18 GitHub stars
A coding-agent skill for multi-phase security audits with independently verified, machine-readable findings
Open sourceGitHub GrowthRelevance 90
larlarua/AutoCVE: +14 GitHub stars
Agent-driven automated CVE discovery platform for source code auditing, vulnerability verification, and report generation.
Open sourceTechCrunchRelevance 90
Hugging Face confirms breach affected internal datasets and credentials, urges users to take action
Hugging Face is urging users to rotate any access tokens stored on the platform and review account activity.
Open sourcehnRelevance 90
Deepsec
Deepsec is a security harness for finding vulnerabilities in your codebase powered by coding agents - vercel-labs/deepsec deepsec deepsec an agent-powered vulnerability scanner that you can run in your own infrastructure, optimized to perform on-demand review ...
Open sourceGitHub GrowthRelevance 90
elder-plinius/T3MP3ST: +76 GitHub stars
autonomous red teaming platform; multi-agent offensive-security meta-harness
Open sourceGitHub GrowthRelevance 90
cloudflare/security-audit-skill: +19 GitHub stars
A coding-agent skill for multi-phase security audits with independently verified, machine-readable findings
Open sourceGitHubRelevance 90
DevenderSEO/leaklatch
Git-aware secret & .env leak guard for TypeScript/Node — sub-second pre-commit scanning with the lowest false-positive rate.
Open sourceHF Daily PapersRelevance 90
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perform best in real-world targets. Existing evaluation protocols assess and optimize for predefined goals such as...
Open sourceHugging FaceRelevance 90
RavichandranJ/Dolphin3-Cyber-8B-GGUF
Fine-tuned for Offensive Security • Defensive Security • Vulnerability Research • Exploit Development
Open sourceTechCrunchRelevance 90
Microsoft patches record number of security vulnerabilities, citing its use of AI
Microsoft's monthly release of security fixes, dubbed Patch Tuesday, resolved a record 570 security vulnerabilities across the company's product line, thanks to discoveries with AI.
Open source36KrRelevance 90
Anthropic揭秘AI四大失控行为:泄密、删账、改分,还差点骗过人类
给足了AI权限,它会不会使坏? Anthropic真的把这个问题,做成了一场实验。 他们把全行业最强的十几个AI模型,一个个扔进模拟的公司和实验室。给代码权限,给财务权限,给评估权限,然后看会发生什么。 结果,四种AI「使坏」模式浮出了水面: Gemini 3.1 Pro暗改训练流程; GPT-5.5帮创始人瞒下投资人的钱; Claude系模型给同行的答卷偷偷改分; Opus 4.5走投无路,教一个员工替自己往外捅料。 7月13日,Anthropic对齐科学团队(Alignment Science)公开了这个实验报...
Open sourcex manual globalRelevance 75
Kimi K3 and Sol show a cost-quality split in cybersecurity benchmarks
Based on internal evals: Kimi K3 is top-tier at cybersecurity. There is chatter on X that Moonshot benchmark-overfit. These are stealth evals. Model has raw IQ. Sol is a leap ahead in cyber capability at a significantly higher cost, but quite remarkable still.
Open source