GitHub GrowthRelevance 90
NanoNets/Graft: +207 GitHub stars
Turbocharge Claude Code, Cursor, Codex, Gemini & every coding agent: faster, cheaper, with contextual understanding specific to your codebase.
Open sourceGitHub GrowthRelevance 90
openai/codex-security: +35 GitHub stars
OpenAI's Codex Security CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities. npm: https://www.npmjs.com/package/@openai/codex-security
Open sourceHF Daily PapersRelevance 90
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CLR), a training-free framework that reallocates test-time compute from additional solution sampling to targeted verificati...
Open sourceHF Daily PapersRelevance 90
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reve...
Open sourcehnRelevance 90
Why Software Factories Fail (or: harness engineering is not enough)
Why Software Factories Fail (or: harness engineering is not enough)
Open sourceRedditRelevance 90
The primitive for software factory should be a roadmap, not kanban or chat? (I will not promote)
Every agentic factory I see right now converges on the same shapes: a kanban board, a workflow graph, or a Slack-like chat. They all tell you status but none of them show you how the work is actually going. Is the agent stuck in research? Did it backtrack? Is ...
Open sourceGitHub GrowthRelevance 90
TestSprite/testsprite-cli: +32 GitHub stars
Official TestSprite CLI — AI-powered automated testing from your terminal
Open sourceHF Daily PapersRelevance 90
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and run...
Open sourceHF Daily PapersRelevance 90
ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders
The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified instructions, agents are increasingly expected to transform incomplete product intent into working software...
Open sourceGitHubRelevance 90
boringmarketer/kimi-first
Claude Code skill: Kimi types, Claude thinks & verifies — route implementation work-orders to the Kimi Code CLI (k3). A port of @steipete's codex-first (github.com/steipete/agent-scripts). By @boringmarketer · boringmarketing.com
Open sourceGitHubRelevance 90
tmustier/pi-queue-steer
Cursor-style visible follow-up queue with inline editing for Pi
Open sourceHF Daily PapersRelevance 90
Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code
Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their strictness makes generation more difficult. Off-the-shelf compilers can provide useful feedback post-generation, but does not guide intermediate ge...
Open sourceQwenRelevance 90
Qwen-Coder-Qoder: Customizing a Fast-Evolving Frontier Model for Real Software
Learn More about Qoder Explore Qoder for Enterprise Introduction Today, we are pleased to introduce Qwen-Coder-Qoder, a customized model tailored to elevate the end-to-end agentic coding experience on Qoder. Built upon the Qwen-Coder foundation, Qwen-Coder-Qod...
Open source