Open Engineering Radar · Domain Registry · Opportunity Library · LLM Summary

Tencent/AICGSecEval

Commercial score 62 · HOLD

Enterprise DevSecOps teams adopting AI coding assistants (Cursor, Copilot, Claude Code) lack an operational way to continuously measure the security qualit

agentaigcbenchmarkcodesecurityllm

Capability

Provides a repository-level benchmark and evaluation framework for measuring the security quality of AI-generated code, integrating task construction from real GitHub projects and CVE patches, project-context code generation, and hybrid static-plus-dynamic security assessment across C/C++, PHP, Java, Python, and JavaScript.

Target user

AI security researchers, academic groups studying LLM/agent code generation, enterprise security teams benchmarking AI coding assistants, and developers integrating AI tooling into secure SDLCs.

Pain point

Enterprise DevSecOps leaders and AI-adopting engineering organizations cannot operationally measure the security quality of AI-generated code inside their SDLC, because A.S.E is a research-grade benchmark that requires 100GB+ disk, Docker infrastructure, multiple API keys, multi-hour batch runs, and CLI fluency to produce an evaluation — leaving them to either ship AI-generated code unscreened for vulnerabilities or build a parallel benchmarking program internally.

Commercial opportunity

Zero payment signals: customer_count=0, payment_signals=0, negative_results=0, no closed-won deals, no waitlist, no preorders. The $499-2,500/mo SaaS / $2,500-10,000 setup / $10,000-25,000 industry pilot pricing model is logical but entirely hypothetical — anchored to comparable DevSecOps tooling tiers rather than validated buyer conversations.

Best MVP

Managed AICGSecEval-as-a-Service: hosted multi-tenant evaluation API that accepts a repo URL + target LLM/agent, returns SARIF + JSON scorecard (CWE/OWASP hits, false-positive rate, dynamic PoC pass/fail). MVP scope = (1) Docker-in-Docker worker pool running invoke.py with a credential proxy so customers never see raw API keys, (2) GitHub App that triggers evaluation on PR open, (3) hosted leaderboard refresh on each new foundation model release, (4) SOC-2-ready audit log. Ship in 7 days as a single-tenant pilot behind a waitlist.