Executive Summary
As we continue developing and benchmarking Rakshak at Rudraksh AGI, our goal is simple: build an autonomous security engine that can test web applications, reason over complex technical vulnerabilities like a human security engineer, and generate clear, action-oriented security reports without manual hand-holding.
While Rakshak is currently undergoing active engineering and benchmark testing prior to its public launch, we believe in publishing evidence-based evaluations of our progress. On September 9, 2026, we published our v1 benchmark evaluation. In that test against OWASP Juice Shop, Rakshak scored 8.05 out of 10. It demonstrated excellent reasoning (10/10 ReAct score) and successfully confirmed SQL injection vulnerabilities. However, that test also exposed important technical gaps in our engine:
- No XSS Detection: Our pipeline relied on traditional URL-parameter scanners, which could not execute JavaScript inside modern Single-Page Applications (SPAs). As a result, 0 XSS vulnerabilities were found.
- Untested Vulnerability Classes: Command injection testing via
commixwas never dispatched due to process termination. - Tool Pipeline Drops: 43% of planned tool steps either timed out or were skipped because of connector limits.
Over the next ten days, we rebuilt key parts of our tool pipeline, binary resolution, and browser orchestration. On September 19, 2026, we ran our official Rakshak V2 benchmark against the exact same test application. This report presents the complete technical findings—explaining how Rakshak V2 improved as an engine, and detailing the Before-Fix and After-Fix security assessment of the target application.
Target Environment & Scope
To evaluate autonomous security testing without introducing synthetic bias, our evaluation harness target was OWASP Juice Shop running inside an isolated Docker container mapped to http://localhost:3000 (Session ID: f8a19164-9edd-45ea-a6da-4808eb6ba77e). OWASP Juice Shop represents a production-grade modern web application architecture:
- Frontend: Angular Single-Page Application (SPA) with client-side JavaScript routing.
- Backend: Node.js, Express REST APIs, and JSON web tokens (JWT).
- Database: SQLite 3.
- Network Scope: 1 containerized target, exposed ports 3000/tcp and 8888/tcp.
How Rakshak V2 Works: Multi-Agent ReAct Engine
Rakshak V2 operates through a continuous five-step autonomous loop powered by a ReAct (Reason + Act) supervisor engine:
- Target Reconnaissance: Dispatches
nmap,httpx,wafw00f, andapi_discoverto map ports, HTTP headers, technologies, and endpoints. - Targeted Fuzzing & Security Probes: Automatically selects and dispatches security tools (
sqlmap,commix,xsstrike,ffuf, and customxss_scanner). - ReAct Reasoning & Adaptive Replanning: Analyzes stdout/stderr streams in real time. If port 80 fails or a tool times out, the engine adapts its plan dynamically without aborting.
- Proof-of-Concept Evidence Verification: Confirms candidate vulnerabilities with deterministic payload execution before filing alerts.
- Structured Report Generation: Compiles executive summaries, OWASP Top 10 mappings, CWE tags, and SARIF/JSON logs.
Comparison A — Previous Benchmark (v1) vs Current Benchmark (v2)
This comparison measures the progress of Rakshak's engine capabilities between our September 9 test (v1) and September 19 test (v2).
| Evaluation Metric | Previous Benchmark (Sep 09) | Current Benchmark V2 (Sep 19) | Engine Delta |
|---|---|---|---|
| Composite Engine Score | 8.05 / 10 | 8.75 / 10 | +0.70 (+8.7%) |
| Total Findings Discovered | 7 | 16 | +9 (+128%) |
| Critical Severity Findings | 0 | 7 | +7 (NEW) |
| High Severity Findings | 2 | 4 | +2 (+100%) |
| Medium Severity Findings | 2 | 2 | Confirmed (Same) |
| Low Severity Findings | 1 | 1 | Confirmed (Same) |
| XSS Vulnerability Detection | 0 (Not run) | 7 (3 DOM, 3 Header, 1 Parameter) | +7 (NEW) |
| SQL Injection Detection | 3 | 3 | Confirmed & Enhanced |
| Command Injection Detection | 0 (Not run) | 1 | +1 (NEW) |
| Tools Executed | 10 | 16 | +6 (+60%) |
| Plan Steps Completed | 8–9 | 25 | +16 (+178%) |
| ReAct Reasoning Rounds | 7 | 20 | +13 (+185%) |
| Reconnaissance Category Score | 8 / 10 (1.20) | 9 / 10 (1.35) | +0.15 weighted |
| Vulnerability Detection Score | 7 / 10 (1.05) | 9 / 10 (1.35) | +0.30 weighted |
| Report Generation Score | 4 / 10 (0.40) | 8 / 10 (0.80) | +0.40 weighted |
| Coverage Scope Score | 5 / 10 (0.25) | 8 / 10 (0.40) | +0.15 weighted |
Detailed Analysis of Engine Changes (v1 → v2)
1. Custom Playwright Headless Browser XSS Scanner (xss_scanner)
- Before: In v1,
xsstrikefailed against Angular SPAs because it only scanned static HTTP response bodies, missing client-side DOM execution. - Now: Rakshak V2 introduces a custom Playwright Chromium tool (
xss_scanner). It renders the live DOM, injects payloads into input fields and hash routes, and detects browser popups and script execution. - Technical Impact: Discovered 7 XSS findings (3 DOM XSS, 3 HTTP Header XSS, 1 Parameter XSS) in 81 seconds of runtime.
2. Integrated Command Injection Fuzzing (commix)
- Before: Command injection was planned in v1 but never executed.
- Now: Dispatched
commixagainst API endpoints. - Technical Impact: Confirmed 1 critical OS command injection vulnerability.
3. Tool Configuration & Timeout Hardening (sqlmap & nikto)
- Before:
sqlmaptimed out at 60s, andniktofailed on closed port 80. - Now:
sqlmaptimeout expanded to 180s with--time-sec=10, whileniktoport mapping was corrected to port 3000. - Technical Impact: Confirmed boolean-based blind, error-based (SQLite >= 3.9 JSON path), and time-based blind SQLi.
4. Catalog Execution Tracking & Report Pipeline Fix
- Before: In v1,
run_tool()did not notify the execution tracker, causing an LLM timeout and skipping final report generation (scoring 4/10). - Now: Integrated state tracking into
run_tool(): plan → queue → invoke → complete → capture output. - Technical Impact: Reporting completed cleanly with full CWE/OWASP mappings (boosting report score from 4/10 to 8/10).
Comparison B — Before-Fix vs After-Fix Security Assessment
This comparison details the security posture of the OWASP Juice Shop target application before and after security remediation.
Scoring Methodology & Meaning of Numbers
- Infrastructure Base Score (Max 100): OWASP Juice Shop scored 85 / 100 (20 pts Attack Surface, 5 pts Configuration due to 5 missing security headers, 20 pts Exposure, 20 pts Encryption, 20 pts Patch Status).
- Vulnerability Penalties:
- 7 Critical Findings × -15 pts = -105 pts
- 4 High Findings × -8 pts = -32 pts
- 2 Medium Findings × -3 pts = -6 pts
- 1 Low Finding × -1 pt = -1 pt
- Total Penalties: -144 pts
- Final Before-Fix Security Score: 0 / 100 (85 - 144 = -59, floored at 0). Expected for an intentionally vulnerable target.
Initial Findings Log & Summary (16 Findings Total)
[2026-09-19T06:47:07] xss_scanner started — target: http://localhost:3000
[2026-09-19T06:47:12] Found 6 injectable URL patterns inside Angular DOM
[2026-09-19T06:47:25] [VULN] DOM XSS: http://localhost:3000/#/search?q=<script>alert('xss')</script>
[2026-09-19T06:47:31] [VULN] DOM XSS: http://localhost:3000/#/search?q=<img src=x onerror=alert('xss')>
[2026-09-19T06:47:38] [VULN] DOM XSS: http://localhost:3000/#/search?q=<iframe src="javascript:alert('xss')">
[2026-09-19T06:47:45] [VULN] Header XSS: User-Agent header injection confirmed
[2026-09-19T06:47:52] [VULN] Header XSS: Referer header injection confirmed
[2026-09-19T06:47:58] [VULN] Header XSS: X-Forwarded-For header injection confirmed
[2026-09-19T06:48:28] xss_scanner completed — 7 findings (81s runtime)
[2026-09-19T06:48:30] sqlmap started — target: GET /rest/products/search?q=
[2026-09-19T06:48:35] [VULN] SQL Injection: SQLite >= 3.9 error-based JSON path confirmed
[2026-09-19T06:49:00] commix completed — [Confirmed] OS Command Injection on target API
Findings List Before Fixes:
- 7 CRITICAL — Cross-Site Scripting (XSS): 3 DOM XSS on
/#/search?q=, 3 HTTP Header XSS (User-Agent, Referer, X-Forwarded-For), 1 parameter XSS on/rest/track-order/{id}(CWE-79, OWASP A03:2021). - 3 CRITICAL — SQL Injection: Boolean-based blind, error-based, time-based blind on
GET /rest/products/search?q=targeting SQLite 3 (CWE-89, OWASP A03:2021). - 1 CRITICAL — Command Injection: OS command execution confirmed via
commix(CWE-78, OWASP A03:2021). - 2 MEDIUM — Security Misconfigurations: 5 missing security headers (CSP, HSTS, Referrer-Policy, Permissions-Policy, X-XSS-Protection) plus 21 Nikto web scan issues.
- 1 LOW — Open Port Exposure: 3000/tcp and 8888/tcp exposed directly.
Remediation & After-Fix Re-Assessment Results
- Remediation: Added HTTP security headers middleware (CSP, HSTS, X-XSS-Protection), resolved binary catalog paths, and hardened execution timeouts.
- After-Fix Re-Assessment: Post-remediation scan verified the presence of all required security headers, eliminating configuration penalties and returning +15 points to the Configuration score category.
Before-Fix vs After-Fix Visual Workflow
1. Initial Assessment
16 Total Findings
7 Critical (XSS, SQLi, Cmdi)
Security Score: 0 / 100
2. Applied Remediation
CSP & HSTS Middleware
Input Fuzzing Hardening
Tool Timeout Fixes
3. Post-Fix Assessment
Headers Confirmed
Misconfigurations Resolved
Verified Post-Fix Baseline
Limitations & Engineering Roadmap
- Controlled Environment: Benchmark was conducted on OWASP Juice Shop in Docker. Enterprise environments with WAFs may require additional evasion profiles.
- Tool Dependencies: Detection relies on underlying tools in the swarm pipeline.
- Timeout:
nucleitimed out during this test run. - Roadmap: Future iterations will focus on dynamic REST/GraphQL API schema reverse-engineering and federated multi-cloud agent swarms.
Conclusion
Building and benchmarking Rakshak V2 has demonstrated that real browser DOM orchestration (via Playwright) combined with resilient ReAct reasoning bridges the gap between manual human pentesting and automated scanners.
Official Enterprise Benchmark Attribution
Organization: Rudraksh AGI | Product: Rakshak V2 | Official Portal: rakshak.rudrakshai.in
Lead AI Architect: Aditya Kumar Mishra | GSTIN: 24KNKPM5455A1ZX | Contact: support@rudrakshai.in
Previous Benchmark Reference: v1 Benchmark Report (Sep 09, 2026)