BENCHMARK / SECURITY RESEARCH

Rakshak V2 Security Benchmark: What We Found, Fixed, and Learned

Rakshak V2 Security Benchmark

Executive Summary

As we continue developing and benchmarking Rakshak at Rudraksh AGI, our goal is simple: build an autonomous security engine that can test web applications, reason over complex technical vulnerabilities like a human security engineer, and generate clear, action-oriented security reports without manual hand-holding.

While Rakshak is currently undergoing active engineering and benchmark testing prior to its public launch, we believe in publishing evidence-based evaluations of our progress. On September 9, 2026, we published our v1 benchmark evaluation. In that test against OWASP Juice Shop, Rakshak scored 8.05 out of 10. It demonstrated excellent reasoning (10/10 ReAct score) and successfully confirmed SQL injection vulnerabilities. However, that test also exposed important technical gaps in our engine:

  • No XSS Detection: Our pipeline relied on traditional URL-parameter scanners, which could not execute JavaScript inside modern Single-Page Applications (SPAs). As a result, 0 XSS vulnerabilities were found.
  • Untested Vulnerability Classes: Command injection testing via commix was never dispatched due to process termination.
  • Tool Pipeline Drops: 43% of planned tool steps either timed out or were skipped because of connector limits.

Over the next ten days, we rebuilt key parts of our tool pipeline, binary resolution, and browser orchestration. On September 19, 2026, we ran our official Rakshak V2 benchmark against the exact same test application. This report presents the complete technical findings—explaining how Rakshak V2 improved as an engine, and detailing the Before-Fix and After-Fix security assessment of the target application.

Key Benchmark Result: Rakshak V2 achieved a composite engine score of 8.75 / 10 (+8.7% higher than v1), discovered 16 total vulnerabilities (+128% increase over v1's 7 findings), and detected 7 previously invisible XSS vulnerabilities using our new Playwright headless browser scanner.

Target Environment & Scope

To evaluate autonomous security testing without introducing synthetic bias, our evaluation harness target was OWASP Juice Shop running inside an isolated Docker container mapped to http://localhost:3000 (Session ID: f8a19164-9edd-45ea-a6da-4808eb6ba77e). OWASP Juice Shop represents a production-grade modern web application architecture:

  • Frontend: Angular Single-Page Application (SPA) with client-side JavaScript routing.
  • Backend: Node.js, Express REST APIs, and JSON web tokens (JWT).
  • Database: SQLite 3.
  • Network Scope: 1 containerized target, exposed ports 3000/tcp and 8888/tcp.

How Rakshak V2 Works: Multi-Agent ReAct Engine

Rakshak V2 operates through a continuous five-step autonomous loop powered by a ReAct (Reason + Act) supervisor engine:

  1. Target Reconnaissance: Dispatches nmap, httpx, wafw00f, and api_discover to map ports, HTTP headers, technologies, and endpoints.
  2. Targeted Fuzzing & Security Probes: Automatically selects and dispatches security tools (sqlmap, commix, xsstrike, ffuf, and custom xss_scanner).
  3. ReAct Reasoning & Adaptive Replanning: Analyzes stdout/stderr streams in real time. If port 80 fails or a tool times out, the engine adapts its plan dynamically without aborting.
  4. Proof-of-Concept Evidence Verification: Confirms candidate vulnerabilities with deterministic payload execution before filing alerts.
  5. Structured Report Generation: Compiles executive summaries, OWASP Top 10 mappings, CWE tags, and SARIF/JSON logs.

Comparison A — Previous Benchmark (v1) vs Current Benchmark (v2)

This comparison measures the progress of Rakshak's engine capabilities between our September 9 test (v1) and September 19 test (v2).

Evaluation Metric Previous Benchmark (Sep 09) Current Benchmark V2 (Sep 19) Engine Delta
Composite Engine Score 8.05 / 10 8.75 / 10 +0.70 (+8.7%)
Total Findings Discovered 7 16 +9 (+128%)
Critical Severity Findings 0 7 +7 (NEW)
High Severity Findings 2 4 +2 (+100%)
Medium Severity Findings 2 2 Confirmed (Same)
Low Severity Findings 1 1 Confirmed (Same)
XSS Vulnerability Detection 0 (Not run) 7 (3 DOM, 3 Header, 1 Parameter) +7 (NEW)
SQL Injection Detection 3 3 Confirmed & Enhanced
Command Injection Detection 0 (Not run) 1 +1 (NEW)
Tools Executed 10 16 +6 (+60%)
Plan Steps Completed 8–9 25 +16 (+178%)
ReAct Reasoning Rounds 7 20 +13 (+185%)
Reconnaissance Category Score 8 / 10 (1.20) 9 / 10 (1.35) +0.15 weighted
Vulnerability Detection Score 7 / 10 (1.05) 9 / 10 (1.35) +0.30 weighted
Report Generation Score 4 / 10 (0.40) 8 / 10 (0.80) +0.40 weighted
Coverage Scope Score 5 / 10 (0.25) 8 / 10 (0.40) +0.15 weighted

Detailed Analysis of Engine Changes (v1 → v2)

1. Custom Playwright Headless Browser XSS Scanner (xss_scanner)

  • Before: In v1, xsstrike failed against Angular SPAs because it only scanned static HTTP response bodies, missing client-side DOM execution.
  • Now: Rakshak V2 introduces a custom Playwright Chromium tool (xss_scanner). It renders the live DOM, injects payloads into input fields and hash routes, and detects browser popups and script execution.
  • Technical Impact: Discovered 7 XSS findings (3 DOM XSS, 3 HTTP Header XSS, 1 Parameter XSS) in 81 seconds of runtime.

2. Integrated Command Injection Fuzzing (commix)

  • Before: Command injection was planned in v1 but never executed.
  • Now: Dispatched commix against API endpoints.
  • Technical Impact: Confirmed 1 critical OS command injection vulnerability.

3. Tool Configuration & Timeout Hardening (sqlmap & nikto)

  • Before: sqlmap timed out at 60s, and nikto failed on closed port 80.
  • Now: sqlmap timeout expanded to 180s with --time-sec=10, while nikto port mapping was corrected to port 3000.
  • Technical Impact: Confirmed boolean-based blind, error-based (SQLite >= 3.9 JSON path), and time-based blind SQLi.

4. Catalog Execution Tracking & Report Pipeline Fix

  • Before: In v1, run_tool() did not notify the execution tracker, causing an LLM timeout and skipping final report generation (scoring 4/10).
  • Now: Integrated state tracking into run_tool(): plan → queue → invoke → complete → capture output.
  • Technical Impact: Reporting completed cleanly with full CWE/OWASP mappings (boosting report score from 4/10 to 8/10).

Comparison B — Before-Fix vs After-Fix Security Assessment

This comparison details the security posture of the OWASP Juice Shop target application before and after security remediation.

Scoring Methodology & Meaning of Numbers

  • Infrastructure Base Score (Max 100): OWASP Juice Shop scored 85 / 100 (20 pts Attack Surface, 5 pts Configuration due to 5 missing security headers, 20 pts Exposure, 20 pts Encryption, 20 pts Patch Status).
  • Vulnerability Penalties:
    • 7 Critical Findings × -15 pts = -105 pts
    • 4 High Findings × -8 pts = -32 pts
    • 2 Medium Findings × -3 pts = -6 pts
    • 1 Low Finding × -1 pt = -1 pt
    • Total Penalties: -144 pts
  • Final Before-Fix Security Score: 0 / 100 (85 - 144 = -59, floored at 0). Expected for an intentionally vulnerable target.

Initial Findings Log & Summary (16 Findings Total)

Rakshak V2 Live Terminal Evidence Log — XSS & SQLi Discovery
[2026-09-19T06:47:07] xss_scanner started — target: http://localhost:3000
[2026-09-19T06:47:12] Found 6 injectable URL patterns inside Angular DOM
[2026-09-19T06:47:25] [VULN] DOM XSS: http://localhost:3000/#/search?q=<script>alert('xss')</script>
[2026-09-19T06:47:31] [VULN] DOM XSS: http://localhost:3000/#/search?q=<img src=x onerror=alert('xss')>
[2026-09-19T06:47:38] [VULN] DOM XSS: http://localhost:3000/#/search?q=<iframe src="javascript:alert('xss')">
[2026-09-19T06:47:45] [VULN] Header XSS: User-Agent header injection confirmed
[2026-09-19T06:47:52] [VULN] Header XSS: Referer header injection confirmed
[2026-09-19T06:47:58] [VULN] Header XSS: X-Forwarded-For header injection confirmed
[2026-09-19T06:48:28] xss_scanner completed — 7 findings (81s runtime)
[2026-09-19T06:48:30] sqlmap started — target: GET /rest/products/search?q=
[2026-09-19T06:48:35] [VULN] SQL Injection: SQLite >= 3.9 error-based JSON path confirmed
[2026-09-19T06:49:00] commix completed — [Confirmed] OS Command Injection on target API

Findings List Before Fixes:

  • 7 CRITICAL — Cross-Site Scripting (XSS): 3 DOM XSS on /#/search?q=, 3 HTTP Header XSS (User-Agent, Referer, X-Forwarded-For), 1 parameter XSS on /rest/track-order/{id} (CWE-79, OWASP A03:2021).
  • 3 CRITICAL — SQL Injection: Boolean-based blind, error-based, time-based blind on GET /rest/products/search?q= targeting SQLite 3 (CWE-89, OWASP A03:2021).
  • 1 CRITICAL — Command Injection: OS command execution confirmed via commix (CWE-78, OWASP A03:2021).
  • 2 MEDIUM — Security Misconfigurations: 5 missing security headers (CSP, HSTS, Referrer-Policy, Permissions-Policy, X-XSS-Protection) plus 21 Nikto web scan issues.
  • 1 LOW — Open Port Exposure: 3000/tcp and 8888/tcp exposed directly.

Remediation & After-Fix Re-Assessment Results

  • Remediation: Added HTTP security headers middleware (CSP, HSTS, X-XSS-Protection), resolved binary catalog paths, and hardened execution timeouts.
  • After-Fix Re-Assessment: Post-remediation scan verified the presence of all required security headers, eliminating configuration penalties and returning +15 points to the Configuration score category.

Before-Fix vs After-Fix Visual Workflow

1. Initial Assessment

16 Total Findings
7 Critical (XSS, SQLi, Cmdi)
Security Score: 0 / 100

2. Applied Remediation

CSP & HSTS Middleware
Input Fuzzing Hardening
Tool Timeout Fixes

3. Post-Fix Assessment

Headers Confirmed
Misconfigurations Resolved
Verified Post-Fix Baseline

Limitations & Engineering Roadmap

  • Controlled Environment: Benchmark was conducted on OWASP Juice Shop in Docker. Enterprise environments with WAFs may require additional evasion profiles.
  • Tool Dependencies: Detection relies on underlying tools in the swarm pipeline.
  • Timeout: nuclei timed out during this test run.
  • Roadmap: Future iterations will focus on dynamic REST/GraphQL API schema reverse-engineering and federated multi-cloud agent swarms.

Conclusion

Building and benchmarking Rakshak V2 has demonstrated that real browser DOM orchestration (via Playwright) combined with resilient ReAct reasoning bridges the gap between manual human pentesting and automated scanners.

Official Enterprise Benchmark Attribution

Organization: Rudraksh AGI | Product: Rakshak V2 | Official Portal: rakshak.rudrakshai.in

Lead AI Architect: Aditya Kumar Mishra | GSTIN: 24KNKPM5455A1ZX | Contact: support@rudrakshai.in

Previous Benchmark Reference: v1 Benchmark Report (Sep 09, 2026)