genai
SECURITY LAB
Beta — hiring assessments are in testing now and go live on 9 September 2026.
For hiring teams

Hire pentesters on proof, not résumés

Send candidates timed, proctored assessments built from real labs. They exploit live targets; you get auto-graded scores and a per-candidate report — evidence for your panel, not a verdict. Coverage runs across the OWASP LLM Top 10, MCP (Model Context Protocol), and agentic security — with web, API, and cloud tasks in the same sitting. A modern AI system gets breached through the layers around it; the assessment should too.

Auto-gradedProctoredAI + web, API & cloud

$199 per candidate, à-la-carte — no subscription. Packs of 10 and 25 bring it to ≈$180 and ≈$148 each.

0
manual grading
100%
live targets
10
OWASP categories
What you can screen for

The whole surface, not just the model

A candidate who can break a prompt but not an API is only half useful — so you can test both in one sitting.

AI & LLM

Prompt injection, extraction, unsafe output handling — the OWASP LLM Top 10.

Agentic & MCP

Tool abuse, over-broad scopes, agent trust boundaries and MCP servers.

Web & API

Access control, injection, business-logic flaws and API authorization.

Cloud

Misconfiguration, identity and privilege paths around the workload.

Assess

A real exploit, not a quiz

Candidates face purpose-built vulnerable applications running live for their session — a distraction-free runner, a countdown timer, and evidence submission. No multiple choice, no take-home guesswork.

genaisecuritylab.com/a/appsec
Task 3 of 641:12
Attack · LLM01

System prompt extraction

Make the live assistant reveal the restricted coupon value, then submit it as evidence.

Paste the extracted secret…
Compare

Auto-graded, comparable scores

The grader replays each candidate's exploit to confirm it actually worked. Everyone sits the same task mix under the same server-side grading, so you can read their scores, targets solved and OWASP coverage side by side.

genaisecuritylab.com/assessments/results
Candidate scores
RA
Rafael A.4 of 6 solved64
MW
Mei W.4 of 6 solved61
JK
Jonas K.3 of 6 solved46
TO
Tomas O.2 of 6 solved31
SD
Sara D.1 of 6 solved14
Decide

A report your whole panel can read

Each candidate gets a shareable report: headline score, a task-by-task breakdown with the evidence they submitted, OWASP coverage, and an activity timeline with proctoring flags. It stops there — no shortlist, no recommendation. Your panel makes the call.

genaisecuritylab.com/assessments/report
64
Rafael Alvarez
4 of 6 targets solved
Direct injectionSolved
Prompt extractionSolved
Tool abuseSolved
MCP scope escalationSolved
PII disclosurePartial
Cloud privilege pathFailed
How a screen runs

Choose labs by category and difficulty — or start from a role template — invite candidates by link, and share the reproduced results with your panel.

Every sitting is time-boxed against fresh per-session targets, with randomized scenario selection and server-side grading, so there is no flag to copy and no candidate works on another candidate's target. Proctoring reports basic signals — tab focus and paste anomalies — honestly, and is not invasive surveillance. Every LLM lab is tagged to an OWASP LLM Top 10 item.

Scores and evidence are one input to a human screening decision, never the decision itself: the platform issues no hire recommendation and no pass/fail on a candidate. All names and scores in the mockups above are sample data.

Screen your next AI security hire on real work

Create your first assessment, or see a sample candidate report on a demo.

$199 per candidate · packs of 10 and 25 · no subscription required.