AI / LLM Security Testing Checklist
A working checklist for testing an LLM-backed application end to end. It’s the AI Hacking 101 notes squeezed into “did I actually do this?” form — tick-boxes grouped by phase, from scoping through to writing it up.
Only run this against a target you have written authorization for, inside an agreed scope. Several sections (agentic tools, ticket close/escalate, load testing) change state or cost money — get explicit sign-off before you touch them.
- LLMmap — minimal-query model fingerprinting (“nmap for LLMs”).
- P4RS3LT0NGV3 — encoding/obfuscation + prompt-mutation workbench.
- Burp Suite — inspect the chat API and its response metadata.
- Small scripts for determinism and rate-limit probing.
0. Scope & rules of engagement
- Written authorization / signed RoE in hand
- In-scope endpoints, accounts, and data confirmed
- Out-of-scope systems and actions written down
- Rate-limit / load ceiling agreed (avoid DoS and “denial of wallet”)
- Point of contact + a kill-switch for anything touching real data
1. Threat model the app first
- Threat actors listed — malicious user, curious insider, competitor, automated bots, criminal actors
- Assets listed — model parameters, training data, private RAG data, user PII, agent tools, API keys / secrets
- Attack surface mapped — input channels, retrieval/RAG layer, tools & plugins, supply chain, logging, admin/debug endpoints, training data
- Risks prioritized — PII/sensitive-data leakage, model/data theft, jailbreak, internal access
2. Recon & fingerprinting
Model & behavior
- Identify model family/version — identity probes + LLMmap
- Estimate determinism / temperature — send one identical prompt N times, count unique replies
-
Estimate context window and where input truncates (long
A×N prompt) - Probe tokenizer / unicode handling (literal separators, mixed-script strings)
-
Detect tools / agentic capabilities; probe a hypothetical
web_fetchfor SSRF surface
Retrieval & prompt
- Confirm whether RAG is used — ask it to cite sources; watch response metadata and latency
- Plant a canary doc in a test index, then ask for it
- Try to surface the system prompt / hidden instructions
- Probe input handling — special chars, unicode, whether Base64/hex is decoded
- Fingerprint moderation — “is phrase X allowed?”, then paraphrase / obfuscate the same idea
Infrastructure
-
Inspect the chat API in Burp for metadata (
rag_used,tickets_used,response_time_ms, …) -
Find rate limits empirically (watch for
429/ “too many requests”) - Look for admin/debug endpoints, trace IDs, or debug flags surfaced to the user
- Multi-turn memory probe — ask it to remember a note, recall it a turn later
The system prompt is not a security control. Treat anything you extract from it (URLs, emails, key formats, org IDs) as a lead to chase, not as the finding itself.
3. Prompt injection
Direct injection
- Ignore-previous-instructions / system-prompt extraction
- “Debug mode” / “maintenance mode” prints instructions
- Role switching (e.g. “you are now a configuration editor”)
- Continuation and verbatim-repetition tricks
-
Tag spoofing —
<SYSTEM>…</SYSTEM>,===END SYSTEM PROMPT===,[INST] - Developer impersonation / authoritative “the user is admin” command
- Translation obfuscation (pig latin, other languages)
Indirect injection
- Plant a payload in app-controlled content the model later reads — ticket titles/bodies, uploaded docs, emails
- Ask the bot about that item so it retrieves and summarizes it
- Confirm whether retrieved content is treated as authoritative over the system prompt
Indirect injection rides in as “trusted” business data, so it isn’t scrutinized like something typed into the chat box — it’s usually far more effective than a direct prompt, and it’s the path a real attacker takes.
Multi-turn injection
- Build a persona over several benign turns, then escalate to sensitive asks
- Reference earlier turns to legitimize the request
- Check whether conversation memory overrides earlier refusals
Obfuscation & encoding
- Base64 / Base32 / hex / URL / unicode-escape payloads
- ROT13 / Caesar / Atbash, Morse and emoji encodings
- Leetspeak, custom-symbol, reversed text, case-flipping, whitespace steganography
- Record which encodings the pipeline decodes vs. treats as literal text
4. Jailbreaks
- Persona jailbreaks — DAN, “unshackled AI”, “evil assistant”
- Mode tricks — Opposite Mode, Chaos Mode
- “Rewrite your guidelines” / helpful-assistant contradiction framing
- Threats & coercion (“I’ll overload your tokens”)
- Log what’s refused vs. what leaks brand-damaging or false output
5. Harmful & off-topic output
- Direct harmful-phrase repetition (profanity, hostile phrases)
- The same phrase wrapped in a jailbreak (Opposite / Chaos mode)
- Non-English / phonetic spelling to slip the filter
- Off-topic requests (recipes, travel, code) — does it stay on-brand or comply then redirect?
6. RAG / retrieval abuse
- Direct query for internal/dev docs — API, security, and access details
- Name a specific internal document and request it verbatim
- “Fishing” — error codes, config-file names, admin-panel and webhook questions
- “Pray-and-spray” keyword dump (keys, api, tokens, admin, secret, credentials…)
- Canary-token retrieval — does a planted doc steer output? are sources leaked verbatim?
- Compare unauthenticated vs. authenticated retrieval rates
-
Watch for leaked infra — admin URLs, API keys, webhook URLs, DB paths,
.env/config.json
Treat any partial hit as a foothold. Even a sliver of useful info is worth keeping — note what came back and how you asked, then reshape the next prompt to dig down that same path.
7. Access control & excessive agency
-
Conversation IDOR — request another user’s
conversation_idthrough the bot - Ticket IDOR — reference an arbitrary ticket number
- Unauthorized actions — close / escalate a ticket with no auth or ownership
- Confirm ownership and authorization checks on every tool the agent can call
These tests change state. Only run close / escalate / delete against test data you’re explicitly cleared to touch, and confirm nothing persists to real users.
8. Map findings to a framework
-
Map each finding to the OWASP LLM Top 10:
- LLM01 Prompt Injection
- LLM02 Sensitive Information Disclosure
- LLM03 Supply Chain
- LLM04 Data & Model Poisoning
- LLM05 Improper Output Handling
- LLM06 Excessive Agency
- LLM07 System Prompt Leakage
- LLM08 Vector & Embedding Weaknesses
- LLM09 Misinformation
- LLM10 Unbounded Consumption
- Map techniques and tactics to MITRE ATLAS
- Record success rates and reproducible prompts for the report