Skip to content

Rishi Rai

All Writings

Security Threat Modeling for LLM Applications

Map trust boundaries and contain prompt injection, tool misuse, retrieval poisoning, data leakage, and resource abuse.

Security guide

For: Software, security, and AI engineers reviewing the threat model of LLM-enabled applications.

2026-10-02

Security Engineering for LLM Applications

LLM applications combine probabilistic text generation with deterministic systems that store data, call APIs, and take actions. The model may influence a workflow, but it is not a security boundary. It cannot reliably distinguish trusted instructions from hostile text, and it should never be the final authority for access control, transaction approval, or data release.

A secure design assumes that user input, retrieved documents, web content, tool results, and model output can all be adversarial. The surrounding application must enforce policy with conventional controls: identity, authorization, validation, isolation, rate limits, audit trails, and explicit failure handling.

Start With a Threat Model

Map assets, actors, entry points, trust boundaries, and irreversible side effects before choosing mitigations. Assets may include tenant documents, conversation history, system instructions, credentials, model quotas, tool permissions, and business decisions. Actors include authenticated users, compromised accounts, malicious document authors, external content providers, insiders, and automated abuse.

Trace data across the complete request path rather than reviewing only the prompt. A typical workflow crosses several boundaries:

  1. Authenticate the caller and establish tenant, role, and request purpose.
  2. Validate input and classify the operation by risk.
  3. Retrieve context under tenant-scoped authorization and provenance rules.
  4. Construct messages with trusted instructions separated from untrusted content.
  5. Ask the model for a constrained proposal, not unrestricted authority.
  6. Validate output and authorize every tool action in application code.
  7. Encode or sanitize the final response for its destination.
  8. Record security-relevant decisions without logging sensitive payloads by default.

For each transition, ask what an attacker can control, what the component trusts, and what happens if the model returns the worst syntactically valid result. Rank threats by impact and reachability. A support summarizer and an agent that can transfer funds should not share the same acceptance criteria.

Treat Prompt Injection as Untrusted Data

Prompt injection is not solved by adding a sentence that says to ignore malicious instructions. Direct injection arrives from the user; indirect injection hides in retrieved files, emails, tickets, websites, images, or tool responses. Both exploit the model's tendency to follow text that resembles instructions.

Use layered containment. Delimit external content and label its origin, but do not rely on delimiters as enforcement. Minimize the context supplied to the model, remove unnecessary active content, and avoid revealing secrets or privileged instructions that the model does not need. Constrain the model to a narrow output schema. Most importantly, enforce authorization after generation, outside the model.

The safe question is not whether the model can detect every injection. It is whether an injection can cross a deterministic boundary and cause harm.

Injection classifiers and model-based detectors can reduce obvious attacks, but they have false positives and false negatives. Use them as risk signals for blocking, review, or reduced capability, never as the only control protecting a privileged action.

Authorize Tools at Execution Time

Tool definitions exposed to a model should be minimal, typed, and capability-specific. Avoid a generic HTTP tool, raw SQL tool, shell tool, or broad cloud credential. The application should derive identity and tenant from the authenticated session, not from model-supplied arguments. Validate argument schemas, object ownership, allowed state transitions, and spending or volume limits for every call.

proposal = model.generate(allowed_tools=tools_for(session.role))
call = parse_and_validate(proposal.tool_call)

assert call.name in policy.allowed_tools(session.role)
resource = load_resource(call.resource_id)
assert resource.tenant_id == session.tenant_id
assert policy.permits(session.user_id, call.name, resource)

if policy.requires_approval(call.name, resource):
    return request_human_confirmation(call)

return execute_with_scoped_credential(call, idempotency_key=request.id)

Use short-lived credentials scoped to one tool and tenant. Add idempotency keys, timeouts, concurrency limits, and transaction boundaries. High-impact actions should present the exact resolved action to a human for confirmation. Confirmation must not be a vague approval of model prose; it should name the target, parameters, and consequences.

Defend Retrieval Against Poisoning

Retrieval-augmented generation creates a content supply chain. An attacker who can add or rank a document may influence answers without touching application code. Enforce write authorization at ingestion, record source and revision, scan supported file types, and separate public, internal, and restricted indexes. Retrieval filters must apply tenant and access-control predicates before ranking, not after documents have entered the prompt.

Use provenance in both the model context and user-visible citations. Prefer authoritative sources for sensitive questions, limit how much any one source can dominate, and detect suspicious instruction-like passages. Content freshness and deletion must propagate to embeddings and caches. For high-risk decisions, retrieval should provide evidence for deterministic policy or human review rather than directly trigger an action.

Prevent Data Leakage and Enforce Tenant Isolation

Minimize data before sending it to any model provider. Redact secrets and unnecessary personal data, define retention and training settings contractually, and keep provider credentials out of prompts. Do not assume a system message can prevent disclosure once sensitive text is in context. Apply output inspection for known secret formats and policy-sensitive fields, while recognizing that detection cannot recover data already sent to an unauthorized provider.

Tenant isolation must hold in storage, retrieval, caches, logs, queues, evaluation datasets, and observability. Include tenant identity in every key and query, preferably with database-level row policies or physically separate stores for stronger requirements. Never accept a tenant identifier generated by the model. Cache keys must include tenant, user authorization scope, model configuration, and relevant document revisions.

documents = vector_search(
  query=validated_query,
  filter={
    "tenant_id": session.tenant_id,
    "classification": {"$in": session.allowed_classes}
  }
)

assert all(doc.tenant_id == session.tenant_id for doc in documents)

Handle Model Output as Hostile

Model output can contain script, unsafe URLs, malformed markup, formula injection, path traversal strings, or plausible but invalid code. Parse structured output with a strict schema and reject unknown fields. Encode text for the destination context. Use an allowlist sanitizer for rendered markup, parameterized queries for databases, safe APIs for file paths, and isolated sandboxes for code execution.

Do not pass generated content through dynamic evaluation, shell interpolation, or templates with unsafe escaping. When displaying citations or links, resolve them against known retrieved sources rather than trusting URLs invented by the model. If validation fails, fail closed for actions and request a corrected response under a bounded retry policy.

Control Abuse, Cost, and Availability

Rate-limit by user, tenant, IP risk, model, and expensive tool, with limits weighted by tokens or estimated cost rather than request count alone. Bound input size, retrieved context, output tokens, tool iterations, retries, and total workflow duration. Apply backpressure and circuit breakers when providers or downstream tools degrade. Quotas should reserve capacity for critical operations and prevent one tenant from exhausting shared resources.

Repeated failures, broad retrieval queries, unusual tool sequences, and rapid tenant enumeration are useful abuse signals. Avoid silently retrying costly or side-effecting actions. Return explicit failure states and preserve idempotency.

Build Useful Audit Trails

Record authenticated actor, tenant, prompt and model version, retrieved document identifiers, policy decisions, tool proposals, approvals, executions, validation failures, and final disposition. Protect logs from modification and restrict access. Store hashes or references instead of raw sensitive prompts when possible, and define retention periods. Audit data should support incident reconstruction without becoming a second ungoverned corpus of customer content.

Test Controls and Red-Team the Workflow

Security tests should verify enforcement, not whether the model happens to behave. Unit-test authorization functions, tenant filters, output parsers, sanitizers, quota accounting, and audit emission. Integration tests should substitute adversarial model responses and confirm that invalid tool names, cross-tenant identifiers, excessive arguments, repeated calls, and malformed output are rejected.

Maintain a versioned red-team corpus covering direct and indirect injection, encoded instructions, multilingual attacks, poisoned retrieval, secret extraction, role confusion, unsafe rendering, denial of wallet, and chained tool abuse. Include benign lookalikes to measure false positives. Run cases whenever prompts, models, tools, retrieval, or policy changes.

  • Can retrieved text cause a tool call that the user did not request?
  • Can a valid user reference another tenant's object by changing an identifier?
  • Can output escape its rendering context or introduce an unapproved link?
  • Can retries duplicate a side effect or bypass a quota?
  • Do logs reveal secrets, full documents, or credentials?
  • Does revoking access invalidate retrieval results and caches promptly?

Evaluate detection rate, false-positive rate, unauthorized action rate, cross-tenant exposure, and control coverage by risk tier. Treat any confirmed tenant crossover or unauthorized high-impact action as a release blocker rather than averaging it into a broad quality score.

Common Mistakes and Trade-offs

  • Trusting the system prompt: instructions guide behavior but do not enforce policy.
  • Giving one agent broad credentials: compromise then reaches every connected system.
  • Filtering only user input: retrieved content and tool output remain attack paths.
  • Logging everything: observability can create a durable leakage channel.
  • Blocking every anomaly: excessive false positives drive unsafe workarounds.

Isolation, approval, and narrow tools add latency and engineering cost. Aggressive filters can reduce usefulness. Make these trade-offs explicit by risk tier: permit more automation for reversible, low-impact tasks, and require stronger authorization, confirmation, and evidence for consequential actions.

Conclusion

LLM application security depends on containing an unreliable decision-maker inside reliable boundaries. Threat-model the full data path, distrust every text channel, authorize tools deterministically, isolate tenants, validate outputs, constrain resource use, and preserve auditable decisions. Red-team the workflow as a system, then block releases on control failures that can produce real harm.