What we test
We test AI-enabled systems end to end, including agents, tool-calling, and MCP integrations, when the engagement is scoped and staffed for it. We do not sell autonomous AI-run testing as a product, and no AI performs the assessment on our end. A Red Sentry pentester runs every test on this page.
Prompt injection and indirect prompt injection (instructions smuggled in through user input, documents, web pages, or upstream data)
Sensitive data leakage and system prompt exposure
Model abuse scenarios (getting the system to do something it should refuse)
RAG and source exposure (retrieving records, files, or context the user should never see)
Tenant isolation and authorization gaps in AI-driven workflows
Unsafe tool or function calling
Excessive tool permissions
MCP and connector risk, a growing attack surface, covered in detail below
Insecure API integrations and insecure output handling (unvalidated model output flowing into your app, browser, or downstream systems)
Business logic abuse (chaining an AI weakness with an app, API, or cloud flaw into real impact)
AGENTIC RISK
MCP and tool-calling: an emerging attack surface
MCP (Model Context Protocol) and similar tool-calling architectures connect an LLM to external tools, data sources, files, APIs, commands, and business workflows. That connection is exactly why it matters. It turns model output into something that can act, not just something that talks.
When MCP servers, connectors, or tool-calling are part of your environment, we evaluate whether those integrations are least-privileged, properly authorized, auditable, and resistant to prompt or tool manipulation, including:
Tool-description injection and tool-output or resource-based prompt injection
Malicious or overtrusted MCP servers, and cross-server trust issues
Excessive tool permissions and unsafe tool or function execution
Credential or token exposure
Conversation or context leakage
Authorization gaps between the user, the model, MCP servers, tools, and downstream systems
Unsafe access to files, APIs, commands, or business workflows
Logging, monitoring, approval-flow, and auditability gaps
This work is scoped and staffed to match your architecture — we confirm tester depth before the engagement is scoped.
Why it matters
AI features can expose sensitive data, trigger actions no one approved, bypass controls you thought were enforced, or open a fresh path into the apps, APIs, cloud services, and internal workflows behind them.
The blast radius is not the chatbot. It is everything the chatbot is connected to.
Three ways this shows up in real engagements:
A support assistant that can be steered into returning another customer's records
An agent with a connected tool that can be pushed into an action the user was never authorized to take
Model output that gets trusted downstream and becomes a classic injection into your own app
HOW WE WORK
How we scope it
Scope depends on your system, not a fixed package. Before testing, we work with you to map:
Tenant model (single vs multi-tenant)
How the engagement runs
Our methodology references the OWASP Top 10 for LLM Applications, MITRE ATLAS, and NIST AI RMF. AI/LLM Security Testing can also support SOC 2, ISO 27001, HIPAA, GDPR, and other compliance evidence needs when AI systems are part of the assessed environment.
Frameworks are a reason to test, not a test on their own. We perform the technical testing and give you evidence you can hand to your auditor. We do not certify or attest compliance, and no single test makes you compliant by itself.
Related Services
Web Application Pentesting (WAPT)
For the app your AI feature lives inside.
API Penetration Testing
For the endpoints your model and integrations expose.
Cloud Configuration Review
For the cloud environment hosting the model and data.
Source Code Review
For the AI-enabled and AI-generated code behind the feature.
Threat Modeling
To map attack paths before you build or before you test.
Red Teaming
Full adversary emulation when you want end-to-end, not scoped.
