Web Security
SSRF Risks in AI Agents and URL-Fetching Features
Defensive guidance for securing AI-agent URL fetching against SSRF with destination policy, DNS validation, egress controls, redirects, and response limits.
Server-side URL fetching becomes an SSRF risk when an application retrieves a destination chosen by a user, model, document, tool result, or autonomous plan. A model-selected URL is still untrusted input. The safe design is a URL-fetch gateway that validates the proposal, applies destination and network policy, controls the request, and treats the response as untrusted data.
The safe URL-fetch boundary
Use this conceptual flow:
User or Agent → URL Proposal → Scheme and Syntax Validation → Destination Policy → DNS Resolution → IP and Network Policy → Controlled Egress → HTTP Fetcher → Redirect Revalidation → Response Limits and Type Validation → Application or Model
Each arrow is an enforcement boundary. The model may explain why a URL is useful, but trusted application code decides whether a request is allowed. Secure Tool Calling for LLM Applications describes the broader proposal-and-policy pattern.
Why AI expands the surface
URLs can arrive in prompts, retrieved pages, emails, tool output, generated citations, crawler queues, preview features, and agent plans. They may be syntactically valid while pointing at a loopback service, private network, control-plane endpoint, or an unexpected tenant. Prompt instructions such as “only fetch public pages” are useful guidance but are not a network boundary.
Define the feature first. A fixed partner integration can use a strict host allowlist. An open-web research tool needs a carefully designed proxy and stronger monitoring. A document previewer may need no outbound network at all. Do not grant open-web reach merely because a model can produce URLs.
Restrict schemes and parse canonically
Allow only protocols required by the product, commonly HTTPS and sometimes HTTP for a controlled legacy case. Reject arbitrary schemes and embedded credentials. Parse with the platform’s standards-compliant URL type; do not validate with string prefixes or ad-hoc regular expressions. Normalize the parsed hostname, port, and path before policy comparison, and define how internationalized names and unusual textual forms are handled.
A visible hostname is not the complete destination. Resolve it using a controlled resolver and evaluate the resulting addresses. Treat a URL that changes meaning after parsing, canonicalization, or resolution as a policy failure, not an opportunity to guess intent.
Destination and IP policy
Where possible, allowlist exact hosts or registered partner domains. For open-web features, maintain a deny policy for loopback, link-local, private and reserved ranges, internal service names, and cloud metadata or control-plane endpoints. Keep the policy in a trusted service rather than in model context.
Hostname validation alone is insufficient. DNS can return an address that is not acceptable to the feature, and the answer can change between validation and connection. Use an egress proxy or controlled resolver where feasible, revalidate at connection time, and prevent the fetcher from reaching internal routing domains. Network controls complement, but do not replace, application authorization.
Redirects are new destinations
A permitted first URL can redirect elsewhere. Disable redirects when the feature does not need them. Otherwise cap the number of hops, parse and validate every Location target, reapply scheme and IP policy, and record the chain. Do not let a client-side redirect policy substitute for server enforcement. Cross-protocol or cross-host redirects deserve an explicit decision.
Egress architecture
A defensible gateway is:
- Authenticate the requesting user, workload, tenant, and task.
- Validate URL syntax and the supported scheme.
- Apply destination allowlist and deny policy.
- Resolve through controlled DNS and reject disallowed address ranges.
- Send traffic through a restricted proxy or network segment.
- Strip ambient cookies, authorization headers, cloud credentials, and internal headers.
- Limit redirects, response size, duration, concurrency, and content types.
- Validate the response before it enters application context or a model.
- Record the decision and outcome without logging secrets.
Sandboxing AI Agents and Generated Code can contain process and filesystem effects, but a sandbox with unrestricted network egress is not automatically safe. The gateway and network still need policy.
Response and credential controls
Fetchers are resource consumers. Set connection and total timeouts, maximum bytes, decompression limits, concurrency budgets, and content-type rules. Stream only when the consumer can enforce limits. Reject or quarantine unexpected content rather than passing it directly into a parser or model. A successful HTTP status does not prove that the body is safe or relevant.
Never forward browser cookies, bearer tokens, cloud credentials, service-account headers, or internal tracing headers to an arbitrary destination. If a partner requires authentication, use a dedicated credential reference bound to that host and purpose, and keep custody in the server-side broker described in Secure Credential Handling for AI Agents. Do not place credentials in the URL.
Safe URL fetch gateway matrix
| Stage | Threat | Control | Failure evidence | Verification test |
|---|---|---|---|---|
| Proposal | Model or user chooses unsafe target | Treat as untrusted; bind principal and task | Rejected proposal event | Negative URL cases |
| Parse | Ambiguous or unsupported syntax | Standard parser; scheme and host rules | Parse or policy denial | Canonicalization tests |
| Resolve | Private or changing address | Controlled DNS; range checks; revalidation | Resolved IP and decision | Resolver tests |
| Egress | Internal reachability | Proxy, firewall, segmented network | Network policy denial | Reachability test |
| Redirect | Allowed URL crosses boundary | Disable or validate every hop | Redirect chain | Multi-hop test |
| Fetch | Resource exhaustion | Time, byte, concurrency, type limits | Timeout or size event | Large/slow response tests |
| Response | Active or poisoned content | Quarantine, parse safely, provenance | Content validation result | Malformed-content tests |
| Credentials | Secret disclosure | Strip ambient headers; scoped broker | Redaction/audit event | Header inspection |
Logging and incident response
Record the requesting principal, agent or tool, original URL, normalized destination, resolved target, policy version, redirect chain where appropriate, response classification, and result. Use correlation IDs across the agent, gateway, proxy, and downstream service. Redact authorization headers, cookies, signed URLs, and sensitive response bodies.
An incident review should answer which agent requested the fetch, which tenant and task were involved, what destination was contacted, which policy allowed it, what data returned, and whether any internal service was exposed. Prepare a disable switch for the tool, preserve evidence, and invalidate cached or memory-stored responses if they were contaminated.
Common failure modes
- Allowing any HTTP or HTTPS URL because it looks public.
- Checking a hostname but not its resolved addresses.
- Validating only the first redirect.
- Forwarding the user’s cookies or cloud credentials.
- Relying on a prompt instruction instead of an egress boundary.
- Allowing unbounded response size, decompression, or concurrency.
- Sending fetched HTML directly to a renderer or model without provenance.
- Treating a container or sandbox as permission to reach every network.
- Logging complete URLs that contain tokens or private query data.
Practical checklist
- Define whether the feature needs partner-only or open-web access.
- Accept only required schemes and parse with a standard URL parser.
- Normalize host, port, and address representation before comparison.
- Use exact allowlists where the business case permits them.
- Reject loopback, link-local, private, reserved, internal, and metadata destinations.
- Control DNS and revalidate the connection destination where feasible.
- Disable redirects or validate every hop with the same policy.
- Place fetches behind a proxy or restricted egress segment.
- Strip ambient cookies, authorization, cloud, and internal headers.
- Bound time, bytes, decompression, concurrency, and content types.
- Treat responses as untrusted data with provenance.
- Log decisions with redaction and test denial, redirect, DNS, and recovery paths.
Sources
- OWASP Server-Side Request Forgery Prevention Cheat Sheet
- OWASP Web and API Security guidance
- CISA secure cloud metadata-service guidance
- MDN URL and Fetch API documentation
- Runtime and network-proxy security documentation
The durable rule is simple: let an agent propose a destination, but let a separate, testable gateway decide whether the server may contact it. Application validation, network egress policy, response limits, and credential isolation must reinforce one another.
Separate fetchers by purpose
Do not give every agent the same network capability. A citation fetcher may need public HTTPS and a small response budget. A partner synchronizer may need one registered domain and authenticated requests. A webhook verifier should accept inbound events rather than fetch arbitrary URLs. Separate adapters, credentials, queues, and egress policies so a bug in one feature does not become a universal network capability.
Validation after the response
The response is another trust boundary. Preserve its source, retrieval time, content type, size, and policy decision. Parse documents in a constrained worker, disable active content where possible, and apply the rules in Secure Rendering of LLM Output before showing fetched text in a browser. Retrieved instructions must remain untrusted context; they cannot authorize a second fetch or tool call.
Verification and change control
Keep a test inventory for denied address ranges, alternate ports, malformed URLs, redirects, DNS changes, oversized responses, unsupported media, and credential stripping. Run it whenever the runtime, HTTP client, resolver, proxy, cloud network, or agent tool schema changes. Review allowlist changes like code: identify the business owner, affected tenant scope, expiry or review date, and rollback procedure. A gateway is only dependable when its decisions are observable and its policy can be changed quickly during an incident.