Web Security
File Upload Security for AI Document Pipelines
Defensive guidance for securing AI document uploads before parsing, OCR, chunking, embedding, indexing, or LLM retrieval.
Uploaded files are untrusted input. Authentication proves who submitted a request; it does not prove that a document is safe to parse, index, or show to a model. A successful upload, OCR run, or PDF conversion is not a security verdict. Treat file ingestion as a chain of trust with explicit gates.
AI document ingestion trust pipeline
RECEIVE → CLASSIFY → QUARANTINE → INSPECT → PARSE IN ISOLATION → VALIDATE EXTRACTION → ATTACH PROVENANCE → AUTHORIZE INDEXING → INDEX → MONITOR / DELETE
The upload gateway authenticates the user and resolves the tenant or workspace from server-side membership. It then applies size and type policy before writing to quarantine. Inspection and parsing occur in constrained workers. Only an approved result is promoted to a retrieval index.
Authorize the upload
Check who may upload and which tenant, project, or knowledge base receives the file. Never trust a tenant ID supplied by a browser or model without validating membership and resource ownership. A user can belong to several tenants, so an explicit workspace choice must be bound to an authenticated session and checked on every object operation. The same rule applies to replacement and deletion.
Give every upload a server-generated file ID. Preserve the original filename only as metadata. Storage keys must not be constructed from user-controlled paths. Keep raw objects private, use tenant-scoped prefixes or buckets, and issue short-lived signed access only when a user is authorized.
Size, type, and resource policy
Set limits for compressed bytes, decompressed bytes, page count, image dimensions, archive entries, nesting depth, parser CPU, memory, wall-clock time, and resulting token or embedding volume. A small archive can expand into millions of files; a long document can create an unexpectedly expensive indexing job. Enforce limits before and during extraction, and terminate workers that exceed them.
Filename extensions and declared MIME types are hints, not proof. Compare them with magic bytes or format signatures using maintained libraries. Detection is imperfect, so unsupported or ambiguous files should remain quarantined. Do not accept a type merely because a browser labelled it correctly.
Quarantine and inspection
Do not index immediately after upload. Store the object in quarantine, record its hash, and run malware or content inspection appropriate to the environment. Scanning reduces known-malware risk but cannot establish that a document is semantically truthful or free of malicious instructions. HTML, SVG, office files, and PDFs may contain active or embedded content; convert them to a safe representation or reject them when rich behavior is unnecessary.
Archives need separate controls: bounded extraction, entry-count limits, nested-archive limits, canonical destination paths, and rejection of unsupported links or special files. Never let extraction write outside a worker-owned temporary directory.
Isolate parsers and converters
PDF parsers, OCR engines, image codecs, office converters, and archive libraries are security-sensitive dependencies. Run them in a dedicated low-privilege worker or sandbox with a temporary filesystem, no host home-directory access, minimal network egress, and CPU/memory/process limits. Keep parser versions tracked and patched. A parser crash should fail one job, not the web process or neighboring tenant.
Extracted text remains untrusted
Converting PDF to text removes some binary complexity but not malicious meaning. A document can contain indirect prompt-injection instructions, false facts, or poisoned knowledge. Store extracted text with its source document ID and trust status. When assembling model context, label retrieved material as untrusted data and apply the controls described in indirect prompt injection guidance and RAG security guidance. Parsing is not authorization.
Provenance and promotion
Attach uploader, tenant, original name, detected type, cryptographic hash, timestamps, parser and scanner versions, scan status, source interaction, and promotion decision. A controlled promotion gate should require successful inspection, extraction validation, authorization, and—where appropriate—human review. Keep rejected and superseded states visible for investigation.
Versions, deletion, and derived data
Use a stable document identity plus immutable versions. Re-uploading the same bytes should be deduplicated intentionally; changed bytes should create a new version and invalidate stale chunks. Deletion must propagate to raw storage, extracted text, chunks, embeddings, retrieval caches, and generated summaries. Immediate erasure may be limited by backups or asynchronous indexes, so document the lifecycle and verify completion.
Upload security control matrix
| Stage | Risk | Control | Evidence | Failure action |
|---|---|---|---|---|
| Receive | unauthorized tenant upload | membership and object authorization | principal, tenant, file ID | reject and alert |
| Classify | spoofed type or oversized input | signature checks and byte/resource limits | detected type, measured sizes | quarantine |
| Inspect | malware or active content | scanner and safe-format policy | scan version/result | reject or hold |
| Parse | vulnerable converter or exhaustion | isolated worker and quotas | parser version, exit status | terminate job |
| Validate | malformed or surprising extraction | schema, page/text limits, sampling | validation result | quarantine |
| Promote | unsafe context enters RAG | explicit approval gate and provenance | promotion event | block indexing |
| Retrieve | leakage or poisoned instructions | tenant filter and untrusted-context handling | retrieval audit | revoke source |
| Delete | stale derivatives remain | propagation workflow and reconciliation | deletion report | retry and escalate |
Logging and testing
Record security events without dumping document contents: principal, tenant, file ID, size and detected type, hash, scanner result, parser version, state transitions, and policy decisions. Test oversized and corrupt files, wrong signatures, decompression limits, unsupported formats, parser failures, quarantine outages, cross-tenant access, duplicate versions, deletion propagation, and resource exhaustion. Include adversarial documents in controlled tests without distributing offensive payloads.
Practical checklist
- Resolve tenant and upload permission server-side.
- Generate storage names; keep originals as metadata.
- Enforce compressed and expanded resource limits.
- Compare declared type with signatures and reject ambiguity.
- Quarantine before scanning, parsing, or indexing.
- Isolate parsers with least privilege and restricted egress.
- Preserve hashes, versions, provenance, and state transitions.
- Treat extracted text as untrusted context.
- Filter retrieval by tenant before context assembly.
- Provide a promotion, revocation, and deletion workflow.
- Test parser failure, cross-tenant leakage, and cleanup.
Sources
OWASP File Upload Cheat Sheet, OWASP ASVS, CISA Secure by Design, and maintained parser or file-format security documentation provide the baseline. PermsAI’s pipeline and matrix are an architectural synthesis for AI ingestion.
Data flow boundaries
Separate the upload API, quarantine store, parsing worker, indexing service, and retrieval API where practical. Each boundary should carry an authenticated job identity and tenant context, not a free-form string copied from document metadata. Queue messages should contain an opaque file ID and constrained operation, while workers re-read authorization state before acting. A compromised parser must not obtain database credentials, model keys, or another tenant’s object-store access.
Content inspection should also consider embedded links, external references, macros, and unexpected encodings. Decide whether the product needs to preserve these features. If not, flatten documents to a safe text or image representation before indexing. If they are required, render them in an isolated preview path and keep the original private. Never expose a generated preview from a public executable directory.
Operationally, monitor queue age, rejection rates, parser crashes, repeated retries, indexing volume, and sudden changes in a tenant’s upload pattern. Alerting is useful evidence, but it does not replace authorization. During an incident, identify every session that retrieved a poisoned document, suspend promotion for the source, invalidate derived chunks, and rotate any credential that the parser could access. This makes the ingestion pipeline recoverable rather than a one-way import.