Available as an add-on on paid plans.
Enabling storage on a task
Storage is configured when you create or update a task. Setstorage.enabled to true to capture files. Set storage.extraction to true to also extract structured data from those files.
Captured files appear in storage, not the output schema. Don’t add fields like file names or file counts; they duplicate what storage returns and can drift from what was actually captured. The storage object is the source of truth for captured files and downloads.
Retrieving storage items
Pass?include=storage when you fetch a task run and its captured files come back embedded in the response, each with its extracted data and a pre-signed download url inline:
Storage item fields
Unique identifier, prefixed with
stor_.Original file name as it appeared on the source.
MIME type (
application/pdf, image/png, text/csv, etc.). See Supported files for supported types.Size in bytes.
attachment for a file you provide as task input for the agent to use, extraction for a file you provide as task input that Deck extracts data from directly (skipping the agent), or output for a file the agent captures during the run.When the storage item was created.
Structured data extracted from the file, if extraction is enabled.
Extraction outcome:
failure if the file’s data couldn’t be extracted, otherwise success (including files with no extraction). Returned by the list and get-by-id endpoints.Signed download URL.
A JSON object of additional key-value details Deck records about the file, or
null when there are none. Which keys appear depends on the file and how it was captured, so treat it as informational and don’t depend on specific keys being present. Returned by the list and get-by-id endpoints.The task run that produced this storage item.
Downloading files
Both list and get-by-id responses include a pre-signedurl you can use to download the raw file. URLs are time-limited; if one expires, re-fetch the item to get a fresh one.
Supported files
These rules apply to every file Deck stores: files an agent captures during a run (purpose: "output"), files you provide as task input (attachment), and files Deck processes directly (extraction).
File types
The following types can be captured and returned through storage. The Extraction column shows which can also be parsed into structured JSON when extraction is enabled.File size
Each captured file can be up to 200 MB. Files you provide as task input have a smaller limit. See Sending a file.Providing files as input
Tasks can accept files as input. The file is uploaded to storage and malware-scanned, then either handed to the agent or extracted directly by Deck, depending on the field’spurpose:
Both purposes share the same field shape and upload behavior, and differ in what happens after the upload. A single task can declare attachment file inputs or extraction file inputs, not both. Creating a task whose input schema mixes both is rejected with an
input_invalid error.
Defining the field
Define a file field in the input schema. The shape is the same for both purposes; only thepurpose constant changes.
Sending a file
Provide the file inline as base64 in the task run input:Attachments
Withpurpose: "attachment", the agent receives the file at run time and uses it on the source: uploading it to a portal, attaching it to a form, or referencing it while completing the task.
The run stays queued until every attachment passes the scan, then transitions to running and dispatches to the agent. If a file is flagged, the run fails with an attachment_invalid error on the task run object. Listen for task_run.failed or fetch the run to handle it.
Deck replaces the base64 data in the stored input with a storage_id reference so the raw bytes aren’t carried through the run. The file becomes a storage item with purpose: "attachment", alongside the output files the agent captures, and appears under the Input tab on the task run in the Console.
Extraction
Withpurpose: "extraction", Deck processes the file directly against the task’s extraction_schema and returns the structured result on the run. The agent doesn’t execute. Use this when you have a document and want JSON back, with no source interaction.
The run transitions to running as soon as the file is uploaded, and finalizes when extraction completes. It doesn’t sit in queued waiting on the scan.
The task must have storage and extraction enabled, with an extraction_schema defining the result shape:
credential_id or source_id to link the extraction to a user or source, even though the agent doesn’t execute.
purpose: "extraction" carrying the extracted data, matching the fields defined in extraction_schema. See Document extraction for guidance on writing extraction schemas.
Reusing a file across runs
Once a file is uploaded, you can reference it on a later run instead of sending the bytes again. Pass thestorage_id in place of data, keeping the same purpose:
purpose: "extraction". Deck verifies the file belongs to your organization, then copies it for the new run so each run keeps its own input files.
Document extraction
When extraction is enabled, Deck parses the captured files and populates theextraction field on each storage item with structured JSON.
Extraction is available for PDF, CSV, Excel (.xlsx, .xls), plain text, JSON, and images (PNG, JPEG, WebP, TIFF). The extracted data depends on the document. A utility bill produces different fields than a hotel receipt.
Custom extraction schemas
Use theextraction_schema field on the task’s storage config to define exactly what fields you want extracted. Deck uses this schema to guide parsing.
Guiding extraction with a prompt
extraction_prompt is optional free-text guidance applied alongside the schema. Set it on the task’s storage config to give the extraction model instructions that the schema’s field names can’t capture on their own: how to disambiguate similar fields, what format to return a value in, where on the document to look, what to assume when a value is missing, or fixes for mistakes you’ve seen it make.
extraction_schema. It’s configured once on the task and applies to every file that task extracts. When omitted, extraction runs on the schema alone.
Extraction example
A utility bill extraction might return:Extraction errors
Each storage item reports its extraction outcome in aresult field: failure if the file’s data couldn’t be extracted, otherwise success. Note that success also covers files with no extraction, so it doesn’t on its own mean data was extracted.
When extraction fails, result is failure and extraction stays null, but the raw file is still kept and available for download. Check result to tell which files failed.
If any file in a task run fails extraction, the run completes with a failure result and an extraction_failed error in the errors array indicating how many files were affected. Successfully extracted files in the same run still return their extraction data.
Deduplication
Deduplication tells Deck to skip files that match one captured by a previous run for the same task and credential, so recurring tasks only return new documents. You define a set of fields that uniquely identify a document. Deck reads those fields from each captured file and compares them against prior captures. If every field matches, the new file is dropped: it’s not stored, nostorage.created event fires, and it’s not extracted even if extraction is enabled.
Configuration
Setdeduplication to true and provide a deduplication_schema on the task’s storage config:
Each property must declare a
type, one of string, integer, number, or boolean. Nested objects and arrays aren’t supported, so pick top-level scalar fields.
Property names are arbitrary; you make them up, and they’re just keys for the result. The description is what tells Deck where to find the value on each document, so describe each field precisely. For example, "Account number printed at the top of the bill" works better than a vague "account".
Choosing fields
The fields you list together form the dedup key. Two files match only if every field is identical. A few rules of thumb:- Pick fields that stay stable for the same logical document. A monthly bill should have the same account number and billing period every time it’s fetched.
- Avoid volatile fields. File names, fetch dates, and page numbers will produce false negatives, since the same document looks new every time.
- Pick enough fields to be unique. A single field like
vendor_namewill collide across unrelated invoices from the same vendor. - Two or three fields is usually right.
Field combinations by document type
- Utility bills
- Invoices
- Receipts
- Bank and credit-card statements
Account plus billing period:
Errors
deduplication_schema must be present with at least one property whenever deduplication is true. Starting a task run without it returns:
deduplication to false (or omit it entirely).