Protect documents
Upload a PDF or DOCX file, run a document job, and download the redacted PDF, the protected text and the entities.
A document job reads a PDF or DOCX file. It returns the redacted PDF, the protected text and the entities. It runs in the background at the batch weight.
Document job workflow
- Upload the file:
POST /v2/uploads - Start a document job with the upload ID:
POST /v2/jobs - Poll the job until its status is succeeded:
GET /v2/jobs/{id} - Download the artifacts, then delete the job.
Send the file bytes with application/pdf or the DOCX media type. A document is at most 10 MB.
Upload the file and start the job
UPLOAD=$(curl -s https://api.getshinrai.com/v2/uploads -H "Authorization: Bearer $SHINRAI_API_KEY" \
-H "Content-Type: application/pdf" --data-binary @contract.pdf | python3 -c 'import sys, json; print(json.load(sys.stdin)["id"])')
JOB=$(curl -s https://api.getshinrai.com/v2/jobs -H "Authorization: Bearer $SHINRAI_API_KEY" -H "Content-Type: application/json" \
-H "Idempotency-Key: contract-4815" \
-d '{"kind": "document", "inputs": [{"kind": "file", "source": {"upload": "'$UPLOAD'"}, "language": "de"}]}' \
| python3 -c 'import sys, json; print(json.load(sys.stdin)["id"])')The answer is 202 with the job and a Location header. A retry with the same Idempotency-Key and body returns the same job and is not charged again.
Poll and download artifacts
curl -s https://api.getshinrai.com/v2/jobs/$JOB -H "Authorization: Bearer $SHINRAI_API_KEY"
curl -s https://api.getshinrai.com/v2/jobs/$JOB/artifacts/protected -H "Authorization: Bearer $SHINRAI_API_KEY" -o contract.redacted.pdf
curl -s https://api.getshinrai.com/v2/jobs/$JOB/artifacts/text -H "Authorization: Bearer $SHINRAI_API_KEY" -o contract.protected.txt
curl -s -X DELETE https://api.getshinrai.com/v2/jobs/$JOB -H "Authorization: Bearer $SHINRAI_API_KEY"Use the artifact URLs returned by the status response. Job and artifact access is restricted to the account that created the job.
| Artifact | Content |
|---|---|
protected | The redacted PDF: image-only pages at 144 dpi, with a black box over every protected entity and no text layer |
text | The protected text of the document |
entities | JSON: the entities with offsets into the extracted text |
mapping | JSON: the original values and their replacements, only on request |
The redacted PDF has no text layer. Take the protected text from the text artifact. Ask for the mapping only when you must restore values: it contains the original values.
What the service keeps
- The upload is deleted when the last job that reads it ends.
- An upload created with
POST /v2/uploads?keep=truelives 24 hours, extended by every job that reads it. - Deleting the job also deletes a kept upload as soon as no other job reads it.
- The results stay for 24 hours. Delete the job to remove them earlier.
- Other accounts cannot read your jobs: they get 404.
Handle failures safely
Do not forward the original file when protection fails.
| Status | Meaning |
|---|---|
415 | The file is not a PDF or DOCX file, or its bytes do not match its media type. |
413 | The file is larger than 10,000,000 bytes. |
402 | Your balance holds too few records. |
422 | The request is invalid, for example because the upload expired. |
429 | The job storage of your account is full, and limit_name names the limit. Delete finished jobs. |
429 | The job queue is full. Wait for the time in the Retry-After header. |
503 | The job service is unavailable, or its storage is full. Retry after the time in the Retry-After header. |
- A document that cannot be read makes the job fail. The job then has the status failed and an error code.
- Retry only when the error says that a retry can succeed, and wait for the time in the Retry-After header.
Cost
A document job costs the records of its extracted text at the batch weight 0.5. Your balance must hold at least one record when you start the job.