Documentation
Guide
Bulk operations
Everything on this page is about volume: staging hundreds of rows at once, reviewing them efficiently, checking a file before you commit to it, and getting a legacy spreadsheet into the registry without creating a mess you then have to unpick.
Bulk import
There is no separate “bulk” endpoint: the ordinary import flow is the bulk flow. One upload stages every data row in the file as its own reviewable record, and the whole batch moves through mapping, review, agreement and promotion together. A file of one row and a file of two thousand rows take exactly the same path.
Two consequences worth planning around. The mapping decision is made once for the batch and applies to every row, so fixing a column mapping once re-runs the transform across the whole batch. And the agreement is signed once per batch, so batch boundaries are also accountability boundaries: split by data source or by the person who can attest to the data, not arbitrarily.
Rows are staged encrypted, keeping the original source row, the proposed eCRF values, any reviewer corrections and a per-field confidence score. Nothing in the batch touches the live patient dataset until promotion. See Data import for the mechanics of each step.
Bulk review & promotion
The review step is designed for triage rather than row-by-row reading:
- Flagged rows are shown first, and the summary chips give the count of flagged, pending and approved rows.
- Approve all clean rows approves everything without a flag in one action. Rows carrying a flag are left alone and always need an explicit decision.
- Rows can be listed filtered by review status, so you can work through
needs_reviewuntil it is empty.
How “approve all” actually works
There is no server-side bulk-approve endpoint: the interface pages through the staged rows in pages of 200, issuing one update per row. On a large batch that takes a visible moment; let it finish rather than reloading the page.
Promotion then writes every approved or corrected row. It is idempotent by content hash: a row identical to one already promoted is skipped, a changed row updates the existing patient, and a new row is inserted. Running promotion twice on the same batch is therefore safe, and is the normal way to finish a batch worked through in several sittings — the job only flips to completed once no rows remain pending, under review, or approved-but-not-written.
Templates
The templates endpoint lists the eCRF import template available for the current study, with its name, a description and the full column list — every field key the importer recognises, in order. Diff your spreadsheet's headers against it before you upload anything.
curl http://localhost:8001/api/v1/data/templates \
-H "Authorization: Bearer $TOKEN" \
-H "X-Study-ID: tiger"
# → {"templates":[{"name":"ecrf_tiger",
# "description":"Tiger eCRF import template — all sections, …",
# "columns":["unique_id","gender","age", …]}]}
# the workbook itself (dropdowns + help text):
curl -OJ http://localhost:8001/api/v1/data/templates/ecrf_tiger/download \
-H "Authorization: Bearer $TOKEN" -H "X-Study-ID: tiger"Any signed-in role may list and download templates; the same workbook is on the first step of the import wizard. The human-readable version of the field list, with labels, types and allowed values, is the data dictionary.
The validation endpoint
Before committing to a full import you can post a file to the validation endpoint for a dry run. It parses the file, runs the same transform the staging path runs over every row, and reports the rows that would be flagged, the columns it does not recognise, and the row count. Nothing is stored, no job is created and the harmonizer is not called, so it is cheap enough to use repeatedly while you get a file into shape.
curl -X POST http://localhost:8001/api/v1/data/validate \
-H "Authorization: Bearer $TOKEN" \
-H "X-Study-ID: tiger" \
-F "file=@site-export.csv" \
-F "source_type=csv"
# → {"total_records": 412,
# "valid_records": 405,
# "failed_records": 7,
# "errors": [{"record_index":38,"record_id":"","field":"age",
# "severity":"error","error":"214.0 above max 120"},
# {"record_index":51,"field":"gender","severity":"error",
# "error":"'M' not in allowed options for gender"}],
# "warnings": [{"record_index":0,"field":"pat_age","severity":"warning",
# "error":"Column is not a recognized eCRF field — it will need
# mapping (AI) or is ignored (template). Its values are
# therefore NOT checked by this dry run."}]}What this check does not do
It cannot check a column it does not recognise. Unrecognised headers come back as warnings and their values are left unchecked — on the template path those columns are ignored anyway, but on the AI-assisted path the harmonizer may map one, and its values are then coerced and possibly flagged during staging.
It also has no value mappings, which come from the harmonizer or from you: a source that writes M where the vocabulary wants Male is reported here as unmapped even though the real run would canonicalise it. The dry run is therefore exact on the template path and pessimistic on the AI path — never optimistic.
failed_records counts rows carrying at least one flag; on the real run those rows are staged as needs_review, not dropped. Only the first 500 findings are listed; the counts cover every row.
Validation accepts the same formats as import (CSV and XLSX) and the same 50 MB ceiling, and requires the same institution-administrator role. It is the one part of the import flow a read-only study administrator may run.
History
The history view lists your institution's recent import and export jobs side by side. For each import you get its source type and name, its status, the total, imported, skipped and failed counts, its harmonization score and its timestamps; for each export, the format, status, whether PHI was included, the record count, the file size, the filters and the expiry.
It is scoped to your effective institution and fails closed: an account with no institution sees nothing rather than everything. Study administrators, not masquerading, see across institutions. Paging is one-based, maximum 100 per page, with one page number for both lists.
curl "http://localhost:8001/api/v1/data/history?page=1&page_size=20" \
-H "Authorization: Bearer $TOKEN" \
-H "X-Study-ID: tiger"
# → {"imports":[{"id":"…","source_type":"csv","source_name":"site-export.csv",
# "status":"completed","total_records":412,
# "imported_records":409,"skipped_records":3,
# "failed_records":0,"harmonization_score":97.6, …}],
# "exports":[{"id":"…","export_format":"csv","status":"complete",
# "include_phi":false,"record_count":312,"expires_at":"…", …}]}Limits
| Limit | Value | Where it applies |
|---|---|---|
| Upload size | 50 MB | Import and validation endpoints. Over it, HTTP 413. |
| Upload size (browser) | 25 MB | The wizard's pre-flight check, before anything is sent. |
| Rows per file | Not limited by the application | No row cap in the import path; the practical ceiling is the 50 MB upload limit and your browser's upload behaviour. |
| Staged rows per page | 200 maximum (50 by default) | Listing a batch's rows for review. |
| History page size | 100 maximum (20 by default) | The import/export history view. |
| Distinct values sampled per column | 50 | What the harmonizer sees when proposing value mappings; a column with more distinct values may need mappings added by hand. |
| Export expiry | 7 days | After which a completed export is no longer downloadable. |
| Request rate | 100 requests per minute | The default tier, used by the data-exchange endpoints; login and the analytics export are tighter at 10 per minute. Exceeding a limit returns 429 with Retry-After. |
No documented processing SLA
Import staging and export generation both run synchronously inside the request, so how long they take depends on the size of the file and — for AI-assisted imports — on the harmonization step. No completion time is defined or promised. Plan large migrations as a task you supervise, not a background job you can walk away from.
Migrating a legacy spreadsheet
The workflow with the fewest ways to go wrong:
Get the column list first
Download the template workbook and compare its column list against your spreadsheet's headers, then read the data dictionary for the fields you are unsure about — particularly the allowed values of coded fields, where most rework comes from.
Make one file per source, not one big file
If your data comes from several places — a surgical log, a pathology extract, a follow-up sheet — keep them separate. Partial uploads are supported: required-field checks only apply to the sections a file actually touches. Each file becomes its own batch with its own attestation.
Make sure every row has a study number
Include the anonymised study number column in every file. It is what ties rows across files to the same patient and what makes re-imports idempotent. Without it the platform falls back to hashing the whole row, which means an edited row will be treated as a new patient on the next run.
Dry-run, then stage a small slice as a test batch
Post the file to the validation endpoint to confirm it parses and to see the unrecognised columns. Then take the first 20–50 rows, upload them with test batch ticked, and walk the whole flow through to promotion: it shows exactly how your real data will be mapped, and test records are excluded from dashboards and analytics. But a promoted test record is a real row in your institution's dataset, and exports do not filter it out. Keep the slice small and remember it when you reconcile counts.
Fix the mapping once, at the batch level
Resolve problems on the mapping step rather than row by row: one value mapping fixes every row that uses that value, and saving the plan re-runs the transform across the batch. Use inline corrections only for genuine one-off data issues.
Import the real batches, largest last
Once the test batch promotes cleanly, upload the real files. Keep each under the size limit; split a large export by year or by ward rather than trying one giant file. Review, sign, promote — and check the summary counts before the next file.
Reconcile, then export to verify
Afterwards, export the same records back out as CSV and compare the counts and a sample of rows against your source. That round trip catches mapping mistakes that looked fine in the review screen. The history view keeps the record of what was imported when, with its checksum and counts.
If your institution is still in the parallel run
Promotion is blocked with a 412 until your institution clears its cutover gate: the legacy system is still the source of truth, and a direct write would make the two diverge. You can still upload, map, review and sign — the batch is kept in full and can be promoted once the gate opens. Do the work early, promote after cutover.