Documentation
Guide
Data import
How a spreadsheet of existing patients becomes reviewed, attested eCRF records — and why nothing reaches the live dataset until a human has approved every row and signed for the batch.
Supported files
The import API accepts two file types: CSV and XLSX. Anything else is refused with Unsupported format. Supported: csv, xlsx. The declared content type must also be on an allow-list and consistent with the file extension — a .csv announced as a spreadsheet is rejected with “File extension does not match the declared content type”.
- Maximum size: 50 MB at the API. The wizard applies a stricter 25 MB guard in the browser before uploading, rejecting larger files locally with “File is larger than 25 MB. Split it or contact an administrator.”
- Empty files and files with a header row but no data rows are rejected.
- Encoding: CSV is read as UTF-8 and a byte-order mark is tolerated, so a file exported from Excel as “CSV UTF-8” works.
- Headers are normalised before anything else: lower-cased, with spaces and dashes turned into underscores. Cell values are trimmed, and a row that is blank in every column is dropped.
Only CSV and XLSX
Some screens still offer CDISC ODM, REDCap or FHIR as an import format; the API rejects them. Flatten such data to CSV first.
The import template
The template is an XLSX workbook generated on demand for the current study, with one column per eCRF field key, drop-downs seeded with the controlled vocabulary for each coded field, and the field help text carried across. Because its column headers are the field keys, a file built from the template is mapped by identity — no guessing, and no AI step.
Two ways to get it:
- In the import wizard, choose Download the template on the first step and press Download blank template.
- On the Data Exchange page, use the template download in the import panel.
Both call the same endpoint. Any signed-in role may download it; you do not need import rights to look at the field list.
GET /api/v1/data/import/template
X-Study-ID: tiger
Authorization: Bearer <access_token>
→ 200 application/vnd.openxmlformats-officedocument.spreadsheetml.sheet
Content-Disposition: attachment; filename=ecrf_import_template_tiger.xlsxAn optional dataset_version query parameter selects a specific dataset version of the template. The list of available templates, with their full column list, is also available as JSON — see Bulk operations.
The wizard, step by step
The wizard has six steps and keeps its position in the URL, so refreshing or coming back later resumes the same job. No patient data is kept in browser storage — only the opaque job id travels in the query string, and every row is re-fetched from the server.
Choose path
Pick between an AI-assisted upload — a spreadsheet in any layout, whose columns and coded values the harmonizer proposes a mapping for — and the template, where you fill in a blank workbook whose columns are already the eCRF field keys. Both paths end in the same review-and-sign flow; the template path skips the guessing. Nothing has been uploaded yet.
Upload
Drop or pick the file. Before anything leaves the browser the wizard checks the extension and the 25 MB limit, parses the file locally and shows you the first eight rows with their headers, so you can catch an obviously wrong file — a shifted header row, an export of the wrong sheet — without a round trip.
You can tick test batch here. A test batch is promoted as normal but its records are flagged, and flagged records are excluded from the dashboard aggregations and analytics. Use it for a dry run of a migration.
On upload the server files the job under your effective institution — your own, or the masquerade target if you are masquerading. No field in the request can name a different institution, and an account with no institution cannot import at all.
Map fields
The server has now parsed the file and produced a mapping plan: for each source column, the eCRF section and field it should feed, a confidence score and a one-line rationale; and for coded columns, a value-level map from each distinct source value to the controlled vocabulary. Columns it could not place are listed as unmapped.
Review it and correct anything wrong. Saving a changed plan re-runs the deterministic transform over every staged row from its stored source values — you never have to re-upload to fix a mapping. Setting a column's section to empty drops it back to the unmapped list, which is how you tell the system to ignore a column.
Review
Every row is shown with its source values beside the proposed eCRF values, flagged rows first. You can correct a value inline, approve a row, or reject it. Corrections are stored separately from the proposal, so the original import is never overwritten and the reviewer's change is auditable.
Approve all clean rows approves everything without a flag in one action; rows carrying a flag always need an explicit decision from you. You cannot leave this step with rows still pending.
Agreement
Read the attestation, tick the acknowledgement and type your full name as an electronic signature. This is the gate: the batch cannot be promoted until it is signed. See below for what is recorded.
Import (promote)
Promotion writes the approved rows into your institution's patient dataset and shows a summary of how many were inserted, skipped and failed. If your institution is still in the parallel run against the legacy system, promotion is refused with a 412 and an explanation — the staged, reviewed and signed batch is kept and can be promoted after cutover.
Field mapping rules
Mapping happens in two layers, and only the second one is deterministic transformation:
- Column → field. In template mode this is identity: a header that matches a known field key is used as-is, and anything else is listed as unmapped and ignored. In AI-assisted mode the harmonizer proposes the mapping from the column headers plus a capped sample of distinct values per column (at most 50), never from whole rows.
- Value → controlled vocabulary. For coded fields the plan can also carry a per-value map, so a source that writes “M” or “male” lands on the vocabulary term the eCRF expects.
Each transformed value is then coerced against the target schema descriptor for that field, deterministically and with no model:
- Booleans accept the obvious tokens in either direction (
true/yes/1/y/t/present/positiveandfalse/no/0/n/f/absent/negative). Anything else is flagged rather than guessed. - Numbers are parsed with thousands separators stripped, truncated to a whole number where the field is an integer field, and checked against the field's minimum and maximum. A value outside the range is still written but flagged, so you can see and correct it. A boolean is never accepted as a number.
- Coded fields are matched case-insensitively against the allowed options. A value that is not in the vocabulary after the value-map has been applied is not written at all — it is flagged.
- Empty cells are simply not written. They are not zeros and not empty strings.
Partial uploads are fine
Required-field checks are scoped to the sections your file actually touches. Uploading only demographics does not flag every required field in pathology and follow-up as missing — only required fields in a section you supplied at least one value for.
In practice the check is dormant: the import target schema marks no field as required, so no upload can raise a missing_required flag. The twenty-one fields the data dictionary lists as mandatory are marked in the form definition, not enforced on import — do not treat a clean import as evidence they are populated.
Validation & the errors report
Staging assigns each row a review status and a list of flags. A row with no flags is pending (clean, awaiting approval); a row with any flag is needs_review and must be dealt with explicitly. Four flag codes exist:
| Flag | Meaning | What to do |
|---|---|---|
unmapped_value | A value could not be coerced to the field's type, or is not in its controlled vocabulary. | Add a value mapping on the mapping step, or correct the value inline on the review step. |
out_of_range | A number is non-numeric, or falls below the field minimum or above its maximum. | Check the unit — a height in metres against a field expecting centimetres is the usual cause. |
missing_required | A required field in a section you supplied data for has no value. Not raised today — no field is marked required on import. | Fill it in as a correction, or accept the record knowing the section will be incomplete. |
ambiguous | A column was mapped to a field name the target schema does not contain. | Fix the mapping plan; the column is otherwise ignored. |
The job also carries a harmonization score: the percentage of staged rows that came through with no flags.
If you want to check a file before staging it, there is a dry-run validation endpoint that parses the file, counts the rows and reports every column that is not a recognised eCRF field, without creating a job or calling the harmonizer. It is the one call that returns a flat errors list — each entry with a row index, a field, a message and a severity. See Bulk operations.
Read the flags on the rows, not the job's errors endpoint
A staged job also exposes an errors endpoint, and for a job created by the wizard it always comes back empty: it reads a job-level field the wizard path never fills in. The findings live on the staged rows themselves, as the flags above, and you see them on the review step or by fetching the rows. An empty errors report is not a clean batch.
The agreement step
Signing the submitter agreement is what turns a reviewed batch into a promotable one. You attest, on behalf of your institution, that the data is accurate and comes from your institution's own records; that you have reviewed the automated harmonization and accept the value mappings and corrections applied; and that you are authorised to import it into your institution's dataset.
Three things are written in a single transaction:
- a sign-off record carrying your user id, the name you typed as your signature, the acknowledgement flag, the verbatim attestation text you accepted, your IP address, your browser, the timestamp and the batch it covers;
- an append-only consent record against the versioned agreement document, carrying the number of staged rows as context;
- the job's status advancing to awaiting promotion, with a pointer to the sign-off.
If any part fails, the whole transaction is rolled back — a batch can never reach the promotable state without a durable, matching sign-off and consent pair.
Countersigned PDF
A PDF rendering of the signed agreement is dispatched after the transaction commits, best-effort: the sign-off is durable whether or not the render succeeds. The render task is currently a stub, so the response may return no PDF reference — the legal record is the sign-off and consent rows, not the PDF.
Promotion
Promotion is the only step that writes to the live patient dataset. Everything before it lives in a separate staging table, encrypted at rest, visible only to reviewers scoped to that institution.
Before any write, three preconditions are checked and fail closed:
- the batch has a signed agreement;
- the batch has an institution;
- every row being promoted is
approvedorcorrected— the query filters to those two statuses and each row is asserted again before it is written.
For each row the reviewer's corrections are merged over the proposed values, the result is grouped into eCRF sections, and a content hash is computed. The row is then matched against existing patients by a stable key derived from the row itself, and one of three things happens:
- insert — no existing patient for that key;
- update — a patient exists and the content hash has changed;
- skip — a patient exists and the hash is identical, so there is nothing to write.
That is what makes re-running an import safe: the key is the mapped study number when the file has one, otherwise a hash of the whole source row, so importing the same spreadsheet twice updates or skips rather than duplicating. Imported patients are tagged with their own source system, in a separate identity space from the legacy migration.
Read the summary counts carefully
The three outcomes above are not what the summary counts report. In the wizard's result panel, Promoted counts inserts only and Skipped is updates and unchanged rows together — so a run that rewrote a hundred existing patients shows a hundred “skipped”. The job record in the history view splits them differently again: there, imported is inserts plus updates. Only the per-row promotion run log distinguishes insert from update from skip, which is what to read when you need to know whether a re-import overwrote data someone had edited by hand.
Promoted records are created with status draft, and their completion percentage is recalculated from the sections the import actually filled. A row that fails is recorded with its error and the run continues — one bad row does not abort the batch.
How the job status flips to completed. Promotion itself only stamps the counts. The job is marked completed when no open rows remain — every non-deleted row either promoted or explicitly rejected. Anything still pending, under review, or approved but not successfully promoted keeps the job in awaiting promotion, so you can run promotion again; re-running is safe by the same hash rule.
Permissions
| Action | Minimum role |
|---|---|
| Download the import template, list templates | Any signed-in role |
| View a job, its mapping plan, its staged rows, its errors | Any signed-in role, scoped to that institution |
| Start an import, validate a file, edit the mapping plan, correct or approve rows, sign the agreement, promote the batch | Institution administrator or higher |
Scope is enforced on every one of these calls, not just the first: a job belonging to another institution answers 404/403 even if you know its id, and the check fails closed for a job with no institution. Masquerading narrows you to the target institution for the whole flow.
Study administrators are the exception to “or higher”
The study administrator role sits above institution administrator but is read-only: it can run a dry-run validation and export, but cannot start an import, edit a mapping, correct a row, sign the agreement or promote a batch.
Audit trail
Each import leaves several independent trails:
- an audit log entry when the batch is staged, with the mode, the file name, the row count and the resulting status; and another when it is promoted, with the result counts — both carrying your identity, role, IP address, browser and the study;
- a per-row run log written under one run id for the promotion, recording for every row whether it was an insert, an update, a skip or an error, together with the content hash before and after;
- the sign-off and consent records from the agreement step;
- the staged rows themselves, which keep the original source row, the proposed values, the reviewer's corrections and a per-field confidence score — encrypted, and retained with the job.
The file's SHA-256 checksum and byte size are stored on the job, so a batch can be tied back to the exact file that produced it.
Common errors
| What you see | Why | Fix |
|---|---|---|
Unsupported format. Supported: csv, xlsx | A format other than CSV or XLSX was declared. | Convert the file to CSV or XLSX and retry. |
File extension does not match the declared content type | The client announced a type inconsistent with the extension. | Re-save the file with the correct extension; do not rename an XLSX to .csv. |
Uploaded file is empty / “No data rows found” | Zero bytes, or a header row with nothing under it. Fully blank rows are dropped before counting. | Check you exported the right sheet and range. |
413 file too large | Over 50 MB at the API (25 MB in the wizard). | Split the spreadsheet — see Bulk operations. |
You must belong to an institution to import data | The account has no institution on its record. | Ask a study administrator to attach your account to a site. |
Job status is failed right after upload | Parsing or harmonization raised. The reason is stored on the job as an agent note rather than returned as a 500. | Open the job to read the note; usually a malformed CSV or an unreadable workbook. |
409 on promotion, mentioning the agreement | The batch has not been signed. | Complete the agreement step and promote again. |
412 — “institution is in the parallel run” | Your institution has not cleared its cutover gate, so direct patient writes are blocked. | Nothing to do. The batch is held and promoted after cutover. |
Many unmapped_value flags on one column | The source uses codes the vocabulary does not contain (for example numeric codes for a text-coded field). | Add the value mappings once on the mapping step; the transform re-runs over every row. |
| Promotion reports many skipped | Skipped covers two different outcomes, only one of which means nothing happened. | Do not read it as “no data was written” — see Promotion for what the counts actually mean and how to tell the two apart. |