Platform showcase — live registry figures, the node map and the provenance record
Documentation

Guide

Data export

Getting records out of the registry in a form your statistician, your EDC or your local system can read — and what the platform records when you do.

Formats

Four export formats are supported: csv, xlsx, cdisc_odm and fhir_bundle; anything else is rejected by the request schema with a 422. All four export the same records — the same scope and the same filters — and differ only in shape.

CSV

One flat row per patient, with the plaintext aggregate columns of the record: study and subject identifiers, site, status and completeness, tumour type, location and histology, clinical and pathological stage components, neoadjuvant therapy, surgical approach, resection margin, nodes harvested and positive with the node ratio, complication grade and count, ICU and hospital days, the three mortality flags, readmission and transfusion flags, follow-up months, recurrence and vital status. Booleans are written as Yes/No, nested values as JSON. This is the format to use for a statistics package: it is the smallest, the fastest and the easiest to read into R, Stata or pandas. An export matching no records returns a file containing No records found rather than an empty file.

XLSX

An Excel workbook whose Overview sheet carries the same flat columns as the CSV. When the export includes PHI, the workbook gains one sheet per eCRF section, each with the study number in the first column and that section's fields across. Use it for review and circulation. If the server cannot produce a real workbook it falls back to CSV, which you can tell from the file name — it ends _xlsx_fallback.csv.

CDISC ODM XML

A CDISC ODM v2.0 Snapshot document: one SubjectData element per patient keyed by the study number, with each eCRF section rendered as a StudyEvent containing a Form, an ItemGroup and one ItemData element per field. Hand it to a clinical data management system or an EDC that speaks ODM, when the receiving side wants structure and provenance rather than a table. Without PHI, the aggregate fields are grouped back into their sections; with PHI, the decrypted section contents are used instead.

FHIR Bundle

A FHIR R4 Bundle of type collection in JSON, with a Patient resource per record carrying the study number, subject id and site id as identifiers, plus — where the data exists — a Condition for the tumour with its clinical and pathological staging, a Procedure for the resection, and Observation resources for nodes harvested, positive nodes, Clavien-Dindo grade, hospital stay, follow-up duration and vital status. Use it to load records into a FHIR server or a hospital integration engine. The codings carry human-readable display text rather than fully bound SNOMED codes, so a receiving system will usually need its own mapping step; with PHI included, sex and age are added to the Patient resource.

Scoping

The institution an export covers is resolved on the server from your account. You cannot widen it.

  • Site users (every role that can enter data) are pinned to their own institution. Naming a different institution in the filters is a 403, not a silent no-op.
  • Study administrators may name an institution to scope the export to it, or omit it entirely to produce a global export spanning every institution in the study. Only this tier can do that.
  • An account with no institution attached cannot export at all — it fails closed with “Your account is not linked to an institution” rather than being handed the global dataset.
  • Masquerade narrows an administrator to the target institution for the duration.

The same check runs again when you read or download an existing job, so knowing a job id is not enough: a global export job (one with no institution) is treated as the most sensitive job in the system and is readable only by a non-masquerading cross-institution role.

Job lifecycle

  1. Create

    Post the format, whether PHI is included, and any filters. A job row is created with status generating, the export runs, and the job comes back complete with a record count, a file size and an expiry — or failed. The call is synchronous: the response you get already reflects the outcome.

    A failed job carries a failure_reason. It names the class of failure and the job id to quote, but not the underlying error — that stays in the server logs. If a large export fails repeatedly, narrow it with filters and give an administrator the job id.

  2. Poll

    Fetch the job by id to see its status, record count, size, filters and expiry. If the expiry has passed, the status you get back is expired.

  3. Download

    Download returns the file as an attachment. A job that is not complete answers 400 with its current status; an expired job answers 410; a PHI export answers 403 unless you are a study administrator — every time it is downloaded, not only when it was created.

    The file is regenerated at download time from the current database, using the job's stored format, scope and filters, not from a stored artefact. Two downloads of the same job days apart can therefore differ. Treat the file you downloaded, not the job, as your snapshot.

bashCreate, poll and download an export
# 1. create
curl -X POST http://localhost:8001/api/v1/data/export \
  -H "Authorization: Bearer $TOKEN" \
  -H "X-Study-ID: tiger" \
  -H "Content-Type: application/json" \
  -d '{"export_format":"csv","include_phi":false,
       "filters":{"tumor_type":"adenocarcinoma","min_completion_pct":50}}'
# → {"id":"<job-id>","status":"complete","record_count":312,...}

# 2. poll
curl http://localhost:8001/api/v1/data/export/<job-id> \
  -H "Authorization: Bearer $TOKEN" -H "X-Study-ID: tiger"

# 3. download
curl -OJ http://localhost:8001/api/v1/data/export/<job-id>/download \
  -H "Authorization: Bearer $TOKEN" -H "X-Study-ID: tiger"

Completed exports expire seven days after they are created. Past that point the job is not downloadable and you create a new one. Your institution's recent import and export jobs are listed together in the history view — see Bulk operations.

Filters

All filters are optional and combine with AND. With one exception, unrecognised or malformed values are ignored rather than rejected, so check the filters echoed back on the job to confirm what was actually applied. The exception is pathologic_stage, which is validated up front and answers 422 on a value it does not know.

FilterApplies to
date_from, date_toRecord creation timestamp (ISO 8601) — when the record was entered, not the date of surgery.
tumor_type, tumor_location, surgery_approach, clinical_stageExact match on the recorded value, unvalidated.
pathologic_stageA pathway-qualified token: pIIIA and ypIIIA are different populations and the same group name exists on two AJCC 8th tables, so an unqualified IIIA is a 422. ambiguous is also valid. It matches only records the staging engine derived, not a stage typed as free text.
completion_statusRecord status: draft, complete, verified or locked.
min_completion_pctMinimum completeness — for “analysable only” extracts.
institution_idStudy administrators only; for anyone else, a 403.

Records are ordered by study number, and soft-deleted records are never exported.

Test records are included

There is no filter for test data and the exporter applies none: a batch imported as a test batch is promoted with a test flag on every record, and those records export like any other. The analytics dashboards exclude them, so an export and a dashboard reading the same population can disagree. Reconcile the export against record_count before you analyse it.

Watermarking & audit

Every export writes an audit entry when the job is created and another every time the file is downloaded. Both record your identity, your role, your IP address, your browser, the study, the format and the number of records returned; the filters are on the creation entry only, so read the two together. Downloads are audited separately so that a file fetched repeatedly, or long after it was created, is visible as such.

Watermarking. The record-level CSV export produced by the administrative analytics query builder is streamed with a watermark footer appended after the last row. It records the authenticated operator's id and email — never narrowed by masquerade — the masquerade target if any, the active institution, the export timestamp, a hash of the query that produced it, and the row count. A leaked file can therefore be traced to the operator, the institution it was taken on behalf of, and the moment it was taken. It ends with a fixed string, small_cell_suppression: k=5, which says nothing about the file above it: the query export is record-level, and k = 5 suppression is an aggregate control that is not applied to it.

Data Exchange exports are audited, not watermarked

The four Data Exchange formats described on this page carry no embedded watermark — the trail is the audit log, not the file. Treat an exported file as an uncontrolled copy the moment it leaves the platform: store it on institutional infrastructure covered by your data agreement, and do not forward it.

PHI handling

Exports come in two strengths, and the difference is not cosmetic:

  • Without PHI (the default) the export contains the plaintext aggregate columns only. The encrypted sections of the record are never decrypted.
  • With PHI (include_phi: true) the encrypted sections — demographics, diagnostics, surgery, pathology, complications and follow-up — are decrypted and flattened into the output, prefixed with their section name so they cannot collide with an aggregate column. Requesting this requires a study administrator role, as does every download of that job. The file name carries a _with_phi marker.

A non-PHI export is still patient data

The default export is pseudonymised, not anonymised. It contains one row per patient with their study number, their site, their stage, their node counts, their complications and their vital status. The k = 5 small-cell suppression that protects the analytics dashboards is an aggregate control and does not apply to record-level exports. A rare combination of stage, approach and site can identify a patient to someone who knows the case. Handle every export under your institution's data-protection agreement, and share only the minimum necessary.

Two limits before you plan a large extract. Export generation is synchronous and buffers the whole result in memory, so narrow very large exports by date range or completeness; and because the file is rebuilt at download time, downloading a large job costs that work again.

To export the aggregate figures behind the dashboards rather than records, see the export section of the analytics guide.