Platform showcase — live registry figures, the node map and the provenance record
Documentation

Guide

Analytics

What each figure on the analytics dashboard means, which records are behind it, what gets hidden to protect small groups, and which statistical tests are used when cohorts are compared.

Scope & access

Every analytics response is scoped on the server from your account, never from a parameter you send. There are two tiers:

  • Institution users — institution administrators and all clinical roles, including coordinators, fellows and trainees. You see your institution against the global benchmark, and nothing else. No other institution, country or region breakdown is returned to you. Passing an institution_id that is not your own is rejected with a 403, and an account with no institution on its record is refused rather than silently given global figures.
  • Study administrators (study_admin and project_admin) — cross-institution scope. Omitting the institution parameter returns global aggregates; supplying one scopes to that institution. This tier also has the Cohort Comparison workbench in the sidebar, which builds arbitrary cohorts from region, country, institution and clinical dimensions and compares them with the tests described below.

Masquerade. When an administrator is acting on behalf of an institution, analytics narrow to that institution for the whole session — the masquerade target wins over the administrator's own home institution. The cohort workbench is deliberately not available while masquerading: those endpoints answer 403 "Exit masquerade to compare cohorts". The rule is that a masquerading admin sees exactly what the institution sees.

All analytics read the plaintext aggregate columns of the patient record. They never decrypt an encrypted section, and they always exclude soft-deleted records. Batches uploaded as a test batch are flagged on the record and excluded from the dashboard aggregations, the administrative analytics and every cohort-comparison query — so a validation run of a spreadsheet does not move anyone's numbers.

Metric definitions

Many fields were never captured for many records, so every rate is calculated over the records that actually carry the field, and the API returns that denominator alongside the value (the n you see next to a figure). A metric with no denominator is returned as no value — never as 0 %.

MetricDefinitionNumerator / denominatorSource fieldCaveats
Total patientsRecords in scope for the selected period.Count of non-deleted records / —patients rowsThe period filter applies to created_at, i.e. when the record was entered.
Completion rateMean record completeness across the cohort.Mean of completion_pct / records in scopecompletion_pctA record's own completion is completed sections out of a fixed denominator of eight canonical sections.
Mean LN yieldAverage number of lymph nodes harvested.Mean of total_ln_harvested / records where that column is not nulltotal_ln_harvestedRecords without a node count are excluded from the mean, not counted as zero.
LN distributionMean, median, 25th and 75th percentile, plus a histogram in bins of five (0–4, 5–9, … 45–49, 50+).Percentiles over total_ln_harvested / records with a node counttotal_ln_harvestedPercentiles are computed in the database with percentile_cont.
Complication rate (assessed)Share of assessed patients with any recorded complication.Records with complication_count > 0 or a non-empty grade / records where complication_assessed is truecomplication_assessed, complication_count, complication_gradeWhere an institution has no assessed patients the API returns overall_status: "not_computable" with no number: that site is left out of every ranking rather than scored at 0 %.
Reoperation rateReturn to theatre. This is the headline morbidity figure.reoperation true / reoperation not nullreoperationA genuine population rate: the eCRF asks this as a Yes/No question, so a “No” is a recorded answer rather than an unticked box, and a patient nobody answered for drops out of the denominator instead of counting as a success.
Severe (≥ IIIa) share of graded patientsOf the patients who were given a Clavien-Dindo grade, how many were graded ≥ III.Grade >=III / records carrying any gradecomplication_gradeConditional, not a population rate. The denominator is “patients who were graded”, which is itself a selected group — a grade tends to be filled in when something happened. It answers “when a complication was graded, how often was it severe?”, and it must never be read as the share of ALL patients who had a severe complication. The API ships the full label with the number for this reason.
Complications by gradeSplit into two buckets only: Clavien-Dindo I–II and ≥ III.Records with that grade / records carrying any gradecomplication_gradeThe source form records the coarse split, not IIIa/IIIb/IVa/IVb/V. Reporting finer buckets would invent precision the data does not have.
30-day mortalityDeath within 30 days of the index operation.Records flagged true / records where the flag is not nullthirty_day_mortalityRecords with an unknown status are excluded from the denominator. A cohort where nobody has a known status returns no value, not 0 %.
90-day mortalityDeath within 90 days of the index operation.Records flagged true / records with a known statusninety_day_mortalitySame known-status rule as 30-day mortality.
In-hospital mortalityDeath before discharge.Records flagged true / records with a known statusin_hospital_mortalitySame known-status rule.
R0 resection rateMicroscopically complete resection.Records with margin R0 / records with any recorded marginresection_marginFeeds the R0 axis of the quality profile.
Volume trendRecords entered per month (or quarter/year), for your institution against the average institution.Count of records grouped by date_trunc(created_at) / — ; the global line divides each bucket's total by the number of institutions that contributed to itcreated_atThis is a data-entry curve, not a surgical activity curve. The response labels its axis basis: "created_at" for exactly this reason.
Quality profileSix normalised axes for the radar chart: volume, LN yield, R0 rate, low reoperation, low mortality, completion.Each axis is rescaled to 0–100 (see below)Derived from the metrics aboveThe scaling is presentational: volume is scored against twice the global average per institution, LN yield against 40 nodes, and the two “low” axes are inverted so that further out is better. Do not read an axis as a clinical rate. The morbidity axis scores the reoperation rate.
Completion by section / by userShare of records with each section marked complete, and per-user data-entry counts with their mean completeness.Records with the section flagged complete / records in scopecompletion_status, entered_byInstitution-scoped. It is a data-management view, not a performance measure.

Metrics this registry cannot report

Several familiar surgical metrics have a column in the schema but no data behind it in the migrated registry — length of stay, ICU days, 30-day readmission, blood transfusion, and the clinical and pathological stage groups. Surgery dates are likewise not held in plaintext: dates are de-identified by design, which is why the volume trend uses the record-entry timestamp rather than the operation date, and why length-of-stay analyses are absent rather than shown as zero. If a metric is missing from your dashboard it is because the registry has nothing to compute it from.

Field-wise profile

/dashboard/analytics/profile compares your institution against the whole registry on every plaintext eCRF field with data, one row per field, grouped by eCRF section. The catalogue holds 54 fields and is derived from the patient model rather than hand-maintained, so a field cannot silently drop out when the schema moves. A field with no values anywhere in the registry is listed under unavailable_fields with the reason, never drawn as an empty chart.

Each row carries:

  • Both figures with their denominators — mean, SD, median and IQR with a 95 % CI for a continuous field; events over n with a Wilson interval for a rate; the category breakdown for a categorical one.
  • The right test for the field's kind — Mann-Whitney U for continuous, Fisher / chi-square for rates and categories — plus an effect size with its own interval. A test is withheld whenever either side of it is withheld, because a published statistic against a published denominator reconstructs the hidden cell.
  • The chart, chosen by the API and not by the page. The response names one of box_with_points, histogram_overlay, rate_bar_with_ci, stacked_100_bar, grouped_bar, line_trend or none, and ships the data that chart needs in the same object — a histogram arrives with its bins already k = 5 masked. The same decision therefore holds in the CSV export and in any other consumer, so two views of one field cannot disagree about how to draw it.
  • A trend on created_at, yearly or quarterly, drawn only once at least three periods survive suppression.

Read the Holm column, not the raw p

This page runs one test per field, so a single page view is a few dozen simultaneous comparisons rather than one. At the conventional 0.05 threshold, roughly one in twenty tests comes back “significant” when nothing is going on — scan thirty comparisons looking for the interesting ones and you should expect to find one or two purely by chance. Reporting only the raw p-value would therefore hand every institution a finding.

The response reports p_value and p_value_holm side by side. The Holm-Bonferroni adjustment is computed across the whole family of tests in that response, and the family size is stated on the wire as tests_run, so the correction always matches the table actually in front of you. When you are scanning the page for something worth investigating, the Holm column is the one to read; the raw p is there for a hypothesis you specified before you opened the page.

Survival analysis

/dashboard/analytics/survival plots Kaplan-Meier curves for two endpoints — overall survival (event: vital status is dead) and recurrence-free survival (event: recurrence recorded) — optionally stratified by one of nine variables: pathologic T, N and stage group, resection margin (R0 against R1/R2 pooled), tumour histology, neoadjuvant therapy, lymphadenectomy extent, and two derived lymph-node bands (yield below or at least 15 nodes, and ratio band). The bands are computed in the query and never stored.

Each curve is reported with:

  • Greenwood confidence bands, log-log transformed rather than plain linear. A linear band on a survival probability runs outside [0, 1] near the ends of the curve and misstates its own coverage; the log-log band cannot. This is the same default R's survfit(conf.type = "log-log") and SAS PROC LIFETEST use.
  • Median survival by the standard reverse-KM rule, reported as not reached rather than as a number when the curve never crosses 0.5, with its own CI from the same band.
  • An at-risk table under the x-axis. This is not decoration: the right-hand tail of a KM curve is often a handful of patients, and without the numbers below it a step down looks like a finding rather than one person leaving the risk set.
  • A log-rank test across strata, with Holm-adjusted pairwise comparisons.

The k = 5 two-sided gate applies to strata, and it is stricter here than a simple size rule. A stratum with fewer than 5 patients is dropped, and so is one whose events — or whose non-events — number between 1 and 4. A dropped stratum is not plotted, not tested and not counted; a warning names it instead. The reason is that a KM curve is not a summary statistic: it draws every event as a visible step, so a curve built on two events publishes those two patients' timings directly onto the screen.

Where the time axis comes from

The de-identified plaintext extract has no clinical clock of its own: surgery_date is absent from every record by design. The events were always present — vital status and recurrence are recorded — but the durations were not, and a Kaplan-Meier estimator without a duration is impossible.

The durations are now derived by migration/scripts/backfill_survival_intervals.py from the encrypted date fields — date of surgery, last known date of follow-up, date of recurrence and recorded date of death — and what is written back to the plaintext record is only the interval in months: followup_months, recurrence_free_months and a death_event flag. No date is ever written. A duration in months is not one of the 18 HIPAA identifiers, whereas a date is, so deriving the interval and discarding the endpoints is what lets survival analysis proceed without widening the de-identified mirror. The script is operator-gated on the encryption key and runs as a dry-run by default.

Coverage is partial, and the gaps are handled rather than filled. Three cases arise, and the rule for each is what matters:

  • A record with a usable end date carries a follow-up interval and enters the risk set normally.
  • A record with no end date at all — neither a last-follow-up date nor a date of death — yields no interval, and is excluded from the risk set rather than censored at an invented time.
  • A record whose follow-up date falls before the date of surgery is left NULL and reported as inconsistent. It is deliberately not clamped to zero, because a clamped negative interval would enter the curve as an immediate event and bias every estimate downwards.

The dashboard prints the denominator it actually computed over, which is the figure to read — the counts are a property of the data on the day you look, not of the method. A record can fall under more than one exclusion, because the follow-up and recurrence intervals are derived separately, so the reasons are not a partition of the records without an interval.

created_at is never used as a clinical clock

created_at is a data-entry timestamp: when a data manager typed the record in, often months or years after the operation, in an order driven by the migration batch rather than by the patient. Months since created_at would produce a smooth, plausible and entirely fabricated curve. It is not substituted, and where intervals are genuinely missing the endpoint returns an explanatory empty state instead of a curve. The period selector on this page filters which records are in the cohort using created_at; it never moves the survival clock, and the response says so in period_basis.

Small-cell suppression (k = 5)

Aggregates are anonymised with a threshold of k = 5. Any cohort, cell or denominator with fewer than five records is returned as null together with a is_suppressed: true flag; the count itself never leaves the database. On the administrative dashboard's enrolment and completion counts the rule is pushed down into the SQL as a CASE expression, so PostgreSQL nulls the value before it reaches the API layer. Everywhere else the same threshold is applied in the service layer, on the way out.

Suppression rule

masked = CASE WHEN count < 5 THEN NULL ELSE count END is_suppressed = (count < 5)

Boundary: a count of exactly 5 is NOT suppressed — the test is strictly less-than.

For a rate, both sides have to clear the threshold. Publishing a rate hands back both cells of a binary outcome: from 18 out of 20 a reader recovers the 2, so 18/20 is exactly as disclosive as 2/20. The gate is therefore applied to the events and to their complement, and a rate is published only when both clear five. The honest extremes survive — none out of twenty and twenty out of twenty are both publishable, because their complements are 20 and 0. What this withholds is any rate whose minority side is a cell of one to four, which includes every non-trivial rate between n = 5 and n = 9, where no split can put five on both sides. At n = 6, a rate of 83.3 % names a patient.

The same logic withholds one more cell than you might expect in a breakdown. When every category but one has been published, the last one is recoverable by subtraction from the total, so the smallest published category is withheld alongside it.

How it appears in the interface:

  • A tile or chart series shows a dash or “suppressed” instead of a number, and keeps showing the denominator n so you can see how far off the threshold you are.
  • In a category breakdown, the suppressed categories are listed by name with a null count, rather than being silently dropped — so the chart does not appear to claim the category has no cases.
  • In dimension pickers, values with fewer than five records are simply not offered as filter options, and institutions below the threshold are listed with a null patient count.
  • In a cohort comparison, a cohort under the threshold is null for every metric and is excluded from the statistical tests; a warning naming the cohort and the metric is included in the response.

Note

Suppression is a privacy control, not a quality signal. A suppressed cell means “fewer than five records”, which may be one or four — the platform will not tell you which, and the CSV export of the dashboard carries the same nulls as the screen.

Statistics used

The cohort comparison workbench (study administrators) reports, for every metric, a per-cohort estimate with a 95 % confidence interval, one omnibus test across all cohorts, and pairwise comparisons with an effect size. Which test runs is decided by the shape of the metric, not chosen by the user. A cohort with fewer than five records for a metric is excluded first; tests run only when at least two cohorts remain.

Proportions — Wilson score interval

Every rate metric (mortality, reoperation, recurrence, severe complication, R0 resection, lymphovascular and perineural invasion) is a proportion over the records where the underlying field is known. Its interval is the Wilson score interval, chosen because it stays inside 0–100 % and behaves sensibly for small numerators, where the textbook normal-approximation interval does not.

Wilson score interval (95 %)

centre = (p + z²/2n) / (1 + z²/n) half = (z / (1 + z²/n)) · √( p(1−p)/n + z²/4n² ) CI = [ centre − half , centre + half ] clamped to [0, 1]

p = events / n, z = 1.96 for 95 %. With n = 0 there is no interval at all — the API returns null rather than [0, 0].

Comparing two means — Welch's t-test

For continuous metrics (nodes harvested, positive nodes, node ratio, complication count) a pair of cohorts is compared with Welch's unequal-variance t-test. It does not assume the two cohorts have the same spread or the same size, which two registry cohorts rarely do. The p-value is the probability of seeing a difference in means at least this large if the two populations had the same mean. The reported interval is the t-based 95 % interval on the difference of means, using the Welch-Satterthwaite degrees of freedom; if it excludes zero, the difference is significant at the 5 % level.

Welch's t and its CI

t = (x̄a − x̄b) / √( sa²/na + sb²/nb ) CI = (x̄a − x̄b) ± t(crit, ν) · √( sa²/na + sb²/nb ) ν = (sa²/na + sb²/nb)² / [ (sa²/na)²/(na−1) + (sb²/nb)²/(nb−1) ]

x̄ = sample mean, s² = sample variance (n−1 denominator), ν = Welch-Satterthwaite degrees of freedom.

Comparing two distributions — Mann-Whitney U

Alongside the t-test, each pair also gets a two-sided Mann-Whitney U rank-sum test. It compares the ranks rather than the means, so it is not thrown by the long right tail typical of node counts. Read it as: how likely is it that a randomly chosen record from one cohort exceeds one from the other? When the two tests disagree, the distributions differ in shape and the mean is a poor summary — trust the medians and the box plot.

Three or more cohorts — Kruskal-Wallis

With three or more cohorts the omnibus test for a continuous metric is Kruskal-Wallis, the rank-based generalisation of Mann-Whitney. A small p-value says “at least one of these cohorts differs” — it does not say which, which is what the pairwise table is for. With exactly two cohorts the omnibus slot reports the Mann-Whitney result instead.

Categories and rates — chi-square, or Fisher when the cells are thin

Categorical metrics (surgical approach, tumour location, histology, vital status, recurrence type, discharge destination, complication grade) and the events/non-events table behind each rate are tested for homogeneity with a chi-square test without Yates correction, reported with its degrees of freedom. For a two-by-two table — two cohorts, one rate — where any expected cell falls below five, the platform switches to Fisher's exact test, which computes the probability directly rather than relying on the chi-square approximation.

Effect sizes: risk difference, odds ratio, Cohen's d

A p-value tells you whether a difference is distinguishable from noise; the effect size tells you how big it is. For rates, each pair reports an absolute risk difference with a Wald interval and an odds ratio with a Woolf (log) interval. When any cell of the table is zero, a Haldane-Anscombe correction of 0.5 is added to every cell so that the odds ratio and its logarithm stay finite; the response flags that this happened, and such an odds ratio should be read as indicative only.

Risk difference and odds ratio

RD = pa − pb RD CI = RD ± z · √( pa(1−pa)/na + pb(1−pb)/nb ) OR = (a · d) / (b · c) OR CI = exp( ln OR ± z · √(1/a + 1/b + 1/c + 1/d) ) ← Woolf

a, b = events and non-events in cohort A; c, d = the same in cohort B; z = 1.96.

For continuous metrics the effect size is Cohen's d, the difference in means expressed in pooled standard deviations. As a rough reading, 0.2 is small, 0.5 moderate and 0.8 large — but in an unadjusted registry comparison, a large d is a prompt to look for case mix, not a conclusion.

Cohen's d (pooled SD)

d = (x̄a − x̄b) / sp sp = √( [ (na−1)·sa² + (nb−1)·sb² ] / (na + nb − 2) )

Many comparisons — Holm adjustment

Six cohorts produce fifteen pairwise comparisons per metric, and at a 5 % threshold roughly one in twenty will look significant by chance alone. Every pairwise p-value is therefore reported twice: raw, and adjusted by the Holm-Bonferroni step-down procedure. Holm is uniformly more powerful than plain Bonferroni and makes no assumption about how the tests relate to each other. Compare the adjusted value against your threshold; use the raw value only to understand what the adjustment cost you.

Holm-Bonferroni step-down

sort p(1) ≤ p(2) ≤ … ≤ p(m) p_holm(i) = max over j ≤ i of min( 1 , (m − j + 1) · p(j) )

m = the number of comparisons in the family. The running maximum keeps the adjusted values monotone; pairs whose test could not run are excluded from m.

Funnel plot limits — exact binomial

The funnel plot places each institution's rate against its volume, inside control limits drawn around the pooled rate. The limits are exact binomial (Clopper-Pearson) rather than normal approximations, evaluated at the expected number of events for each volume so the curves are smooth and always bracket the pooled rate. Two pairs are drawn: 95 % (roughly two standard deviations) and 99.8 % (roughly three). Institutions below the minimum volume are excluded and counted separately rather than plotted with meaningless limits.

Exact binomial control limits at volume n

k = n · p̄ lower = Beta⁻¹( α/2 ; k , n − k + 1 ) upper = Beta⁻¹( 1 − α/2 ; k + 1 , n − k ) z = ( r − p̄ ) / √( p̄(1−p̄)/n )

p̄ = pooled rate across included institutions, r = the institution's own rate, α = 0.05 for the 95 % pair and 0.002 for the 99.8 % pair.

Reporting conventions

P-values are rounded to four significant digits and percentages to two decimals. Confidence intervals are 95 % throughout: Wilson for proportions, t-based for means, Woolf on the log scale for odds ratios. A comparison is capped at six cohorts and twenty metrics per request.

Reading the charts

Box plot — distribution of a continuous metric

The box spans the interquartile range (25th to 75th percentile) with the median as the line inside it; the whiskers reach the minimum and maximum of the cohort. Compare medians and box widths, not just the means printed beside them: a wide box with a short one next to it means the two cohorts differ in consistency, which no single test statistic conveys. Node counts are typically right-skewed, so the median usually sits left of the mean.

Box plot: interquartile box, median line, whiskers to the extremes. Two cohorts side by side.

Grouped bars with confidence intervals — rates side by side

Each bar is a cohort's rate for one metric and the vertical whisker is its 95 % Wilson interval. The interval is the honest part of the chart: two bars of visibly different height whose intervals overlap substantially are not distinguishable at this sample size. Bars carry their denominator, so a tall bar over n = 6 is immediately recognisable as fragile.

Grouped bars with 95 % confidence interval whiskers. Height is the rate; the whisker is the uncertainty.

Forest plot — odds ratios across comparisons

One row per pairwise comparison: a marker at the odds ratio and a horizontal line for its 95 % interval, on a logarithmic axis with a reference line at 1. An interval crossing 1 means no detectable difference. Because the axis is logarithmic, an odds ratio of 2 and one of 0.5 sit symmetrically about the line — which is the point of plotting it this way.

Forest plot: markers are odds ratios, lines are 95 % intervals, the dashed line is OR = 1 (no difference).

Funnel plot — institution rates against volume

Each dot is one institution: volume on the horizontal axis, rate on the vertical, with the pooled rate as a horizontal line and two pairs of curved control limits (95 % inner, 99.8 % outer) that narrow as volume grows. Small institutions are expected to scatter widely — that is what the funnel shape encodes. A point outside the 99.8 % limits warrants a look at data completeness and case mix before anything else; it is a screening device, not a verdict.

Funnel plot: dots are institutions, the horizontal line is the pooled rate, the curves are the 95 % and 99.8 % exact binomial limits.

Heatmap — category mix across cohorts

A grid of cohorts against categories, shaded by share within the cohort. It answers “is the mix different?” at a glance, which the accompanying chi-square test then quantifies. Cells below the suppression threshold are drawn as empty rather than as zero, so a blank cell means “too few to report”, not “none”.

Heatmap: rows are cohorts, columns are categories, shade is the share within the row. Blank = suppressed (fewer than five).

Limitations

Everything on these dashboards is descriptive. Please read it with four caveats in mind.

  • Observational, not randomised. Cohorts are defined by what was recorded, not by allocation. Differences between them reflect referral patterns, national practice and local protocols as much as anything the surgery did.
  • Unadjusted case mix. No comparison here is risk-adjusted. Age, performance status, comorbidity, tumour stage and neoadjuvant treatment are all captured in the eCRF but none of them enter these calculations. An institution operating on sicker or more advanced patients will look worse on raw outcome rates, and the statistics cannot tell you that is what happened.
  • Missingness is not random. Every rate is calculated over the records where the field is known, and how much is known varies enormously by field and by site. Two institutions with the same true outcome can show very different rates if one records follow-up thoroughly and the other does not. Always read the n beside a figure.
  • Time axis is data entry. Because operation dates are not held in plaintext, every trend is indexed on when the record was entered. A spike in a month usually means a site did a batch import, not that it operated more. The one exception is the survival page, which runs on durations derived from the encrypted dates — and even there the period selector filters the cohort on created_at, never the clock.
  • A corrected rate is still a floor. Fixing a denominator does not fix a numerator. The complication figure now divides by the patients we can prove were assessed, but the source only ever captured a fixed list of major complications, so lower-grade morbidity outside that list is invisible to it.

Not for public reporting

These figures are for internal data-quality review and hypothesis generation within the study. They are not risk-adjusted outcome measures and should not be published, shared as institutional performance data, or used in any comparison outside the study without the analysis committee.

Export

The analytics export endpoint produces a CSV of the figures currently in scope. You choose the period and which blocks to include — summary, volume, complications, mortality. The CSV is generated server-side and streamed back; requesting the PDF format returns a queued acknowledgement rather than a file. Every export is written to the audit log with your identity, role, IP address and the parameters you used.

httpAnalytics CSV export (bearer token required)
POST /api/v1/analytics/export
X-Study-ID: tiger
Authorization: Bearer <access_token>
Content-Type: application/json

{
  "format": "csv",
  "metrics": ["summary", "complications", "mortality"],
  "period": "1y"
}

The export is scoped exactly like the dashboard: institution users get their own institution, and a study administrator gets global figures unless they name an institution. Each rate column is followed by its denominator column, and a rate that is missing or suppressed is written as an empty cell, never as 0 — so a spreadsheet built on this file cannot accidentally average a fabricated zero. The complications block exports only the two grade buckets the source records, and its headline rate uses the same assessed denominator as the dashboard — alongside the reoperation rate and the conditional severe share, each with its own denominator column. Requests to this endpoint are rate-limited to 10 per minute.

The separate administrative record-level CSV export (study administrators, from the admin analytics query builder) is streamed rather than buffered and appends a watermark footer recording the exporting user, the masquerade target if any, the active institution, the timestamp, a hash of the query, the row count and that k = 5 was in force.

To export patient-level records rather than aggregates, see Data export.