Documentation
Guide
Analytics
What each figure on the analytics dashboard means, which records are behind it, what gets hidden to protect small groups, and which statistical tests are used when cohorts are compared.
Scope & access
Every analytics response is scoped on the server from your account, never from a parameter you send. There are two tiers:
- Institution users — institution administrators and all clinical roles, including coordinators, fellows and trainees. You see your institution against the global benchmark, and nothing else. No other institution, country or region breakdown is returned to you. Passing an
institution_idthat is not your own is rejected with a 403, and an account with no institution on its record is refused rather than silently given global figures. - Study administrators (
study_adminandproject_admin) — cross-institution scope. Omitting the institution parameter returns global aggregates; supplying one scopes to that institution. This tier also has the Cohort Comparison workbench in the sidebar, which builds arbitrary cohorts from region, country, institution and clinical dimensions and compares them with the tests described below.
Masquerade. When an administrator is acting on behalf of an institution, analytics narrow to that institution for the whole session — the masquerade target wins over the administrator's own home institution. The cohort workbench is deliberately not available while masquerading: those endpoints answer 403 "Exit masquerade to compare cohorts". The rule is that a masquerading admin sees exactly what the institution sees.
All analytics read the plaintext aggregate columns of the patient record. They never decrypt an encrypted section, and they always exclude soft-deleted records. Batches uploaded as a test batch are flagged on the record and excluded from the dashboard aggregations, the administrative analytics and every cohort-comparison query — so a validation run of a spreadsheet does not move anyone's numbers.
Metric definitions
Many fields were never captured for many records, so every rate is calculated over the records that actually carry the field, and the API returns that denominator alongside the value (the n you see next to a figure). A metric with no denominator is returned as no value — never as 0 %.
| Metric | Definition | Numerator / denominator | Source field | Caveats |
|---|---|---|---|---|
| Total patients | Records in scope for the selected period. | Count of non-deleted records / — | patients rows | The period filter applies to created_at, i.e. when the record was entered. |
| Completion rate | Mean record completeness across the cohort. | Mean of completion_pct / records in scope | completion_pct | A record's own completion is completed sections out of a fixed denominator of eight canonical sections. |
| Mean LN yield | Average number of lymph nodes harvested. | Mean of total_ln_harvested / records where that column is not null | total_ln_harvested | Records without a node count are excluded from the mean, not counted as zero. |
| LN distribution | Mean, median, 25th and 75th percentile, plus a histogram in bins of five (0–4, 5–9, … 45–49, 50+). | Percentiles over total_ln_harvested / records with a node count | total_ln_harvested | Percentiles are computed in the database with percentile_cont. |
| Complication rate (assessed) | Share of assessed patients with any recorded complication. | Records with complication_count > 0 or a non-empty grade / records where complication_assessed is true | complication_assessed, complication_count, complication_grade | Where an institution has no assessed patients the API returns overall_status: "not_computable" with no number: that site is left out of every ranking rather than scored at 0 %. |
| Reoperation rate | Return to theatre. This is the headline morbidity figure. | reoperation true / reoperation not null | reoperation | A genuine population rate: the eCRF asks this as a Yes/No question, so a “No” is a recorded answer rather than an unticked box, and a patient nobody answered for drops out of the denominator instead of counting as a success. |
| Severe (≥ IIIa) share of graded patients | Of the patients who were given a Clavien-Dindo grade, how many were graded ≥ III. | Grade >=III / records carrying any grade | complication_grade | Conditional, not a population rate. The denominator is “patients who were graded”, which is itself a selected group — a grade tends to be filled in when something happened. It answers “when a complication was graded, how often was it severe?”, and it must never be read as the share of ALL patients who had a severe complication. The API ships the full label with the number for this reason. |
| Complications by grade | Split into two buckets only: Clavien-Dindo I–II and ≥ III. | Records with that grade / records carrying any grade | complication_grade | The source form records the coarse split, not IIIa/IIIb/IVa/IVb/V. Reporting finer buckets would invent precision the data does not have. |
| 30-day mortality | Death within 30 days of the index operation. | Records flagged true / records where the flag is not null | thirty_day_mortality | Records with an unknown status are excluded from the denominator. A cohort where nobody has a known status returns no value, not 0 %. |
| 90-day mortality | Death within 90 days of the index operation. | Records flagged true / records with a known status | ninety_day_mortality | Same known-status rule as 30-day mortality. |
| In-hospital mortality | Death before discharge. | Records flagged true / records with a known status | in_hospital_mortality | Same known-status rule. |
| R0 resection rate | Microscopically complete resection. | Records with margin R0 / records with any recorded margin | resection_margin | Feeds the R0 axis of the quality profile. |
| Volume trend | Records entered per month (or quarter/year), for your institution against the average institution. | Count of records grouped by date_trunc(created_at) / — ; the global line divides each bucket's total by the number of institutions that contributed to it | created_at | This is a data-entry curve, not a surgical activity curve. The response labels its axis basis: "created_at" for exactly this reason. |
| Quality profile | Six normalised axes for the radar chart: volume, LN yield, R0 rate, low reoperation, low mortality, completion. | Each axis is rescaled to 0–100 (see below) | Derived from the metrics above | The scaling is presentational: volume is scored against twice the global average per institution, LN yield against 40 nodes, and the two “low” axes are inverted so that further out is better. Do not read an axis as a clinical rate. The morbidity axis scores the reoperation rate. |
| Completion by section / by user | Share of records with each section marked complete, and per-user data-entry counts with their mean completeness. | Records with the section flagged complete / records in scope | completion_status, entered_by | Institution-scoped. It is a data-management view, not a performance measure. |
Metrics this registry cannot report
Several familiar surgical metrics have a column in the schema but no data behind it in the migrated registry — length of stay, ICU days, 30-day readmission, blood transfusion, and the clinical and pathological stage groups. Surgery dates are likewise not held in plaintext: dates are de-identified by design, which is why the volume trend uses the record-entry timestamp rather than the operation date, and why length-of-stay analyses are absent rather than shown as zero. If a metric is missing from your dashboard it is because the registry has nothing to compute it from.
Field-wise profile
/dashboard/analytics/profile compares your institution against the whole registry on every plaintext eCRF field with data, one row per field, grouped by eCRF section. The catalogue holds 54 fields and is derived from the patient model rather than hand-maintained, so a field cannot silently drop out when the schema moves. A field with no values anywhere in the registry is listed under unavailable_fields with the reason, never drawn as an empty chart.
Each row carries:
- Both figures with their denominators — mean, SD, median and IQR with a 95 % CI for a continuous field; events over n with a Wilson interval for a rate; the category breakdown for a categorical one.
- The right test for the field's kind — Mann-Whitney U for continuous, Fisher / chi-square for rates and categories — plus an effect size with its own interval. A test is withheld whenever either side of it is withheld, because a published statistic against a published denominator reconstructs the hidden cell.
- The chart, chosen by the API and not by the page. The response names one of
box_with_points,histogram_overlay,rate_bar_with_ci,stacked_100_bar,grouped_bar,line_trendornone, and ships the data that chart needs in the same object — a histogram arrives with its bins already k = 5 masked. The same decision therefore holds in the CSV export and in any other consumer, so two views of one field cannot disagree about how to draw it. - A trend on
created_at, yearly or quarterly, drawn only once at least three periods survive suppression.
Read the Holm column, not the raw p
This page runs one test per field, so a single page view is a few dozen simultaneous comparisons rather than one. At the conventional 0.05 threshold, roughly one in twenty tests comes back “significant” when nothing is going on — scan thirty comparisons looking for the interesting ones and you should expect to find one or two purely by chance. Reporting only the raw p-value would therefore hand every institution a finding.
The response reports p_value and p_value_holm side by side. The Holm-Bonferroni adjustment is computed across the whole family of tests in that response, and the family size is stated on the wire as tests_run, so the correction always matches the table actually in front of you. When you are scanning the page for something worth investigating, the Holm column is the one to read; the raw p is there for a hypothesis you specified before you opened the page.
Survival analysis
/dashboard/analytics/survival plots Kaplan-Meier curves for two endpoints — overall survival (event: vital status is dead) and recurrence-free survival (event: recurrence recorded) — optionally stratified by one of nine variables: pathologic T, N and stage group, resection margin (R0 against R1/R2 pooled), tumour histology, neoadjuvant therapy, lymphadenectomy extent, and two derived lymph-node bands (yield below or at least 15 nodes, and ratio band). The bands are computed in the query and never stored.
Each curve is reported with:
- Greenwood confidence bands, log-log transformed rather than plain linear. A linear band on a survival probability runs outside [0, 1] near the ends of the curve and misstates its own coverage; the log-log band cannot. This is the same default R's
survfit(conf.type = "log-log")and SASPROC LIFETESTuse. - Median survival by the standard reverse-KM rule, reported as not reached rather than as a number when the curve never crosses 0.5, with its own CI from the same band.
- An at-risk table under the x-axis. This is not decoration: the right-hand tail of a KM curve is often a handful of patients, and without the numbers below it a step down looks like a finding rather than one person leaving the risk set.
- A log-rank test across strata, with Holm-adjusted pairwise comparisons.
The k = 5 two-sided gate applies to strata, and it is stricter here than a simple size rule. A stratum with fewer than 5 patients is dropped, and so is one whose events — or whose non-events — number between 1 and 4. A dropped stratum is not plotted, not tested and not counted; a warning names it instead. The reason is that a KM curve is not a summary statistic: it draws every event as a visible step, so a curve built on two events publishes those two patients' timings directly onto the screen.
Where the time axis comes from
The de-identified plaintext extract has no clinical clock of its own: surgery_date is absent from every record by design. The events were always present — vital status and recurrence are recorded — but the durations were not, and a Kaplan-Meier estimator without a duration is impossible.
The durations are now derived by migration/scripts/backfill_survival_intervals.py from the encrypted date fields — date of surgery, last known date of follow-up, date of recurrence and recorded date of death — and what is written back to the plaintext record is only the interval in months: followup_months, recurrence_free_months and a death_event flag. No date is ever written. A duration in months is not one of the 18 HIPAA identifiers, whereas a date is, so deriving the interval and discarding the endpoints is what lets survival analysis proceed without widening the de-identified mirror. The script is operator-gated on the encryption key and runs as a dry-run by default.
Coverage is partial, and the gaps are handled rather than filled. Three cases arise, and the rule for each is what matters:
- A record with a usable end date carries a follow-up interval and enters the risk set normally.
- A record with no end date at all — neither a last-follow-up date nor a date of death — yields no interval, and is excluded from the risk set rather than censored at an invented time.
- A record whose follow-up date falls before the date of surgery is left NULL and reported as inconsistent. It is deliberately not clamped to zero, because a clamped negative interval would enter the curve as an immediate event and bias every estimate downwards.
The dashboard prints the denominator it actually computed over, which is the figure to read — the counts are a property of the data on the day you look, not of the method. A record can fall under more than one exclusion, because the follow-up and recurrence intervals are derived separately, so the reasons are not a partition of the records without an interval.
created_at is never used as a clinical clock
created_at is a data-entry timestamp: when a data manager typed the record in, often months or years after the operation, in an order driven by the migration batch rather than by the patient. Months since created_at would produce a smooth, plausible and entirely fabricated curve. It is not substituted, and where intervals are genuinely missing the endpoint returns an explanatory empty state instead of a curve. The period selector on this page filters which records are in the cohort using created_at; it never moves the survival clock, and the response says so in period_basis.
Small-cell suppression (k = 5)
Aggregates are anonymised with a threshold of k = 5. Any cohort, cell or denominator with fewer than five records is returned as null together with a is_suppressed: true flag; the count itself never leaves the database. On the administrative dashboard's enrolment and completion counts the rule is pushed down into the SQL as a CASE expression, so PostgreSQL nulls the value before it reaches the API layer. Everywhere else the same threshold is applied in the service layer, on the way out.
Suppression rule
masked = CASE WHEN count < 5 THEN NULL ELSE count END is_suppressed = (count < 5)
Boundary: a count of exactly 5 is NOT suppressed — the test is strictly less-than.
For a rate, both sides have to clear the threshold. Publishing a rate hands back both cells of a binary outcome: from 18 out of 20 a reader recovers the 2, so 18/20 is exactly as disclosive as 2/20. The gate is therefore applied to the events and to their complement, and a rate is published only when both clear five. The honest extremes survive — none out of twenty and twenty out of twenty are both publishable, because their complements are 20 and 0. What this withholds is any rate whose minority side is a cell of one to four, which includes every non-trivial rate between n = 5 and n = 9, where no split can put five on both sides. At n = 6, a rate of 83.3 % names a patient.
The same logic withholds one more cell than you might expect in a breakdown. When every category but one has been published, the last one is recoverable by subtraction from the total, so the smallest published category is withheld alongside it.
How it appears in the interface:
- A tile or chart series shows a dash or “suppressed” instead of a number, and keeps showing the denominator
nso you can see how far off the threshold you are. - In a category breakdown, the suppressed categories are listed by name with a null count, rather than being silently dropped — so the chart does not appear to claim the category has no cases.
- In dimension pickers, values with fewer than five records are simply not offered as filter options, and institutions below the threshold are listed with a null patient count.
- In a cohort comparison, a cohort under the threshold is null for every metric and is excluded from the statistical tests; a warning naming the cohort and the metric is included in the response.
Note
Suppression is a privacy control, not a quality signal. A suppressed cell means “fewer than five records”, which may be one or four — the platform will not tell you which, and the CSV export of the dashboard carries the same nulls as the screen.
Statistics used
The cohort comparison workbench (study administrators) reports, for every metric, a per-cohort estimate with a 95 % confidence interval, one omnibus test across all cohorts, and pairwise comparisons with an effect size. Which test runs is decided by the shape of the metric, not chosen by the user. A cohort with fewer than five records for a metric is excluded first; tests run only when at least two cohorts remain.
Proportions — Wilson score interval
Every rate metric (mortality, reoperation, recurrence, severe complication, R0 resection, lymphovascular and perineural invasion) is a proportion over the records where the underlying field is known. Its interval is the Wilson score interval, chosen because it stays inside 0–100 % and behaves sensibly for small numerators, where the textbook normal-approximation interval does not.
Wilson score interval (95 %)
centre = (p + z²/2n) / (1 + z²/n) half = (z / (1 + z²/n)) · √( p(1−p)/n + z²/4n² ) CI = [ centre − half , centre + half ] clamped to [0, 1]
p = events / n, z = 1.96 for 95 %. With n = 0 there is no interval at all — the API returns null rather than [0, 0].
Comparing two means — Welch's t-test
For continuous metrics (nodes harvested, positive nodes, node ratio, complication count) a pair of cohorts is compared with Welch's unequal-variance t-test. It does not assume the two cohorts have the same spread or the same size, which two registry cohorts rarely do. The p-value is the probability of seeing a difference in means at least this large if the two populations had the same mean. The reported interval is the t-based 95 % interval on the difference of means, using the Welch-Satterthwaite degrees of freedom; if it excludes zero, the difference is significant at the 5 % level.
Welch's t and its CI
t = (x̄a − x̄b) / √( sa²/na + sb²/nb ) CI = (x̄a − x̄b) ± t(crit, ν) · √( sa²/na + sb²/nb ) ν = (sa²/na + sb²/nb)² / [ (sa²/na)²/(na−1) + (sb²/nb)²/(nb−1) ]
x̄ = sample mean, s² = sample variance (n−1 denominator), ν = Welch-Satterthwaite degrees of freedom.
Comparing two distributions — Mann-Whitney U
Alongside the t-test, each pair also gets a two-sided Mann-Whitney U rank-sum test. It compares the ranks rather than the means, so it is not thrown by the long right tail typical of node counts. Read it as: how likely is it that a randomly chosen record from one cohort exceeds one from the other? When the two tests disagree, the distributions differ in shape and the mean is a poor summary — trust the medians and the box plot.
Three or more cohorts — Kruskal-Wallis
With three or more cohorts the omnibus test for a continuous metric is Kruskal-Wallis, the rank-based generalisation of Mann-Whitney. A small p-value says “at least one of these cohorts differs” — it does not say which, which is what the pairwise table is for. With exactly two cohorts the omnibus slot reports the Mann-Whitney result instead.
Categories and rates — chi-square, or Fisher when the cells are thin
Categorical metrics (surgical approach, tumour location, histology, vital status, recurrence type, discharge destination, complication grade) and the events/non-events table behind each rate are tested for homogeneity with a chi-square test without Yates correction, reported with its degrees of freedom. For a two-by-two table — two cohorts, one rate — where any expected cell falls below five, the platform switches to Fisher's exact test, which computes the probability directly rather than relying on the chi-square approximation.
Effect sizes: risk difference, odds ratio, Cohen's d
A p-value tells you whether a difference is distinguishable from noise; the effect size tells you how big it is. For rates, each pair reports an absolute risk difference with a Wald interval and an odds ratio with a Woolf (log) interval. When any cell of the table is zero, a Haldane-Anscombe correction of 0.5 is added to every cell so that the odds ratio and its logarithm stay finite; the response flags that this happened, and such an odds ratio should be read as indicative only.
Risk difference and odds ratio
RD = pa − pb RD CI = RD ± z · √( pa(1−pa)/na + pb(1−pb)/nb ) OR = (a · d) / (b · c) OR CI = exp( ln OR ± z · √(1/a + 1/b + 1/c + 1/d) ) ← Woolf
a, b = events and non-events in cohort A; c, d = the same in cohort B; z = 1.96.
For continuous metrics the effect size is Cohen's d, the difference in means expressed in pooled standard deviations. As a rough reading, 0.2 is small, 0.5 moderate and 0.8 large — but in an unadjusted registry comparison, a large d is a prompt to look for case mix, not a conclusion.
Cohen's d (pooled SD)
d = (x̄a − x̄b) / sp sp = √( [ (na−1)·sa² + (nb−1)·sb² ] / (na + nb − 2) )
Many comparisons — Holm adjustment
Six cohorts produce fifteen pairwise comparisons per metric, and at a 5 % threshold roughly one in twenty will look significant by chance alone. Every pairwise p-value is therefore reported twice: raw, and adjusted by the Holm-Bonferroni step-down procedure. Holm is uniformly more powerful than plain Bonferroni and makes no assumption about how the tests relate to each other. Compare the adjusted value against your threshold; use the raw value only to understand what the adjustment cost you.
Holm-Bonferroni step-down
sort p(1) ≤ p(2) ≤ … ≤ p(m) p_holm(i) = max over j ≤ i of min( 1 , (m − j + 1) · p(j) )
m = the number of comparisons in the family. The running maximum keeps the adjusted values monotone; pairs whose test could not run are excluded from m.
Funnel plot limits — exact binomial
The funnel plot places each institution's rate against its volume, inside control limits drawn around the pooled rate. The limits are exact binomial (Clopper-Pearson) rather than normal approximations, evaluated at the expected number of events for each volume so the curves are smooth and always bracket the pooled rate. Two pairs are drawn: 95 % (roughly two standard deviations) and 99.8 % (roughly three). Institutions below the minimum volume are excluded and counted separately rather than plotted with meaningless limits.
Exact binomial control limits at volume n
k = n · p̄ lower = Beta⁻¹( α/2 ; k , n − k + 1 ) upper = Beta⁻¹( 1 − α/2 ; k + 1 , n − k ) z = ( r − p̄ ) / √( p̄(1−p̄)/n )
p̄ = pooled rate across included institutions, r = the institution's own rate, α = 0.05 for the 95 % pair and 0.002 for the 99.8 % pair.
Reporting conventions
P-values are rounded to four significant digits and percentages to two decimals. Confidence intervals are 95 % throughout: Wilson for proportions, t-based for means, Woolf on the log scale for odds ratios. A comparison is capped at six cohorts and twenty metrics per request.
Reading the charts
Box plot — distribution of a continuous metric
The box spans the interquartile range (25th to 75th percentile) with the median as the line inside it; the whiskers reach the minimum and maximum of the cohort. Compare medians and box widths, not just the means printed beside them: a wide box with a short one next to it means the two cohorts differ in consistency, which no single test statistic conveys. Node counts are typically right-skewed, so the median usually sits left of the mean.
Grouped bars with confidence intervals — rates side by side
Each bar is a cohort's rate for one metric and the vertical whisker is its 95 % Wilson interval. The interval is the honest part of the chart: two bars of visibly different height whose intervals overlap substantially are not distinguishable at this sample size. Bars carry their denominator, so a tall bar over n = 6 is immediately recognisable as fragile.
Forest plot — odds ratios across comparisons
One row per pairwise comparison: a marker at the odds ratio and a horizontal line for its 95 % interval, on a logarithmic axis with a reference line at 1. An interval crossing 1 means no detectable difference. Because the axis is logarithmic, an odds ratio of 2 and one of 0.5 sit symmetrically about the line — which is the point of plotting it this way.
Funnel plot — institution rates against volume
Each dot is one institution: volume on the horizontal axis, rate on the vertical, with the pooled rate as a horizontal line and two pairs of curved control limits (95 % inner, 99.8 % outer) that narrow as volume grows. Small institutions are expected to scatter widely — that is what the funnel shape encodes. A point outside the 99.8 % limits warrants a look at data completeness and case mix before anything else; it is a screening device, not a verdict.
Heatmap — category mix across cohorts
A grid of cohorts against categories, shaded by share within the cohort. It answers “is the mix different?” at a glance, which the accompanying chi-square test then quantifies. Cells below the suppression threshold are drawn as empty rather than as zero, so a blank cell means “too few to report”, not “none”.
Limitations
Everything on these dashboards is descriptive. Please read it with four caveats in mind.
- Observational, not randomised. Cohorts are defined by what was recorded, not by allocation. Differences between them reflect referral patterns, national practice and local protocols as much as anything the surgery did.
- Unadjusted case mix. No comparison here is risk-adjusted. Age, performance status, comorbidity, tumour stage and neoadjuvant treatment are all captured in the eCRF but none of them enter these calculations. An institution operating on sicker or more advanced patients will look worse on raw outcome rates, and the statistics cannot tell you that is what happened.
- Missingness is not random. Every rate is calculated over the records where the field is known, and how much is known varies enormously by field and by site. Two institutions with the same true outcome can show very different rates if one records follow-up thoroughly and the other does not. Always read the
nbeside a figure. - Time axis is data entry. Because operation dates are not held in plaintext, every trend is indexed on when the record was entered. A spike in a month usually means a site did a batch import, not that it operated more. The one exception is the survival page, which runs on durations derived from the encrypted dates — and even there the period selector filters the cohort on
created_at, never the clock. - A corrected rate is still a floor. Fixing a denominator does not fix a numerator. The complication figure now divides by the patients we can prove were assessed, but the source only ever captured a fixed list of major complications, so lower-grade morbidity outside that list is invisible to it.
Not for public reporting
These figures are for internal data-quality review and hypothesis generation within the study. They are not risk-adjusted outcome measures and should not be published, shared as institutional performance data, or used in any comparison outside the study without the analysis committee.
Export
The analytics export endpoint produces a CSV of the figures currently in scope. You choose the period and which blocks to include — summary, volume, complications, mortality. The CSV is generated server-side and streamed back; requesting the PDF format returns a queued acknowledgement rather than a file. Every export is written to the audit log with your identity, role, IP address and the parameters you used.
POST /api/v1/analytics/export
X-Study-ID: tiger
Authorization: Bearer <access_token>
Content-Type: application/json
{
"format": "csv",
"metrics": ["summary", "complications", "mortality"],
"period": "1y"
}The export is scoped exactly like the dashboard: institution users get their own institution, and a study administrator gets global figures unless they name an institution. Each rate column is followed by its denominator column, and a rate that is missing or suppressed is written as an empty cell, never as 0 — so a spreadsheet built on this file cannot accidentally average a fabricated zero. The complications block exports only the two grade buckets the source records, and its headline rate uses the same assessed denominator as the dashboard — alongside the reoperation rate and the conditional severe share, each with its own denominator column. Requests to this endpoint are rate-limited to 10 per minute.
The separate administrative record-level CSV export (study administrators, from the admin analytics query builder) is streamed rather than buffered and appends a watermark footer recording the exporting user, the masquerade target if any, the active institution, the timestamp, a hash of the query, the row count and that k = 5 was in force.
To export patient-level records rather than aggregates, see Data export.