One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

A Strong Linear Baseline for Whole-Heart Cardiac Shape Completion on CT, with an Open Eleven-Structure Statistical Shape Model

arXiv:2608.199322026-08-21

For filling in missing heart structures on CT scans, a simple math formula beat a graph deep-learning model

Different public heart CT datasets label different parts of the heart, so combining them is hard because structures like the atrial appendage or pulmonary veins are often missing as separate pieces. The authors built and released a statistical shape model covering eleven heart structures from 383 CT cases, all matched to the same 11,571 mesh points, and tested several methods for predicting missing structures from the ones that are present. A closed-form conditional-Gaussian formula outperformed a graph-based deep learning model on this completion task.

What they did

  1. No existing public cardiac shape resource kept the atrial appendage, pulmonary veins, and vena cava stumps as separate full surface pieces, so the authors built and released a new eleven-structure shape model from 383 CT cases with 11,571 corresponding mesh points
  2. They compared a closed-form conditional-Gaussian estimator, a mask-conditioned graph variational autoencoder (graph beta-VAE), and nearest-neighbour retrieval, all under one frozen 76-case internal test split and identical evaluation rules
  3. The conditional-Gaussian estimator reached 3.717 mm average per-vertex error, beating the graph beta-VAE (5.248 mm) and nearest-neighbour retrieval (8.931 mm), with a 1.531 mm gap (95% confidence interval 1.384 to 1.711 mm) over the beta-VAE
  4. The same ranking held on two external test sets with expert-labeled structures (58 cases from CARE2026 and 20 cases from MM-WHS), though four structures like the aorta had to be excluded because the reference data wasn't accurate enough there
  5. The authors state clearly that the released model and method are research tools for unifying misaligned cardiac datasets, not validated for clinical use
Fig. 1: End-to-end pipeline and frozen evaluation benchmark. After segmentation QC, joint-label registration, registration/FOV QC, and same-structure R-local repair, the 383 corresponding meshes are divided into a development set used for model selection and a 76-case evaluation set withheld from every fit reported here (the graph architecture predates this partition; Section V). The orange block carries the conditional-mean equation; all comparators use the same split, panels, endpoint, and evaluation cases. CARE2026 and MM-WHS provide separate expert-surface validation without fitting or tuning.
Fig. 1: End-to-end pipeline and frozen evaluation benchmark. After segmentation QC, joint-label registration, registration/FOV QC, and same-structure R-local repair, the 383 corresponding meshes are divided into a development set used for model selection and a 76-case evaluation set withheld from every fit reported here (the graph architecture predates this partition; Section V). The orange block carries the conditional-mean equation; all comparators use the same split, panels, endpoint, and evaluation cases. CARE2026 and MM-WHS provide separate expert-surface validation without fitting or tuning.
TABLE I: Closest multistructure CT geometry resources. “Public package” means a direct public download of a reusable model or mesh cohort, assessed from the cited publications and their linked repositories at the time of writing; boundary tags or cut rings are not counted as separate anatomical surface bodies. None of the prior resources supports missing-label conditioning; this work conditions on any observed structure-block subset.
ResourceAnatomy outside the four chamber wallsPublic package
Hoogendoorn et al. [20]Detailed multiregion whole-heart atlas, including great vesselsAtlas and SSM reported; no reusable model package located
Rodero et al. [2]AO and PA walls; PV and caval openings tagged; LAA body omittedSSM parameters and 1 000 synthetic finite-element meshes
Strocchi et al. [3]AO and PA walls; LAA, PV, SVC and IVC at cut rings24 patient-specific finite-element meshes
This workAO, PA, LAA, PV, SVC and IVC as separate surface blocks in one shared topologyTemplate, per-case displacement fields and reconstruction code
Fig. 2: The eleven-structure whole-heart template mesh (11 571 vertices) in anterior, posterior, left, and right views, each structure a colour-coded surface in fixed vertex correspondence. This single template is warped into every patient’s space to build the SSM.
Fig. 2: The eleven-structure whole-heart template mesh (11 571 vertices) in anterior, posterior, left, and right views, each structure a colour-coded surface in fixed vertex correspondence. This single template is warped into every patient’s space to build the SSM.
TABLE II: The three cohorts and the exclusions that produce them. The released internal metadata supplies age, sex, manufacturer, scanner model, tube voltage, pathology class, and study type; “n/a” attributes are absent from the released metadata and are not estimated, and “none” means no fitting or tuning. The 76-case evaluation list is frozen (split file released with the artifacts). Both CARE2026 exclusions fall in subset B; MM-WHS has no exclusions.
Internal (TotalSegmentator)CARE2026MM-WHS CT
RoleSSM fit; internal eval.indep. validationindep. secondary
Source cases63160 (A/B/G)20
Segmentation QC−206 (vol., extent, centr.) → 425
Affine check−2 (non-orthon.)
Registration QC−26 (>25 mm) → 399
FOV truncation−16
Retained38358 (20/18/20)20
Split/tuning307 dev. (5 folds) / 76 withheld†nonenone
Labelssilver, all 11silver 11, expert 7expert 7
Scanner models11 distinctn/an/a
Median age64 yearsn/an/a
Study mix336 non-cardiac, 47 cardiacn/an/a
Contrast phasen/an/an/a
ECG gating, phasen/a, likely mixedn/an/a
Spacing analysed1 mm isotropic1 mm isotropicper case → 1 mm
†The graph architecture predates the partition and may have seen
cases now in the 76-case evaluation list.
Fig. 3: The completion operator. A selection matrix 𝐏o extracts whichever structure blocks or partial vertex sets are observed. The covariance is fitted once; inference extracts the corresponding covariance blocks and evaluates the conditional mean. Observed coordinates are copied to the output; missing coordinates are predicted. The figure writes the operator in covariance form for readability; the implementation evaluates the low-rank coefficient-space posterior of Eq. (1), whose ridge acts on the posterior precision rather than on 𝐂o​o.
Fig. 3: The completion operator. A selection matrix 𝐏o extracts whichever structure blocks or partial vertex sets are observed. The covariance is fitted once; inference extracts the corresponding covariance blocks and evaluates the conditional mean. Observed coordinates are copied to the output; missing coordinates are predicted. The figure writes the operator in covariance form for readability; the implementation evaluates the low-rank coefficient-space posterior of Eq. (1), whose ridge acts on the posterior precision rather than on 𝐂o​o.
TABLE III: Eleven cardiac structures and template vertex counts. MYO is the left-ventricular wall; the other ten are blood-pool or luminal surfaces. The model therefore carries no right-ventricular or atrial wall and does not represent total myocardial mass.
IDStructureAbbrev.Vertices
1Left ventricleLV1 502
2MyocardiumMYO2 502
3Right ventricleRV2 002
4Left atriumLA1 002
5Right atriumRA1 002
6AortaAO1 202
7Pulmonary arteryPA802
8Left atrial appendageLAA402
9Pulmonary veinsPV602
10Superior vena cavaSVC402
11Inferior vena cavaIVC151
Total11 571
Fig. 4: Completion across six observed-structure panels, one internal TotalSegmentator case (TS, s0365) beside one external CARE2026 case (A_Case1020), each the median case of its cohort by the displayed Cond-G estimator’s mean missing-vertex error over the six panels. Observed structures are grey; the conditional-Gaussian model completes the rest, coloured by per-vertex error (mm). The estimator is the selected configuration of Table V, fitted on the 307 development cases in the R-local domain.
Fig. 4: Completion across six observed-structure panels, one internal TotalSegmentator case (TS, s0365) beside one external CARE2026 case (A_Case1020), each the median case of its cohort by the displayed Cond-G estimator’s mean missing-vertex error over the six panels. Observed structures are grey; the conditional-Gaussian model completes the rest, coloured by per-vertex error (mm). The estimator is the selected configuration of Table V, fitted on the 307 development cases in the R-local domain.
TABLE IV: Per-structure warp-overlap Dice (SyN-warped patient segmentations vs template), CARE2026 (n=58): a registration-consistency measure, not accuracy against a manual reference. Structures marked † (LAA, PV, SVC, IVC) are silver-only. Chambers is the mean of the five per-chamber medians. p25 is the 25th percentile.
StructureMedianp25StructureMedianp25
LV0.9810.979AO0.8240.787
MYO0.9620.944PA0.9690.910
RV0.9740.965LAA†0.9410.938
LA0.9780.977PV†0.9180.891
RA0.9750.971SVC†0.7250.619
IVC†0.5280.461
Chambers: 0.974All 11: 0.889
TABLE V: Frozen 76-case internal evaluation completion MPVED (mm). Each k entry is the case mean over its fixed panels; the four-k column equally weights the four k values, with a case-bootstrap 95% CI. PPCA is the selected K=1,q=160 limit; local Cond-G reuses the exact global prediction by construction.
RepresentationMethodk=1k=3k=5k=9Four-k [95% CI]
RawMean shape11.21811.27811.18910.90911.149 [10.310, 12.040]
Nearest neighbour10.0969.0318.5088.6449.070 [8.621, 9.563]
Cond-G5.4793.9183.1933.0033.898 [3.629, 4.206]
β-VAE6.5965.2384.6934.7485.319 [4.991, 5.701]
PPCA (K=1)5.5243.9583.2483.0643.948 [3.683, 4.251]
Local Cond-G†5.4793.9183.1933.0033.898 [3.629, 4.206]
R-localMean shape11.15711.21511.13010.85811.090 [10.245, 11.990]
Nearest neighbour9.9428.9048.3528.5278.931 [8.503, 9.405]
Cond-G5.3043.7323.0072.8253.717 [3.480, 3.984]
β-VAE6.5225.1724.6214.6775.248 [4.935, 5.613]
PPCA (K=1)5.3483.7673.0612.8823.764 [3.530, 4.027]
Local Cond-G†5.3043.7323.0072.8253.717 [3.480, 3.984]
TABLE VI: Frozen paired family. Differences are challenger minus Cond-G, so positive values favour Cond-G. The R-local four-k primary is unadjusted; the next seven rows are the complete Holm family (m=7). The final row is an identity check rather than empirical evidence.
ContrastDifference (mm)95% CI (mm)Multiplicity result
R-local β-VAE minus Cond-G, four-k primary1.531[1.384, 1.711]Unadjusted p<0.001 (floor)
Raw β-VAE minus Cond-G, four-k1.421[1.264, 1.610]pHolm<0.002
R-local β-VAE minus Cond-G, k=11.218[1.023, 1.439]pHolm<0.002
R-local β-VAE minus Cond-G, k=31.440[1.298, 1.602]pHolm<0.002
R-local β-VAE minus Cond-G, k=51.613[1.466, 1.800]pHolm<0.002
R-local β-VAE minus Cond-G, k=91.852[1.689, 2.047]pHolm<0.002
R-local PPCA (K=1) minus Cond-G, four-k0.047[0.038, 0.056]pHolm<0.002
R-local local minus global Cond-G, four-k0.000[0.000, 0.000]Structural identity
TABLE VII: Additional non-linear baselines, R-local representation, MPVED (mm) with case-bootstrap 95% intervals. The internal column is the frozen 76-case split at four observed-structure counts; the MM-WHS column is exhaustive panels at k=1,3,5 over the seven expert-overlap structures, so only the ordering within a column is comparable. Both arms lie outside the prespecified Holm family and are descriptive; the latent-optimisation setting is not retuned externally. Cond-G, the β-VAE and the floors repeat Table V for reference; the same arms on CARE2026 expert surfaces, a different endpoint, are in Table VIII.
MethodTypeInternal 76MM-WHS (n=20)
Cond-GClosed-form linear3.717 [3.480, 3.984]4.823 [4.446, 5.226]
Latent optimisation‡Published method5.044 [4.746, 5.389]6.248 [5.755, 6.777]
β-VAEFeed-forward deep5.248 [4.935, 5.613]6.616 [6.109, 7.152]
SpiralNet++§Published operator6.710 [6.283, 7.218]8.916 [8.105, 9.741]
Nearest neighbourRetrieval floor8.931 [8.503, 9.405]10.907 [10.145, 11.684]
Mean shapePopulation floor11.090 [10.245, 11.990]13.174 [12.045, 14.419]
TABLE VIII: Additional non-linear baselines on CARE2026 expert surfaces, R-local, biventricular panel, mean ASSD / HD95 / Chamfer RMS (mm) over 58 cases. The endpoint is distance to expert manual segmentations and is not comparable to Table VII. AO and PA are suppressed for every method by the prespecified reference-closeness rule. Cond-G and the β-VAE repeat Table IX for reference.
MethodTypeLARA
Cond-GClosed-form linear3.275 / 7.846 / 4.1383.479 / 9.007 / 4.586
Latent optimisationPublished method3.902 / 8.941 / 4.8224.243 / 10.265 / 5.413
β-VAEFeed-forward deep3.934 / 8.990 / 4.8604.333 / 10.372 / 5.495
SpiralNet++Published operator4.299 / 9.701 / 5.2664.605 / 10.868 / 5.779
TABLE IX: R-local CARE biventricular native prediction-versus-expert surface metrics, pooled (n=58) and by acquisition subset (A/B/G contain 20/18/20 cases). The panel observes LV/MYO/RV; values are mean ASSD / HD95 / Chamfer RMS in mm. AO and PA are omitted after failing the reference-closeness rule.
PopulationMethodLARA
PooledCond-G3.275 / 7.846 / 4.1383.479 / 9.007 / 4.586
Pooledβ-VAE3.934 / 8.990 / 4.8604.333 / 10.372 / 5.495
ACond-G2.634 / 6.643 / 3.3533.035 / 8.085 / 3.989
Aβ-VAE3.103 / 7.553 / 3.9113.780 / 9.373 / 4.811
BCond-G4.010 / 9.642 / 5.0324.028 / 10.598 / 5.351
Bβ-VAE4.878 / 11.412 / 6.0344.706 / 11.572 / 6.011
GCond-G3.253 / 7.432 / 4.1193.427 / 8.496 / 4.494
Gβ-VAE3.915 / 8.247 / 4.7524.549 / 10.290 / 5.715
TABLE X: R-local CARE leave-one-expert-structure-out completion, pooled (n=58) and by acquisition subset (A/B/G contain 20/18/20 cases). Values are mean ASSD / HD95 / Chamfer RMS in mm. AO and PA are omitted because their references are too far.
Missing structurePopulationCond-Gβ-VAE
LVPooled2.362 / 6.407 / 3.1193.899 / 9.623 / 4.942
A1.818 / 4.744 / 2.3583.367 / 8.177 / 4.254
B3.697 / 10.361 / 4.9694.841 / 12.575 / 6.258
G1.704 / 4.510 / 2.2153.582 / 8.411 / 4.446
MYOPooled2.056 / 5.667 / 2.7292.943 / 7.583 / 3.769
A1.695 / 4.541 / 2.2122.526 / 6.257 / 3.171
B2.947 / 8.627 / 4.0363.847 / 10.332 / 5.046
G1.616 / 4.129 / 2.0682.547 / 6.434 / 3.219
RVPooled2.717 / 8.098 / 3.7703.377 / 8.555 / 4.335
A2.868 / 8.349 / 3.9693.016 / 8.073 / 3.994
B3.247 / 10.314 / 4.6194.148 / 10.495 / 5.305
G2.088 / 5.852 / 2.8073.045 / 7.290 / 3.801
LAPooled2.417 / 6.228 / 3.1923.393 / 7.837 / 4.207
A2.169 / 5.919 / 2.8712.577 / 6.451 / 3.266
B3.023 / 7.642 / 3.9194.445 / 10.407 / 5.495
G2.120 / 5.264 / 2.8603.262 / 6.909 / 3.989
RAPooled2.800 / 7.434 / 3.7633.878 / 9.530 / 4.988
A2.577 / 7.013 / 3.4583.474 / 8.729 / 4.448
B3.246 / 9.006 / 4.4414.372 / 10.762 / 5.607
G2.621 / 6.439 / 3.4563.836 / 9.224 / 4.972
TABLE XI: The evidence behind each structure; structures with identical status share a row. “Reference close enough” means the registered reference lies within the fixed distance limits of the expert surface, so a native comparison is meaningful; the internal benchmark scores only structures that are completion targets there. “Cond-G better” is the lower descriptive mean error against the β-VAE at every k. AO’s reference is too far because the expert annotations cover a shorter aortic extent than the whole-aorta reference, not because of misregistration (Section V-D).
StructuresInternal benchmarkCARE2026 expert checkMM-WHS expert check
LV, MYO, RVnever a completion targetreference close enough (leave-one-out panel)expert label, but observed as input
LA, RAnever a completion targetreference close enough (both panels)reference close enough; Cond-G better
AOCond-G better at every kreference too farreference too far
PACond-G better at every kreference too farreference close enough under R-local; Cond-G better
LAA, PV, SVC, IVCCond-G better at every kno expert labelno expert label
TABLE XII: Morphometric retention from the biventricular panel: per-structure volume agreement between the completion and the case’s own registered mesh. MAE is mean absolute error; medAPE is median absolute percentage error; Δ is the paired difference in absolute error against the development-only regression on the three observed chamber volumes, with its 95% case-bootstrap interval (negative favours completion). Vessel and appendage rows are segment volumes under the fixed template crop.
Internal (n=76)CARE2026 (n=58)
StructureMAE (mL)medAPEΔ [95% CI]MAE (mL), medAPE
LA7.998.3%−6.84 [−9.43, −4.43]18.47, 18.5%
RA9.618.9%−7.29 [−10.36, −4.35]20.89, 25.7%
AO13.815.6%−7.99 [−12.75, −3.33]23.08, 13.1%
PA6.748.5%−4.51 [−6.35, −2.68]13.39, 17.5%
LAA0.9310.7%−0.49 [−0.71, −0.27]1.41, 14.8%
PV1.4113.1%−0.01 [−0.31, +0.28]1.55, 13.7%
SVC3.1812.3%−0.60 [−1.30, +0.21]5.52, 22.6%
TABLE XIII: Nominal 95% ellipsoid coverage by out-of-support stratum, R-local, 76 evaluation cases (64 typical, 12 atypical). Strata come from the frozen development 95th-percentile Mahalanobis threshold. Exploratory and post-hoc; brackets are descriptive case-bootstrap intervals.
Cond-Gβ-VAE
Observed paneltypicalatypicaltypicalatypical
LV1.0000.946 [0.868, 0.998]0.1280.105 [0.071, 0.138]
LV, MYO, RV1.0000.944 [0.878, 0.991]0.1290.140 [0.104, 0.174]
Five chambers1.0000.974 [0.953, 0.993]0.1410.139 [0.103, 0.175]
Nine structures1.0000.999 [0.996, 1.000]0.2050.194 [0.110, 0.281]

Why it matters

For researchers trying to combine heart CT datasets that each label different structures, this shows that a simple, fast statistical formula can outperform a more complex deep learning model, meaning simpler baselines deserve testing before building elaborate architectures. It also highlights that deep learning performance claims depend heavily on what baseline they're compared against.

Terms in this paper

  • Statistical Shape Model (SSM) · a 3D model built by aligning many people's organ shapes to the same reference points, capturing the average shape and how it varies
  • vertex correspondence · matching the same anatomical point across different people's meshes so shapes can be directly compared
  • conditional-Gaussian estimator (Cond-G) · a formula that, assuming a normal distribution, calculates the most likely missing values directly from the observed ones
  • graph beta-VAE · a type of deep learning model that compresses and reconstructs 3D mesh data represented as a graph
  • MPVED (mean per-vertex error distance) · the average distance in millimeters between predicted and true surface points, used to measure accuracy

Original abstract (English)

Public cardiac cohorts annotate different subsets of the heart, so shapes from separate sources cannot be pooled without shared correspondence. Among released cardiac shape resources, none we identified carries the atrial appendage, pulmonary veins, and caval stumps as separate blocks in one mesh. Completion benchmarks also compare deep models against a least-squares projection onto shape modes, not the conditional estimator the same fitted model implies. We release an eleven- structure cardiac computed-tomography (CT) statistical shape model, built from 383 automatically labelled cases in 11 571-vertex correspondence, and compare completion estimators under one frozen internal split and endpoint. On a 76-case internal list held out from fitting, a closed-form conditional-Gaussian estimator reconstructed the missing non-chamber structures at 3.717 mm mean per-vertex error, averaged equally over one, three, five, and nine observed structures. A five-refit mask-conditioned graph variational autoencoder reached 5.248 mm and nearest-neighbour retrieval 8.931 mm. The paired difference was 1.531 mm (95% confidence interval 1.384 to 1.711), and the ordering held in a raw-coordinate sensitivity arm. Expert manual labels exist for 58 external CT cases, but our registered reference is close enough to score only five structures. There the closed-form estimator again had lower average surface distance, 95th-percentile Hausdorff distance, and Chamfer error for both completed atria. On a second public benchmark of 20 cases the reference was close enough for three of four completed structures, and the same ordering held there. Four structures have no expert reference. The released model and its completion operator support cohort-unification research on aligned CT, not clinical use.

Authors · Matej Gazda, Jakub Gazda, Juraj Gazda, Peter Drotar

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB

Figures: Matej Gazda et al., arXiv:2608.19932, arxiv-nonexclusive